Inter prediction method and terminal
Patent Information
- Application Number
- CN202210400128.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-15
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2042-04-15
AI Technical Summary
[0004]本申请实施例提供一种帧间预测方法及终端,能够解决对边界像素点预测值修正不准确,进而降低视频编解码效率的技术问题
[0028] In this embodiment, target information is obtained, including the prediction value derivation mode corresponding to the target image frame and/or the prediction value derivation mode corresponding to each first image block in the target image frame. Based on the target information, inter-frame prediction is performed on each first image block. That is, for any first image block in a target image frame, inter-frame prediction is performed on the first image block using the corresponding prediction value derivation mode, thereby improving the accuracy of the predicted values obtained after inter-frame prediction of the image block, and thus improving the video encoding and decoding efficiency.
Smart Images

Figure CN116962686B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of video encoding and decoding technology, specifically relating to an inter-frame prediction method and terminal. Background Technology
[0002] Currently, in the process of video encoding and decoding, all image blocks in the image frame to be encoded or decoded use the same prediction value derivation mode for inter-frame prediction, thereby correcting the prediction values of the boundary pixels of the image blocks.
[0003] However, the motion information of the boundary pixels of different image blocks may be different. Using the same prediction value derivation mode for all image blocks in an image frame may lead to inaccurate prediction values of the corrected boundary pixels, thereby reducing the efficiency of video encoding and decoding. Summary of the Invention
[0004] This application provides an inter-frame prediction method and terminal that can solve the technical problem of inaccurate correction of predicted values of boundary pixels, thereby reducing the efficiency of video encoding and decoding.
[0005] Firstly, an inter-frame prediction method is provided, which includes:
[0006] Obtain target information; perform inter-frame prediction for each first image block;
[0007] Wherein, the target image frame is an image frame to be encoded, and the first image block is an image block to be encoded; or the target image frame is an image frame to be decoded, and the first image block is an image block to be decoded.
[0008] Secondly, an inter-frame prediction method is provided, which includes:
[0009] Acquire first motion information of a first image block and second motion information of a second image block, wherein the first image block and the second image block are adjacent;
[0010] Based on the position information of the first image block and the first motion information, at least one first pixel region associated with the first image block is determined;
[0011] Based on the first motion information and the second motion information, determine the predicted value of each pixel in the first pixel region;
[0012] Based on the predicted value of each pixel, the target predicted value of each pixel in the second pixel region of the first image block is determined;
[0013] Wherein, the first image block is an image block to be encoded, and the second image block is an encoded image block; or, the first image block is an image block to be decoded, and the second image block is a decoded image block.
[0014] Thirdly, an inter-frame prediction device is provided, comprising:
[0015] An acquisition module is used to acquire target information; the target information includes the prediction value export mode corresponding to the target image frame and / or the prediction value export mode corresponding to each first image block in the target image frame.
[0016] The processing module is used to perform inter-frame prediction on each of the first image blocks based on the target information;
[0017] Wherein, the target image frame is an image frame to be encoded, and the first image block is an image block to be encoded; or the target image frame is an image frame to be decoded, and the first image block is an image block to be decoded.
[0018] Fourthly, an inter-frame prediction apparatus is provided, comprising:
[0019] The acquisition module is used to acquire first motion information of a first image block and second motion information of a second image block, wherein the first image block and the second image block are adjacent to each other.
[0020] The first determining module is configured to determine at least one first pixel region associated with the first image block based on the position information of the first image block and the first motion information.
[0021] The second determining module is used to determine the predicted value of each pixel in the first pixel region based on the first motion information and the second motion information.
[0022] The third determining module is used to determine the target predicted value of each pixel in the second pixel region of the first image block based on the predicted value of each pixel.
[0023] Wherein, the first image block is an image block to be encoded, and the second image block is an encoded image block; or, the first image block is an image block to be decoded, and the second image block is a decoded image block.
[0024] Fifthly, a terminal is provided, the terminal including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the first aspect, or implementing the steps of the method as described in the second aspect.
[0025] In a sixth aspect, a readable storage medium is provided, on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect, or implement the steps of the method described in the second aspect.
[0026] In a seventh aspect, a chip is provided, the chip including a processor and a communication interface coupled to the processor, the processor being configured to run a program or instructions to implement the method as described in the first aspect, or to implement the steps of the method as described in the second aspect.
[0027] Eighthly, a computer program / program product is provided, which is stored in a storage medium and is executed by at least one processor to implement the steps of the method as described in the first aspect, or to implement the steps of the method as described in the second aspect.
[0028] In this embodiment, target information is obtained, including the prediction value derivation mode corresponding to the target image frame and / or the prediction value derivation mode corresponding to each first image block in the target image frame. Based on the target information, inter-frame prediction is performed on each first image block. That is, for any first image block in a target image frame, inter-frame prediction is performed on the first image block using the corresponding prediction value derivation mode, thereby improving the accuracy of the predicted values obtained after inter-frame prediction of the image block, and thus improving the video encoding and decoding efficiency. Attached Figure Description
[0029] Figure 1 This is a flowchart of an inter-frame prediction method provided in an embodiment of this application;
[0030] Figure 2 This is one of the application scenario diagrams of an inter-frame prediction method provided in the embodiments of this application;
[0031] Figure 3 This is a second schematic diagram illustrating an application scenario of an inter-frame prediction method provided in this application embodiment;
[0032] Figure 4 This is the third schematic diagram of an application scenario of an inter-frame prediction method provided in this application embodiment;
[0033] Figure 5 This is the fourth illustration of an application scenario for an inter-frame prediction method provided in this application embodiment;
[0034] Figure 6 This is the fifth illustration of an application scenario of an inter-frame prediction method provided in this application embodiment;
[0035] Figure 7This is one of the application scenarios of existing inter-frame prediction methods;
[0036] Figure 8 This is the second illustration of an application scenario for existing inter-frame prediction methods;
[0037] Figure 9 This is a flowchart of another inter-frame prediction method provided in the embodiments of this application;
[0038] Figure 10 This is one of the application scenario diagrams of another inter-frame prediction method provided in the embodiments of this application;
[0039] Figure 11 This is a second schematic diagram illustrating an application scenario of another inter-frame prediction method provided in the embodiments of this application;
[0040] Figure 12 This is a structural diagram of the inter-frame prediction device provided in the embodiments of this application;
[0041] Figure 13 This is a structural diagram of another inter-frame prediction device provided in an embodiment of this application;
[0042] Figure 14 This is a structural diagram of the communication device provided in the embodiments of this application;
[0043] Figure 15 This is a schematic diagram of the hardware structure of the terminal provided in the embodiments of this application. Detailed Implementation
[0044] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0045] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0046] In the embodiments of this application, the attribute decoding device corresponding to the inter-frame prediction method can be a terminal, which can also be called a terminal device or user equipment (UE). The terminal can be a mobile phone, tablet computer, laptop computer, personal digital assistant (PDA), handheld computer, netbook, ultra-mobile personal computer (UMPC), mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device or vehicle-mounted device (VUE), pedestrian terminal (PUE), smart home (home devices with wireless communication functions, such as refrigerators, televisions, washing machines or furniture, etc.), game console, personal computer (PUE). Terminal devices such as computers (PCs), ATMs, or self-service machines; wearable devices include smartwatches, smart bracelets, smart earphones, smart glasses, smart jewelry (smart bracelets, smart necklaces, smart anklets, smart ankle chains, etc.), smart wristbands, smart clothing, etc. It should be noted that the embodiments in this application do not limit the specific type of terminal.
[0047] This application provides an inter-frame prediction method. The inter-frame prediction method provided by this application will be described in detail below with reference to the accompanying drawings, through some embodiments and application scenarios.
[0048] Please see Figure 1 , Figure 1 This is a flowchart of an inter-frame prediction method provided in this application. The inter-frame prediction method provided in this embodiment includes the following steps:
[0049] S101, Obtain target information.
[0050] The inter-frame prediction method provided in this embodiment can be applied to either the encoding or decoding end. For the purpose of explaining the technical solution in detail, the following examples will be based on the application of this inter-frame prediction method to the decoding end.
[0051] In this step, the decoding end obtains target information from the bitstream. This target information includes the prediction value derivation mode corresponding to the target image frame and / or the prediction value derivation mode corresponding to each first image block within the target image frame. The target image frame is the image frame to be encoded, and the first image block is the image block to be encoded; alternatively, the target image frame is the image frame to be decoded, and the first image block is the image block to be decoded. Optionally, a first enable flag corresponding to the target image frame can be set to characterize the prediction value derivation mode corresponding to the target image frame, and a second enable flag corresponding to the first image block can be set to characterize the prediction value derivation mode corresponding to that first image block.
[0052] S102, based on the target information, perform inter-frame prediction for each first image block.
[0053] In this step, after obtaining the target information, the prediction value export mode corresponding to each first image block in the target image frame can be determined based on the target information.
[0054] In one optional implementation, a prediction value derivation mode corresponding to each first image block in the target image frame is determined based on a first enable flag corresponding to the target image frame. In this case, the prediction value derivation mode is the same for each first image block.
[0055] Another alternative implementation is to determine the prediction value derivation mode corresponding to each first image block based on the first enable identifier corresponding to the target image frame and the second enable identifier corresponding to each first image block.
[0056] Another alternative implementation is to determine the prediction value derivation mode corresponding to each first image block based on the second enable flag corresponding to each first image block.
[0057] It should be understood that for the specific technical solution of determining the predicted value derivation mode corresponding to each first image block in the target image frame based on the target information, please refer to the following embodiments.
[0058] In this step, after determining the prediction value export mode corresponding to each first image block, inter-frame prediction is performed on each first image block using the prediction value export mode corresponding to the first image block, so as to obtain the target prediction value of the first image block.
[0059] In this embodiment, target information is obtained, including the prediction value derivation mode corresponding to the target image frame and / or the prediction value derivation mode corresponding to each first image block in the target image frame. Based on the target information, inter-frame prediction is performed on each first image block. That is, for any first image block in a target image frame, inter-frame prediction is performed on the first image block using the corresponding prediction value derivation mode, thereby improving the accuracy of the predicted values obtained after inter-frame prediction of the image block, and thus improving the video encoding and decoding efficiency.
[0060] The following section details how to perform inter-frame prediction for each first image patch based on the target information, provided that the target information includes the predicted value derivation mode corresponding to the target image frame:
[0061] Optionally, the step of performing inter-frame prediction for each first image patch based on the target information includes at least one of the following:
[0062] When the prediction value export mode corresponding to the target image frame is the first export mode, the prediction value export mode corresponding to each first image block in the target image frame is determined to be the first export mode, and the first export mode is used to perform inter-frame prediction on each first image block.
[0063] When the prediction value export mode corresponding to the target image frame is the second export mode, the inter-frame prediction value export method of each first image block in the target image frame is determined to be the second export mode, and the second export mode is used to perform inter-frame prediction on each first image block;
[0064] When the prediction value export mode corresponding to the target image frame is the third export mode, the inter-frame prediction value export method of each first image block in the target image frame is determined to be the third export mode, and the third export mode is used to perform inter-frame prediction on each first image block.
[0065] In this embodiment, the prediction value export mode corresponding to each first image block in the target image frame can be determined solely based on the prediction value export mode corresponding to the target image frame, i.e., the first enable flag.
[0066] In this embodiment, when the prediction value export mode corresponding to the target image frame is the first export mode, the prediction value export mode corresponding to each first image block is determined to be the first export mode; when the prediction value export mode corresponding to the target image frame is the second export mode, the prediction value export mode corresponding to each first image block is determined to be the second export mode; when the prediction value export mode corresponding to the target image frame is the third export mode, the prediction value export mode corresponding to each first image block is determined to be the third export mode.
[0067] In this embodiment, after determining the prediction value export mode corresponding to each first image block, the prediction value export mode is used to perform inter-frame prediction on the first image block.
[0068] For example, setting the first enable flag to 0 indicates the first export mode, setting it to 1 indicates the second export mode, and setting it to 2 indicates the third export mode. In this case, if the first enable flag corresponding to the target image frame is 0, then the predicted value export mode corresponding to each first image block is determined to be the first export mode; if the first enable flag corresponding to the target image frame is 1, then the predicted value export mode corresponding to each first image block is determined to be the second export mode; and if the first enable flag corresponding to the target image frame is 2, then the predicted value export mode corresponding to each first image block is determined to be the third export mode.
[0069] For example, setting the first enable flag to 0 indicates the third export mode, and setting the first enable flag to 1 indicates the first export mode. In this case, if the first enable flag corresponding to the target image frame is 0, then the predicted value export mode corresponding to each first image block is determined to be the third export mode; if the first enable flag corresponding to the target image frame is 1, then the predicted value export mode corresponding to each first image block is determined to be the first export mode.
[0070] The first export mode is a prediction value export mode determined based on the motion information corresponding to each first image block, the position information corresponding to each first image block, and the motion information corresponding to the adjacent blocks of each first image block. For a detailed definition of the first export mode, please refer to the following embodiments.
[0071] The second export mode is a prediction value export mode determined based on the motion information corresponding to each first image block, the preset pixel region corresponding to each first image block, and the motion information corresponding to the adjacent blocks of each first image block. The difference between the second export mode and the first export mode is that the pixel region used in the second export mode is a preset pixel region, while the pixel region used in the first export mode is a pixel region determined based on the position information corresponding to the first image block and the motion information of the adjacent blocks, and / or the position information corresponding to the first image block and the motion information of the first image block. This will not be elaborated further here.
[0072] The third derivation mode is a prediction value derivation mode determined based on the motion information corresponding to each first image block. The third derivation mode does not apply the motion information of adjacent blocks. The third derivation mode is the Overlapped Block Motion Compensation (OBMC) technique. For the specific definition of OBMC, please refer to the following embodiments.
[0073] It should be understood that for regions containing sharp textures in an image, such as text areas, using OBMC for inter-frame prediction in these regions may result in texture blurring, thereby degrading image quality. Therefore, for the first image patch located within a region containing sharp textures, the prediction value derivation mode corresponding to this first image patch is determined to be the third derivation mode, i.e., OBMC technology is not used for inter-frame prediction of this first image patch. This avoids texture blurring after inter-frame prediction, thereby improving image quality.
[0074] The following section elaborates on how to perform inter-frame prediction for each first image patch in the target image frame based on the target information, which includes the predicted value derivation mode corresponding to each first image patch in the target image frame:
[0075] Optionally, the step of performing inter-frame prediction for each first image patch in the target image frame based on the target information includes at least one of the following:
[0076] When the prediction value export mode corresponding to any first image block is the first export mode, the first export mode is used to perform inter-frame prediction on the first image block.
[0077] When the prediction value export mode corresponding to any first image block is the second export mode, the second export mode is used to perform inter-frame prediction on the first image block.
[0078] When the prediction value export mode corresponding to any first image block is the third export mode, the first image block is predicted inter-frame using the third export mode.
[0079] In this embodiment, the prediction value export mode corresponding to each first image block in the target image frame can be determined solely based on the prediction value export mode corresponding to each first image block in the target image frame, i.e., the second enable flag.
[0080] In this embodiment, if the predicted value export mode corresponding to the first image block is the first export mode, the predicted value export mode corresponding to the first image block is determined to be the first export mode; if the predicted value export mode corresponding to the first image block is the second export mode, the predicted value export mode corresponding to the first image block is determined to be the second export mode; if the predicted value export mode corresponding to the first image block is the third export mode, the predicted value export mode corresponding to the first image block is determined to be the third export mode.
[0081] In this embodiment, after determining the prediction value export mode corresponding to each first image block, the prediction value export mode is used to perform inter-frame prediction on the first image block.
[0082] For example, setting the second enable flag to 0 indicates the first export mode, setting it to 1 indicates the second export mode, and setting it to 2 indicates the third export mode. In this case, if the second enable flag corresponding to the first image patch is 0, then the predicted value export mode corresponding to the first image patch is determined to be the first export mode; if the second enable flag corresponding to the first image patch is 1, then the predicted value export mode corresponding to the first image patch is determined to be the second export mode; and if the second enable flag corresponding to the first image patch is 2, then the predicted value export mode corresponding to the first image patch is determined to be the third export mode.
[0083] For example, setting the second enable flag to 0 indicates the third export mode, and setting the second enable flag to 1 indicates the first export mode. In this case, if the second enable flag corresponding to the first image block is 0, then the predicted value export mode corresponding to the first image block is determined to be the third export mode; if the second enable flag corresponding to the first image block is 1, then the predicted value export mode corresponding to the first image block is determined to be the first export mode.
[0084] For example, setting the second enable flag to 0 indicates the first export mode, and setting the second enable flag to 1 indicates the second export mode. In this case, if the second enable flag corresponding to the first image block is 0, then the predicted value export mode corresponding to the first image block is determined to be the first export mode; if the second enable flag corresponding to the first image block is 1, then the predicted value export mode corresponding to the first image block is determined to be the third export mode.
[0085] In this embodiment, by obtaining the prediction value export mode corresponding to each first image block, and then performing inter-frame prediction on each first image block according to the prediction value export mode corresponding to each first image block, multiple prediction value export modes can be used for inter-frame prediction of all first image blocks included in an image block, making the method of inter-frame prediction of image blocks more flexible and improving the accuracy of the prediction values of image blocks.
[0086] Optionally, the step of performing inter-frame prediction for each first image patch in the target image frame based on the target information includes:
[0087] When the prediction value export mode corresponding to the target image frame is the third export mode, the third export mode is used to perform inter-frame prediction for each first image block;
[0088] If the prediction value export mode corresponding to the target image frame is not the third export mode, then inter-frame prediction is performed on each first image block based on the prediction value export mode corresponding to each first image block in the target image frame.
[0089] In this embodiment, the prediction value export mode corresponding to each first image block can be determined according to the prediction value export mode corresponding to the target image frame and the prediction value export mode corresponding to each first image block in the target image frame. That is, the prediction value export mode corresponding to each first image block in the target image frame can be determined according to the first enable flag and the second enable flag.
[0090] In this embodiment, when the prediction value export mode corresponding to the target image frame is the third export mode, the prediction value export mode of each first image block in the target image frame is determined to be the third export mode; when the prediction value export mode corresponding to the target image frame is not the third export mode, the prediction value export mode corresponding to each first image block in the target image frame is obtained, thereby determining the prediction value export mode corresponding to each first image block.
[0091] In this embodiment, after determining the prediction value export mode corresponding to each first image block, the prediction value export mode is used to perform inter-frame prediction on the first image block.
[0092] For example, setting the first enable flag to 0 indicates the third export mode, setting the first enable flag to 1 indicates a non-third export mode, setting the second enable flag to 0 indicates the first export mode, setting the second enable flag to 1 indicates the second export mode, and setting the third enable flag to 2 indicates the third export mode. In this case, if the first enable flag is 0, the predicted value export mode corresponding to each first image block in the target image frame is determined to be the third export mode; if the first enable flag is 1 and the second enable flag is 0, the predicted value export mode corresponding to each first image block is determined to be the first export mode; if the first enable flag is 1 and the second enable flag is 1, the predicted value export mode corresponding to each first image block is determined to be the second export mode.
[0093] For example, setting the first enable flag to 0 indicates the third export mode, and setting the first enable flag to 1 indicates a non-third export mode. Setting the second enable flag to 0 indicates the third export mode, and setting the second enable flag to 1 indicates the first export mode. In this case, if the first enable flag is 0, the predicted value export mode corresponding to each first image block in the target image frame is determined to be the third export mode; if the first enable flag is 1 and the second enable flag is 0, the predicted value export mode corresponding to each first image block is determined to be the third export mode; if the first enable flag is 1 and the second enable flag is 1, the predicted value export mode corresponding to each first image block is determined to be the first export mode.
[0094] In this embodiment, the prediction value export mode corresponding to the target image frame is first determined. If the prediction value export mode corresponding to the target image frame is not the third export mode, the second enable flag corresponding to each first image block is further obtained, thereby reducing the bitstream in the inter-frame prediction process.
[0095] Optionally, when the prediction value derivation mode corresponding to each first image patch in the target image frame is a first derivation mode, the step of performing inter-frame prediction for each first image patch based on the target information includes:
[0096] Acquire the first motion information of the first image block and the second motion information of the second image block;
[0097] Based on the position information of the first image block and the first motion information, at least one first pixel region associated with the first image block is determined;
[0098] Based on the first motion information and the second motion information, determine the predicted value of each pixel in the first pixel region;
[0099] Based on the predicted value of each pixel, the target predicted value of each pixel in the second pixel region of the first image block is determined.
[0100] The first image block is an image block to be encoded, and the second image block is an image block adjacent to the first image block that has already been encoded; or, the first image block is an image block to be decoded, and the second image block is an image block adjacent to the first image block that has already been decoded. The first image block and the second image block satisfy any one of the following conditions:
[0101] 1. The prediction direction of the first image block is different from that of the second image block.
[0102] 2. The prediction direction of the first image block is the same as that of the second image block, but the reference frames to which the prediction directions point are different.
[0103] 3. The prediction direction of the first image block is the same as that of the second image block, and the reference frame to which the prediction direction points is the same, but the motion vector of the first image block is different from that of the second image block.
[0104] In this step, if the first image block and the second image block meet the above conditions, the first motion information of the first image block and the second motion information of the second image block are obtained.
[0105] The aforementioned motion information includes the predicted direction and motion vector. In this step, based on the location of the first image block, prediction can be performed using the motion vector and predicted direction to obtain at least one first pixel region associated with the first image block. For example... Figure 2 As shown, the aforementioned first pixel region can be a rectangular region located above and inside the first image block; or, as... Figure 3 As shown, the first pixel region can be a rectangular region adjacent to the upper edge of the first image block; or, as... Figure 4As shown, the first pixel region can be a rectangular region adjacent to the left of the first image block; or, the first pixel region can be a rectangular region inside the first image block on the left.
[0106] In this step, after determining the first pixel region, the predicted value of each pixel in the first pixel region is determined based on the first motion information and the second motion information. For details on how to determine the predicted value of each pixel in the first pixel region, please refer to subsequent embodiments. It should be understood that each pixel in the first pixel region includes at least two predicted values, one of which is determined based on the first motion information, and the other is determined based on the second motion information.
[0107] The aforementioned second pixel region is a portion of the pixel region within the first image block, and each pixel within this second pixel region is also referred to as a boundary pixel. For easier understanding, please refer to [link to relevant documentation]. Figure 5 and Figure 6 , Figure 5 The diagram shows the position of the second pixel region when the first pixel region is a rectangular region above the interior of the first image block, or when the first pixel region is a rectangular region adjacent to the upper edge of the first image block. Figure 6 The diagram shows the position of the second pixel region when the first pixel region is a rectangular region inside the first image block on the left, or when the first pixel region is a rectangular region adjacent to the left of the first image block.
[0108] In this step, after determining the predicted value of each pixel in the first pixel region, the target predicted value of each pixel in the second pixel region is determined based on the predicted value of each pixel. For specific technical solutions, please refer to the following embodiments.
[0109] In this embodiment, first motion information of a first image block and second motion information of a second image block are obtained; based on the position information and first motion information of the first image block, at least one first pixel region associated with the first image block is determined; according to the first motion information and second motion information, a predicted value for each pixel in the first pixel region is determined; and based on the predicted value for each pixel, a target predicted value for each pixel in the second pixel region of the first image block is determined. In this embodiment, the predicted value for each pixel in the first pixel region is determined based on the first motion information and second motion information. This fully considers the motion differences between the first and second image blocks during the correction of the predicted values of boundary pixels, improving the accuracy of the corrected predicted values of boundary pixels, and thus improving video encoding and decoding efficiency.
[0110] Optionally, determining the predicted value of each pixel in the first pixel region based on the first motion information and the second motion information includes:
[0111] For any first pixel in the first sub-pixel region, the reconstructed value of the first reference pixel corresponding to the first pixel in the first reference frame is determined as the first predicted value of the first pixel.
[0112] The reconstructed value of the second reference pixel corresponding to the first pixel in the second reference frame is determined as the second predicted value of the first pixel.
[0113] In this embodiment, the first pixel region includes a first sub-pixel region. The first sub-pixel region is either a rectangular region above the interior of the first image block, or a rectangular region to the left of the interior of the first image block, or a rectangular region adjacent to the top edge of the first image block, or a rectangular region adjacent to the left edge of the first image block, or a rectangular region formed by the top edge of the first image block and the upper interior of the first image block, or a rectangular region formed by the left edge of the first image block and the left interior of the first image block. In this embodiment, the pixels included in the first sub-pixel region are referred to as first sub-pixels.
[0114] Motion information includes prediction direction, reference frame information, and motion vector. In this embodiment, a first reference frame and a first reference pixel can be determined based on the first motion vector and the first reference frame information. The first reference pixel is a reconstructed pixel located in the first reference frame at the same position as the first sub-pixel region. Then, based on the first reference pixel, the reconstructed value of the pixel in the first reference frame pointed to by the first motion vector is determined as the first prediction value according to the first prediction direction.
[0115] The second reference frame and the second reference pixel can be determined based on the second motion vector and the second reference frame information. The second reference pixel is a reconstructed pixel located in the second reference frame at the same position as the first sub-pixel region. Then, based on the second reference pixel, the reconstructed value of the pixel in the second reference frame pointed to by the second motion vector is determined as the second predicted value according to the second prediction direction.
[0116] Optionally, determining the target predicted value of each pixel in the second pixel region of the first image block based on the predicted value of each pixel includes:
[0117] Calculate the first difference value between the first predicted value and the second predicted value for each first pixel in the first sub-pixel region;
[0118] If the first difference value is greater than the first preset threshold, the target prediction value of each pixel in the second pixel region is determined based on the position information of each pixel in the second pixel region and the first motion information.
[0119] If the first difference value is less than or equal to the first preset threshold, the third and fourth predicted values of each pixel in the second pixel region are weighted and summed based on the preset first weight value combination to obtain the target predicted value of each pixel in the second pixel region.
[0120] In this embodiment, after obtaining the first predicted value and the second predicted value for each first pixel, a first difference value between the first predicted value and the second predicted value is calculated. Optionally, the first difference value may be the absolute value of the difference between the first predicted value and the second predicted value, or the average value between the first predicted value and the second predicted value, or the variance between the first predicted value and the second predicted value, or the root mean square error between the first predicted value and the second predicted value; no specific limitation is imposed here.
[0121] In this embodiment, each pixel in the second pixel region includes a third predicted value and a fourth predicted value. The third predicted value is determined based on the first motion information, and the fourth predicted value is determined based on the second motion information. It should be understood that the method for determining the third predicted value is the same as the method for determining the first predicted value in the above embodiment, and the method for determining the fourth predicted value is the same as the method for determining the second predicted value in the above embodiment, and will not be repeated here.
[0122] In this embodiment, a first preset threshold is also preset. After obtaining the first difference value, for each first pixel, the relationship between the first difference value and the first preset threshold is compared. If the first difference value is greater than the first preset threshold, then for each pixel in the second pixel region, the target predicted value of the pixel is determined based on the pixel's position information and the first motion information. Specifically, taking the pixel as the motion starting point, the motion starting point is offset according to the first motion vector and the first prediction direction in the first motion information to determine the offset pixel. The offset pixel is located in the first reference frame. Further, the reconstructed value of the offset pixel is determined as the target predicted value of the pixel. It should be understood that when the first pixel region is located inside the first image block, the first predicted value of each first pixel in the first pixel region is determined as the target predicted value of the pixel corresponding to the first pixel in the second pixel region.
[0123] If the first difference value is less than or equal to the first preset threshold, the third and fourth predicted values of each pixel in the second pixel region can be weighted and summed using the first weight value combination to obtain the target predicted value of each pixel in the second pixel region. It should be understood that the first weight value combination includes at least one weight recombination, which includes a first weight value and a second weight value. The first weight value corresponds to the third predicted value, and the second weight value corresponds to the fourth predicted value.
[0124] Specifically, the target prediction value for each pixel in the second pixel region can be calculated using the following formula:
[0125] shift = log2(w11 + w12)
[0126] offset = (w11 + w12) / 2
[0127] Pixel(i,j)=(w11×Pixel3(i,j)+w12×Pixel4(i,j)+offset)>>shift
[0128] In this context, Pixel represents the target predicted value, w11 represents the first weight value, w12 represents the second weight value, Pixel3 represents the third predicted value, and Pixel4 represents the fourth predicted value.
[0129] It should be understood that the first image block can be a luminance block or a chrominance block, and the first weight value combination corresponding to the first image block being a luminance block may be different from the first weight value combination corresponding to the first image block being a chrominance block.
[0130] Optionally, determining the predicted value of each pixel in the first pixel region based on the first motion information and the second motion information includes:
[0131] For any second pixel in the second sub-pixel region, the reconstructed value of the third reference pixel corresponding to the second pixel in the third reference frame is determined as the fifth predicted value of the second pixel.
[0132] The reconstructed value of the fourth reference pixel corresponding to the second pixel in the fourth reference frame is determined as the sixth predicted value of the second pixel.
[0133] For any third pixel in the third sub-pixel region, the reconstructed value of the fifth reference pixel corresponding to the third pixel in the fifth reference frame is determined as the seventh predicted value of the third pixel.
[0134] The reconstructed value of the sixth reference pixel corresponding to the third pixel in the sixth reference frame is determined as the eighth predicted value of the third pixel.
[0135] In this embodiment, the first pixel region includes a second sub-pixel region and a third sub-pixel region. The second sub-pixel region is either a rectangular region located above the first image block or a rectangular region located to the left of the first image block. The third sub-pixel region is either a rectangular region adjacent to the top edge of the first image block or a rectangular region adjacent to the left edge of the first image block. In this embodiment, the pixels included in the second sub-pixel region are referred to as second pixels, and the pixels included in the third sub-pixel region are referred to as third pixels.
[0136] As described above, the first motion information includes a first prediction direction, first reference frame information, and a first motion vector. In this embodiment, a third reference frame and a third reference pixel can be determined based on the first motion information. Then, based on the third reference pixel, the reconstructed value of the pixel in the third reference frame pointed to by the first motion vector is determined as the fifth prediction value according to the first prediction direction. It should be understood that the specific technical solution for determining the fifth prediction value of the second pixel in this embodiment is the same as the technical solution for determining the first prediction value of the first pixel described above, and will not be repeated here.
[0137] As described above, the second motion information includes a second prediction direction, second reference frame information, and a second motion vector. In this embodiment, a fourth reference frame and a fourth reference pixel can be determined based on the second motion information. Then, based on the fourth reference pixel, the reconstructed value of the pixel in the fourth reference frame pointed to by the second motion vector is determined as the sixth prediction value according to the second prediction direction. It should be understood that the specific technical solution for determining the sixth prediction value of the second pixel in this embodiment is the same as the technical solution for determining the second prediction value of the second pixel described above, and will not be repeated here.
[0138] In this embodiment, the same technical solution used to determine the fifth predicted value of the second pixel can be used to determine the seventh predicted value of the third pixel, which will not be repeated here.
[0139] In this embodiment, the same technical solution used to determine the sixth predicted value of the second pixel can be used to determine the eighth predicted value of the third pixel, which will not be repeated here.
[0140] Optionally, determining the target predicted value of each pixel in the second pixel region of the first image block based on the predicted value of each pixel includes:
[0141] Based on the fifth, sixth, seventh, and eighth predicted values, determine the second and third difference values for each target pixel.
[0142] When the second difference value and the third difference value meet the preset conditions, the target prediction value of each pixel in the second pixel region is determined based on the position information of each pixel in the second pixel region and the first motion information;
[0143] If the second difference value and the third difference value do not meet the preset conditions, the ninth and tenth predicted values of each pixel in the second pixel region are weighted and summed based on the preset second weight value combination to obtain the target predicted value of each pixel in the second pixel region.
[0144] In this embodiment, after obtaining the fifth and sixth predicted values for each second pixel, and the seventh and eighth predicted values for each third pixel, the second and third difference values of the target pixel can be determined. The target pixel includes both the second and third pixels, thus enabling the determination of the second and third difference values of both the second and third pixels. For a detailed explanation of the technical solution for determining the second and third difference values of the target pixel, please refer to subsequent embodiments.
[0145] It should be understood that each pixel in the second pixel region includes a ninth predicted value and a tenth predicted value, wherein the ninth predicted value is determined based on the first motion information, and the tenth predicted value is determined based on the second motion information. It should be understood that the method for determining the ninth predicted value is the same as the method for determining the first predicted value in the above embodiments, and the method for determining the tenth predicted value is the same as the method for determining the second predicted value in the above embodiments, and will not be repeated here.
[0146] In this embodiment, preset conditions are established. If the second difference value and the third difference value meet the preset conditions, it indicates that the motion pattern of the boundary pixels in the first image block is more similar to that of the first image block. Then, for each pixel in the second pixel region, the target prediction value of the pixel is determined based on the pixel's position information and the first motion information. For specific implementation methods, please refer to the above embodiments, which will not be repeated here.
[0147] If the second and third difference values do not meet the preset conditions, the ninth and tenth predicted values of each pixel in the second pixel region can be weighted and summed using the second weight value combination to obtain the target predicted value of each pixel in the second pixel region. It should be understood that the second weight value combination includes at least one weighted recombination, which includes a third weight value and a fourth weight value. The third weight value corresponds to the aforementioned ninth predicted value, and the fourth weight value corresponds to the aforementioned tenth predicted value. Optionally, the second weight value combination can be the same as the first weight value combination.
[0148] The specific implementation method of weighted summation is the same as the implementation method described above that uses the combination of the first weight values to perform weighted summation on the third and fourth predicted values, and will not be repeated here.
[0149] Optionally, determining the second difference value and the third difference value based on the fifth predicted value, the sixth predicted value, the seventh predicted value, and the eighth predicted value includes:
[0150] The difference between the fifth predicted value and the sixth predicted value is determined as the second difference value, and the difference between the seventh predicted value and the eighth predicted value is determined as the third difference value, or;
[0151] The difference between the fifth predicted value and the seventh predicted value is determined as the second difference value, and the difference between the sixth predicted value and the eighth predicted value is determined as the third difference value.
[0152] In one optional implementation, the difference between the fifth and sixth predicted values is determined as the second difference value, and the difference between the seventh and eighth predicted values is determined as the third difference value.
[0153] In this implementation, the second difference value can be the absolute value of the difference between the fifth predicted value and the sixth predicted value, or the average value between the fifth predicted value and the sixth predicted value, or the variance between the fifth predicted value and the sixth predicted value, or the root mean square error between the fifth predicted value and the sixth predicted value, without any specific limitation.
[0154] The aforementioned third difference value can be the absolute value of the difference between the seventh and eighth predicted values, or the average value between the seventh and eighth predicted values, or the variance between the seventh and eighth predicted values, or the root mean square error between the seventh and eighth predicted values; no specific restrictions are imposed here.
[0155] Another alternative implementation is to determine the difference between the fifth and seventh predicted values as the second difference value, and the difference between the sixth and eighth predicted values as the third difference value.
[0156] In this implementation, the second difference value can be the absolute value of the difference between the fifth predicted value and the seventh predicted value, or the average value between the fifth predicted value and the seventh predicted value, or the variance between the fifth predicted value and the seventh predicted value, or the root mean square error between the fifth predicted value and the seventh predicted value, without any specific limitation.
[0157] The aforementioned third difference value can be the absolute value of the difference between the sixth and eighth predicted values, or the average value between the sixth and eighth predicted values, or the variance between the sixth and eighth predicted values, or the root mean square error between the sixth and eighth predicted values; no specific restrictions are imposed here.
[0158] Optionally, the preset conditions include any one of the following:
[0159] The second difference value is smaller than the third difference value;
[0160] The ratio between the second difference value and the third difference value is less than the second preset threshold;
[0161] The ratio between the third difference value and the second difference value is less than the second preset threshold;
[0162] The second difference value is greater than the third preset threshold, and the ratio between the second difference value and the third difference value is less than the second preset threshold;
[0163] The second difference value is greater than the third preset threshold, and the ratio between the third difference value and the second difference value is less than the second preset threshold.
[0164] One implementation method in this embodiment involves first comparing the magnitude of a second difference value with a third preset threshold. If the second difference value is greater than the third preset threshold, then comparing the ratio of the second difference value to the third difference value with the second preset threshold, or comparing the ratio of the third difference value to the second difference value with the second preset threshold. By first comparing the magnitude of the second difference value with the third preset threshold, the motion differences between the first image block and the second image block are more fully considered during the correction of the predicted values of boundary pixels.
[0165] Optionally, the first pixel region satisfies at least one of the following:
[0166] The first pixel region is the encoded or decoded pixel region consisting of M1 rows and N1 columns of pixels adjacent to the top edge of the first image block;
[0167] The first pixel region is the encoded or decoded pixel region consisting of M2 rows and N2 columns of pixels adjacent to the left of the first image block;
[0168] The first pixel region is an uncoded or undecoded pixel region consisting of M3 rows and N3 columns of pixels located inside and above the first image block;
[0169] The first pixel region is an uncoded or undecoded pixel region consisting of M4 rows and N4 columns of pixels located inside the left side of the first image block;
[0170] The first pixel region is an M5-row, N5-column pixel region consisting of the encoded or decoded pixel region adjacent to the top of the first image block and the unencoded or undecoded pixel region inside the top of the first image block;
[0171] The first pixel region is an M6-row, N6-column pixel region consisting of the encoded or decoded pixel region adjacent to the left of the first image block and the unencoded or undecoded pixel region inside the left of the first image block;
[0172] Where M1, M2, M3, M4, M5, M6, N1, N2, N3, N4, N5, and N6 are all positive integers.
[0173] In one optional implementation, the first pixel region is an uncoded or undecoded pixel region consisting of M3 rows and N3 columns above the first image block. For clarity, please refer to [link to relevant documentation]. Figure 2 ,exist Figure 2 In the scene shown, the first pixel region is a rectangular area consisting of 1 row and 3 columns of pixels above the first image block.
[0174] Another alternative implementation is that the first pixel region is the encoded or decoded pixel region consisting of M1 rows and N1 columns adjacent to the top edge of the first image block. For easier understanding, please refer to [link to relevant documentation]. Figure 3 ,exist Figure 3 In the scene shown, the first pixel region is a rectangular region consisting of 1 row and 3 columns of pixels adjacent to the top edge of the first image block.
[0175] Another alternative implementation is that the first pixel region is the encoded or decoded pixel region consisting of M2 rows and N2 columns adjacent to the left of the first image block. For clarity, please refer to [link to relevant documentation]. Figure 4 ,exist Figure 4 In the scene shown, the first pixel region is a rectangular area consisting of 3 rows and 1 column adjacent to the left side inside the first image block.
[0176] Another alternative implementation is that the first pixel region can be an encoded or decoded pixel region composed of some pixels adjacent to the top of the first image block and an encoded or decoded pixel region composed of some pixels adjacent to the left of the first image block. In this case, the first pixel region is "L" shaped.
[0177] In this embodiment, by defining the positional relationship between the first pixel region and the first image block, the motion differences between the first image block and each image block adjacent to the first image block are fully considered, thereby improving the accuracy of the predicted value of the corrected boundary pixel point.
[0178] This application also provides an inter-frame prediction method. The inter-frame prediction method provided by this application will be described in detail below with reference to the accompanying drawings, through some embodiments and application scenarios. For ease of understanding, some contents related to the embodiments of this application are explained below:
[0179] When the boundary of an image patch does not fit the contour of the current image patch, the motion of the boundary pixels of the current image patch may be consistent with the current image patch or with adjacent image patches. The predicted values of the boundary pixels determined based on the motion information of the current image patch differ significantly from the actual predicted values, thereby reducing the encoding and decoding efficiency of the video. Here, the current image patch can be a block to be encoded, and adjacent image patches can be already encoded; or, the current image patch can be a block to be decoded, and adjacent image patches can be already decoded.
[0180] Currently, OBMC technology can be used to correct the predicted values of boundary pixels in the current image patch, thus solving the aforementioned technical problems. OBMC is an inter-frame prediction method. The following is a detailed explanation of OBMC technology:
[0181] The first scenario is that the inter-frame prediction modes of all pixels in the current block are the same.
[0182] In this scenario, when adjacent image blocks are in inter-frame prediction mode, not intra-frame block copy mode, and the motion mode of adjacent image blocks is inconsistent with the motion mode of the current image block, motion information of the adjacent image blocks is obtained. Please refer to [link to relevant documentation]. Figure 7 Adjacent image blocks can be the image block above the current image block, or the image block to the left of the current image block.
[0183] The motion pattern of an adjacent image block is inconsistent with the motion pattern of the current image block if any one of the following conditions is met.
[0184] 1. The prediction direction of adjacent image blocks is different from the prediction direction of the current image block.
[0185] 2. The prediction direction of adjacent image blocks is the same as that of the current image block, but the reference frames to which the prediction directions point are different.
[0186] 3. The prediction direction of adjacent image blocks is the same as that of the current image block, and the reference frame to which the prediction direction points is the same, but the motion vector of adjacent image blocks is different from that of the current image block.
[0187] After acquiring motion information from neighboring image blocks, a first predicted value is obtained based on the motion information of the current image block; a second predicted value is obtained based on the motion information of neighboring image blocks. The predicted values of the boundary pixels of the current image block are then corrected using the first and second predicted values.
[0188] Specifically, if the current image block is a brightness sub-block, the first and second predicted values can be weighted and summed using the following formula to obtain the predicted value after boundary pixel correction:
[0189] NewPixel(i,j)=(26×Pixel1(i,j)+6×Pixel2(i,j)+16)>>5
[0190] NewPixel(i,j)=(7×Pixel1(i,j)+Pixel2(i,j)+4)>>3
[0191] NewPixel(i,j)=(15×Pixel1(i,j)+Pixel2(i,j)+8)>>4
[0192] NewPixel(i,j)=(31×Pixel1(i,j)+Pixel2(i,j)+16)>>5
[0193] Where i represents the column coordinate of the boundary pixel in the current image block, j represents the row coordinate of the boundary pixel in the current image block, Pixel1 represents the first predicted value of the boundary pixel, Pixel2 represents the second predicted value of the boundary pixel, and NewPixel represents the corrected predicted value of the boundary pixel.
[0194] If the current image patch is a chroma sub-patch, the first and second predicted values can be weighted and summed using the following formula to obtain the predicted value after boundary pixel correction:
[0195] NewPixel(i,j)=(26×Pixel1(i,j)+6×Pixel2(i,j)+16)>>5
[0196] Among them, NewPixel represents the predicted value after correction of the boundary pixel.
[0197] It should be understood that the above formula is applicable to scenarios where the pixel region of the boundary pixel is 4 rows or 4 columns. In other application scenarios, there is no specific limitation on the pixel region of the boundary pixel.
[0198] The second scenario is if the current image block is a coded block and the inter-frame prediction mode is affine mode, or if the current image block is a decoded block and the inter-frame prediction mode is motion vector correction mode.
[0199] In this case, the motion information of the four adjacent image blocks—the top, bottom, left, and right sides—that are adjacent to the current image block are obtained. Please refer to [link / reference needed]. Figure 8 , Figure 8This illustrates the positional relationship between adjacent image blocks and the current image block under the above conditions.
[0200] Based on the motion information of the current image block, a first predicted value is obtained; if any of the following conditions are met between the current image block and its neighboring image blocks, a second predicted value is obtained based on the motion information of the neighboring image blocks.
[0201] 1. The prediction direction of adjacent image blocks is different from the prediction direction of the current image block.
[0202] 2. The prediction direction of adjacent image blocks is the same as that of the current image block, but the reference frames to which the prediction directions point are different.
[0203] 3. The prediction direction of adjacent image blocks is the same as that of the current image block, and the reference frame to which the prediction direction points is the same. However, the absolute value of the difference between the motion vector of the adjacent image block and the motion vector of the current image block is greater than a preset threshold.
[0204] The predicted values of the boundary pixels of the current image block are corrected using the first and second predicted values described above. Specifically, the corrected predicted values of the boundary pixels can be obtained by weighted summing the first and second predicted values using the following formula:
[0205] rem_w(i,j)=(32–w(i)–w(width-i)–w(j)–w(height-j))
[0206] subNewPixel(i,j)=(subPixel2 L (i,j)×w(i)+subPixel2 R (i,j)×w(width-1-i)+subPixel2 T (i,j)×w(j)+subPixel2 B (i,j)×w(height-1-j)+subPixel1×rem_w(i,j)+16)>>5
[0207] Where i represents the column coordinate of the boundary pixel in the current image patch, j represents the row coordinate of the boundary pixel in the current image patch, subNewPixel represents the corrected predicted value of the boundary pixel, and subPixel2 represents the corrected predicted value of the boundary pixel. L subPixel2 R subPixel2 T and subPixel2 BThe second predicted value is determined based on the motion information of adjacent image blocks. The width represents the number of columns of the adjacent image blocks, the height represents the number of rows of the adjacent image blocks, and w represents the preset weight combination. The weight combination corresponding to the current image block being a luminance block is different from the weight combination corresponding to the current image block being a chrominance block.
[0208] It should be understood that the above formula is applicable to scenarios where the pixel region of the boundary pixel is 4 rows or 4 columns. In other application scenarios, there is no specific limitation on the pixel region of the boundary pixel.
[0209] In the process of correcting the predicted values of boundary pixels using OBMC technology, the difference between the motion mode of the current image block and the motion mode of adjacent image blocks was not taken into account. This resulted in inaccurate predicted values of boundary pixels after correction, which in turn reduced the efficiency of video encoding and decoding.
[0210] Given the above situation, how to improve the accuracy of the predicted values of the corrected boundary pixels, and thus improve the efficiency of video encoding and decoding, is a technical problem that needs to be solved.
[0211] To address the aforementioned potential technical problems, this application provides an inter-frame prediction method. Please refer to [link to relevant documentation]. Figure 9 , Figure 9 This is a flowchart of another inter-frame prediction method provided in this application. The other inter-frame prediction coding method provided in this embodiment includes the following steps:
[0212] S901, acquire the first motion information of the first image block and the second motion information of the second image block.
[0213] S902, based on the position information of the first image block and the first motion information, determine at least one first pixel region associated with the first image block.
[0214] S903, based on the first motion information and the second motion information, determine the predicted value of each pixel in the first pixel region.
[0215] S904, based on the predicted value of each pixel, determine the target predicted value of each pixel in the second pixel region of the first image block.
[0216] It should be understood that, in other embodiments, the inter-frame prediction method provided in this application can also be used to generate prediction values for the boundary pixels of each sub-block within a coding block or decoding block. For this implementation, please refer to... Figure 10 , Figure 10 The sub-block shown can be understood as the first image block in this embodiment, and this sub-block is located inside the lower right corner of the encoding block or decoding block. Please refer to [link / reference]. Figure 11 , Figure 11 The sub-block shown can be understood as the first image block in this embodiment, and the sub-block is located in the middle of the encoding block or decoding block.
[0217] In this embodiment, first motion information of a first image block and second motion information of a second image block are obtained; based on the position information and first motion information of the first image block, at least one first pixel region associated with the first image block is determined; according to the first motion information and second motion information, a predicted value for each pixel in the first pixel region is determined; and based on the predicted value for each pixel, a target predicted value for each pixel in the second pixel region of the first image block is determined. In this embodiment, the predicted value for each pixel in the first pixel region is determined based on the first motion information and second motion information. This fully considers the motion differences between the first and second image blocks during the correction of the predicted values of boundary pixels, improving the accuracy of the corrected predicted values of boundary pixels, and thus improving video encoding and decoding efficiency.
[0218] Optionally, the first pixel region includes a first sub-pixel region, which is a portion of the pixel region of the first image block, or the first sub-pixel region is a portion of the pixel region of the second image block, or the first sub-pixel region is a region composed of a portion of the pixel region of the first image block and a portion of the pixel region of the second image block.
[0219] The step of determining the predicted value of each pixel in the first pixel region based on the first motion information and the second motion information includes:
[0220] For any first pixel in the first sub-pixel region, the reconstructed value of the first reference pixel corresponding to the first pixel in the first reference frame is determined as the first predicted value of the first pixel.
[0221] The reconstructed value of the second reference pixel corresponding to the first pixel in the second reference frame is determined as the second predicted value of the first pixel.
[0222] The first reference frame is determined based on the first motion information, and the second reference frame is determined based on the second motion information.
[0223] Optionally, determining the target predicted value of each pixel in the second pixel region of the first image block based on the predicted value of each pixel includes:
[0224] Calculate the first difference value between the first predicted value and the second predicted value for each first pixel in the first sub-pixel region;
[0225] If the first difference value is greater than the first preset threshold, the target prediction value of each pixel in the second pixel region is determined based on the position information of each pixel in the second pixel region and the first motion information.
[0226] When the first difference value is less than or equal to the first preset threshold, the third and fourth predicted values of each pixel in the second pixel region are weighted and summed based on the preset first weight value combination to obtain the target predicted value of each pixel in the second pixel region.
[0227] The first weight value combination includes at least one weight recombination, the weight recombination includes a first weight value and a second weight value, the first weight value corresponds to the third predicted value, the second weight value corresponds to the fourth predicted value, the third predicted value is determined based on the first motion information, and the fourth predicted value is determined based on the second motion information.
[0228] Optionally, the first pixel region includes a second sub-pixel region and a third sub-pixel region, wherein the second sub-pixel region is a portion of the pixel region of the first image block, and the third sub-pixel region is a portion of the pixel region of the second image block;
[0229] The step of determining the predicted value of each pixel in the first pixel region based on the first motion information and the second motion information includes:
[0230] For any second pixel in the second sub-pixel region, the reconstructed value of the third reference pixel corresponding to the second pixel in the third reference frame is determined as the fifth predicted value of the second pixel.
[0231] The reconstructed value of the fourth reference pixel corresponding to the second pixel in the fourth reference frame is determined as the sixth predicted value of the second pixel.
[0232] For any third pixel in the third sub-pixel region, the reconstructed value of the fifth reference pixel corresponding to the third pixel in the fifth reference frame is determined as the seventh predicted value of the third pixel.
[0233] The reconstructed value of the sixth reference pixel corresponding to the third pixel in the sixth reference frame is determined as the eighth predicted value of the third pixel.
[0234] The third and fifth reference frames are determined based on the first motion information, and the fourth and sixth reference frames are determined based on the second motion information.
[0235] Optionally, determining the target predicted value of each pixel in the second pixel region of the first image block based on the predicted value of each pixel includes:
[0236] Based on the fifth, sixth, seventh, and eighth predicted values, a second difference value and a third difference value are determined for each target pixel; the target pixel includes a second pixel and a third pixel.
[0237] When the second difference value and the third difference value meet the preset conditions, the target prediction value of each pixel in the second pixel region is determined based on the position information of each pixel in the second pixel region and the first motion information;
[0238] If the second difference value and the third difference value do not meet the preset conditions, the ninth prediction value and the tenth prediction value of each pixel in the second pixel region are weighted and summed based on the preset second weight value combination to obtain the target prediction value of each pixel in the second pixel region.
[0239] The second weight value combination includes at least one weight recombination, which includes a third weight value and a fourth weight value. The third weight value corresponds to the ninth predicted value, and the fourth weight value corresponds to the tenth predicted value. The ninth predicted value is determined based on the first motion information, and the tenth predicted value is determined based on the second motion information.
[0240] Optionally, determining the second difference value and the third difference value based on the fifth predicted value, the sixth predicted value, the seventh predicted value, and the eighth predicted value includes:
[0241] The difference between the fifth predicted value and the sixth predicted value is determined as the second difference value, and the difference between the seventh predicted value and the eighth predicted value is determined as the third difference value, or;
[0242] The difference between the fifth predicted value and the seventh predicted value is determined as the second difference value, and the difference between the sixth predicted value and the eighth predicted value is determined as the third difference value.
[0243] The inter-frame prediction method provided in this application can be executed by an inter-frame prediction device. This application uses an inter-frame prediction device executing the inter-frame prediction method as an example to illustrate the inter-frame prediction device provided in this application.
[0244] like Figure 12 As shown, the inter-frame prediction device 1200 includes:
[0245] Module 1201 is used to acquire target information;
[0246] The processing module 1202 is used to perform inter-frame prediction for each first image block based on the target information.
[0247] Optionally, the processing module 1202 is specifically used for:
[0248] When the prediction value export mode corresponding to the target image frame is the first export mode, the prediction value export mode corresponding to each first image block in the target image frame is determined to be the first export mode, and the first export mode is used to perform inter-frame prediction on each first image block.
[0249] When the prediction value export mode corresponding to the target image frame is the second export mode, the inter-frame prediction value export method of each first image block in the target image frame is determined to be the second export mode, and the second export mode is used to perform inter-frame prediction on each first image block;
[0250] When the prediction value export mode corresponding to the target image frame is the third export mode, the inter-frame prediction value export method of each first image block in the target image frame is determined to be the third export mode, and the third export mode is used to perform inter-frame prediction on each first image block.
[0251] Optionally, the processing module 1202 is further specifically used for:
[0252] When the prediction value export mode corresponding to any first image block is the first export mode, the first export mode is used to perform inter-frame prediction on the first image block.
[0253] When the prediction value export mode corresponding to any first image block is the second export mode, the second export mode is used to perform inter-frame prediction on the first image block.
[0254] When the prediction value export mode corresponding to any first image block is the third export mode, the first image block is predicted inter-frame using the third export mode.
[0255] Optionally, the processing module 1202 is further specifically used for:
[0256] When the prediction value export mode corresponding to the target image frame is the third export mode, the third export mode is used to perform inter-frame prediction for each first image block;
[0257] If the prediction value export mode corresponding to the target image frame is not the third export mode, the prediction value export mode corresponding to each first image block is determined based on the prediction value export mode corresponding to each first image block in the target image frame, and inter-frame prediction is performed on each first image block.
[0258] Optionally, the processing module 1202 is specifically used for:
[0259] Acquire the first motion information of the first image block and the second motion information of the second image block;
[0260] Based on the position information of the first image block and the first motion information, at least one first pixel region associated with the first image block is determined;
[0261] Based on the first motion information and the second motion information, determine the predicted value of each pixel in the first pixel region;
[0262] Based on the predicted value of each pixel, the target predicted value of each pixel in the second pixel region of the first image block is determined.
[0263] Optionally, the processing module 1202 is further specifically used for:
[0264] For any first pixel in the first sub-pixel region, the reconstructed value of the first reference pixel corresponding to the first pixel in the first reference frame is determined as the first predicted value of the first pixel.
[0265] The reconstructed value of the second reference pixel corresponding to the first pixel in the second reference frame is determined as the second predicted value of the first pixel.
[0266] Optionally, the processing module 1202 is further specifically used for:
[0267] Calculate the first difference value between the first predicted value and the second predicted value for each first pixel in the first sub-pixel region;
[0268] If the first difference value is greater than the first preset threshold, the target prediction value of each pixel in the second pixel region is determined based on the position information of each pixel in the second pixel region and the first motion information.
[0269] If the first difference value is less than or equal to the first preset threshold, the third and fourth predicted values of each pixel in the second pixel region are weighted and summed based on the preset first weight value combination to obtain the target predicted value of each pixel in the second pixel region.
[0270] Optionally, the processing module 1202 is further specifically used for:
[0271] For any second pixel in the second sub-pixel region, the reconstructed value of the third reference pixel corresponding to the second pixel in the third reference frame is determined as the fifth predicted value of the second pixel.
[0272] The reconstructed value of the fourth reference pixel corresponding to the second pixel in the fourth reference frame is determined as the sixth predicted value of the second pixel.
[0273] For any third pixel in the third sub-pixel region, the reconstructed value of the fifth reference pixel corresponding to the third pixel in the fifth reference frame is determined as the seventh predicted value of the third pixel.
[0274] The reconstructed value of the sixth reference pixel corresponding to the third pixel in the sixth reference frame is determined as the eighth predicted value of the third pixel.
[0275] Optionally, the processing module 1202 is further specifically used for:
[0276] Based on the fifth, sixth, seventh, and eighth predicted values, determine the second and third difference values for each target pixel.
[0277] When the second difference value and the third difference value meet the preset conditions, the target prediction value of each pixel in the second pixel region is determined based on the position information of each pixel in the second pixel region and the first motion information;
[0278] If the second difference value and the third difference value do not meet the preset conditions, the ninth and tenth predicted values of each pixel in the second pixel region are weighted and summed based on the preset second weight value combination to obtain the target predicted value of each pixel in the second pixel region.
[0279] Optionally, the processing module 1202 is further specifically used for:
[0280] The difference between the fifth predicted value and the sixth predicted value is determined as the second difference value, and the difference between the seventh predicted value and the eighth predicted value is determined as the third difference value, or;
[0281] The difference between the fifth predicted value and the seventh predicted value is determined as the second difference value, and the difference between the sixth predicted value and the eighth predicted value is determined as the third difference value.
[0282] In this embodiment, target information is obtained, including the prediction value derivation mode corresponding to the target image frame and / or the prediction value derivation mode corresponding to each first image block in the target image frame. Based on the target information, inter-frame prediction is performed on each first image block. That is, for any first image block in a target image frame, inter-frame prediction is performed on the first image block using the corresponding prediction value derivation mode, thereby improving the accuracy of the predicted values obtained after inter-frame prediction of the image block, and thus improving the video encoding and decoding efficiency.
[0283] The inter-frame prediction device provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method embodiments achieve the same technical effect, and will not be described again here to avoid repetition.
[0284] This application also provides an inter-frame prediction method, which can be executed by an inter-frame prediction device. This application uses an inter-frame prediction device executing the inter-frame prediction method as an example to illustrate the inter-frame prediction device provided in this application.
[0285] like Figure 13 As shown, the inter-frame prediction device 1300 includes:
[0286] The acquisition module 1301 is used to acquire the first motion information of the first image block and the second motion information of the second image block;
[0287] The first determining module 1302 is used to determine at least one first pixel region associated with the first image block based on the position information of the first image block and the first motion information.
[0288] The second determining module 1303 is used to determine the predicted value of each pixel in the first pixel region based on the first motion information and the second motion information.
[0289] The third determining module 1304 is used to determine the target predicted value of each pixel in the second pixel region of the first image block based on the predicted value of each pixel.
[0290] Optionally, the second determining module 1303 is specifically used for:
[0291] For any first pixel in the first sub-pixel region, the reconstructed value of the first reference pixel corresponding to the first pixel in the first reference frame is determined as the first predicted value of the first pixel.
[0292] The reconstructed value of the second reference pixel corresponding to the first pixel in the second reference frame is determined as the second predicted value of the first pixel.
[0293] Optionally, the third determining module 1304 is specifically used for:
[0294] Calculate the first difference value between the first predicted value and the second predicted value for each first pixel in the first sub-pixel region;
[0295] If the first difference value is greater than the first preset threshold, the target prediction value of each pixel in the second pixel region is determined based on the position information of each pixel in the second pixel region and the first motion information.
[0296] If the first difference value is less than or equal to the first preset threshold, the third and fourth predicted values of each pixel in the second pixel region are weighted and summed based on the preset first weight value combination to obtain the target predicted value of each pixel in the second pixel region.
[0297] Optionally, the second determining module 1303 is further specifically used for:
[0298] For any second pixel in the second sub-pixel region, the reconstructed value of the third reference pixel corresponding to the second pixel in the third reference frame is determined as the fifth predicted value of the second pixel.
[0299] The reconstructed value of the fourth reference pixel corresponding to the second pixel in the fourth reference frame is determined as the sixth predicted value of the second pixel.
[0300] For any third pixel in the third sub-pixel region, the reconstructed value of the fifth reference pixel corresponding to the third pixel in the fifth reference frame is determined as the seventh predicted value of the third pixel.
[0301] The reconstructed value of the sixth reference pixel corresponding to the third pixel in the sixth reference frame is determined as the eighth predicted value of the third pixel.
[0302] Optionally, the third determining module 1304 is further specifically used for:
[0303] Based on the fifth, sixth, seventh, and eighth predicted values, determine the second and third difference values for each target pixel.
[0304] When the second difference value and the third difference value meet the preset conditions, the target prediction value of each pixel in the second pixel region is determined based on the position information of each pixel in the second pixel region and the first motion information;
[0305] If the second difference value and the third difference value do not meet the preset conditions, the ninth and tenth predicted values of each pixel in the second pixel region are weighted and summed based on the preset second weight value combination to obtain the target predicted value of each pixel in the second pixel region.
[0306] Optionally, the third determining module 1304 is further specifically used for:
[0307] The difference between the fifth predicted value and the sixth predicted value is determined as the second difference value, and the difference between the seventh predicted value and the eighth predicted value is determined as the third difference value, or;
[0308] The difference between the fifth predicted value and the seventh predicted value is determined as the second difference value, and the difference between the sixth predicted value and the eighth predicted value is determined as the third difference value.
[0309] In this embodiment, first motion information of a first image block and second motion information of a second image block are obtained; based on the position information and first motion information of the first image block, at least one first pixel region associated with the first image block is determined; according to the first motion information and second motion information, a predicted value for each pixel in the first pixel region is determined; and based on the predicted value for each pixel, a target predicted value for each pixel in the second pixel region of the first image block is determined. In this embodiment, the predicted value for each pixel in the first pixel region is determined based on the first motion information and second motion information. This fully considers the motion differences between the first and second image blocks during the correction of the predicted values of boundary pixels, improving the accuracy of the corrected predicted values of boundary pixels, and thus improving video encoding and decoding efficiency.
[0310] The inter-frame prediction device provided in this application embodiment can achieve... Figure 7 The various processes implemented in the method embodiments achieve the same technical effect, and will not be described again here to avoid repetition.
[0311] The inter-frame prediction device in this application embodiment can be an electronic device, such as an electronic device with an operating system, or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or other devices besides a terminal. For example, the terminal can include, but is not limited to, the types of terminals listed above; other devices can be servers, network attached storage (NAS), etc., and this application embodiment does not specifically limit the scope of the device.
[0312] Optionally, such as Figure 14As shown, this application embodiment also provides a communication device 1400, including a processor 1401 and a memory 1402. The memory 1402 stores a program or instructions that can be run on the processor 1401. For example, when the communication device 1400 is a terminal, when the program or instructions are executed by the processor 1401, they implement the various steps of the above-described inter-frame prediction method embodiment and can achieve the same technical effect.
[0313] This application embodiment also provides a terminal, including a processor and a communication interface, wherein the processor is used to perform the following operations:
[0314] Obtain target information;
[0315] Based on the target information, inter-frame prediction is performed for each first image block.
[0316] Alternatively, it can be used to perform the following operations:
[0317] Acquire the first motion information of the first image block and the second motion information of the second image block;
[0318] Based on the position information of the first image block and the first motion information, at least one first pixel region associated with the first image block is determined;
[0319] Based on the first motion information and the second motion information, determine the predicted value of each pixel in the first pixel region;
[0320] Based on the predicted value of each pixel, the target predicted value of each pixel in the second pixel region of the first image block is determined.
[0321] This terminal embodiment corresponds to the aforementioned terminal-side method embodiment. All implementation processes and methods of the aforementioned method embodiments can be applied to this terminal embodiment and achieve the same technical effect. Specifically, Figure 15 A schematic diagram of the hardware structure of a terminal to implement an embodiment of this application.
[0322] The terminal 1500 includes, but is not limited to, the following components: radio frequency unit 1501, network module 1502, audio output unit 1503, input unit 1504, sensor 1505, display unit 1506, user input unit 1507, interface unit 1508, memory 1509, and processor 1510.
[0323] Those skilled in the art will understand that the terminal 1500 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1510 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 15The terminal structure shown does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0324] It should be understood that, in this embodiment, the input unit 1504 may include a graphics processing unit (GPU) 15041 and a microphone 15042. The GPU 15041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1506 may include a display panel 15061, which may be configured as a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1507 includes a touch panel 15071 and at least one of other input devices 15072. The touch panel 15071 is also called a touch screen. The touch panel 15071 may include a touch detection device and a touch controller. Other input devices 15072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0325] In this embodiment, after receiving downlink data from the network-side device, the radio frequency unit 1501 can transmit it to the processor 1510 for processing; the radio frequency unit 1501 can also send uplink data to the network-side device. Typically, the radio frequency unit 1501 includes, but is not limited to, antennas, amplifiers, transceivers, couplers, low-noise amplifiers, and duplexers.
[0326] The memory 1509 can be used to store software programs or instructions, as well as various data. The memory 1509 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1509 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1509 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0327] Processor 1510 may include one or more processing units; optionally, processor 1510 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1510.
[0328] The processor 1510 is used to perform the following operations:
[0329] Obtain target information;
[0330] Based on the target information, inter-frame prediction is performed for each first image block.
[0331] Alternatively, processor 1510 is used to perform the following operations:
[0332] Acquire the first motion information of the first image block and the second motion information of the second image block;
[0333] Based on the position information of the first image block and the first motion information, at least one first pixel region associated with the first image block is determined;
[0334] Based on the first motion information and the second motion information, determine the predicted value of each pixel in the first pixel region;
[0335] Based on the predicted value of each pixel, the target predicted value of each pixel in the second pixel region of the first image block is determined.
[0336] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described inter-frame prediction method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0337] The processor is the processor in the terminal described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0338] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described inter-frame prediction method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0339] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0340] This application also provides a computer program / program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described inter-frame prediction method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0341] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0342] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0343] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. An inter-frame prediction method, characterized in that, include: Obtain target information; The target information includes the prediction value export mode corresponding to the target image frame and / or the prediction value export mode corresponding to each first image block in the target image frame; Based on the target information, inter-frame prediction is performed on each of the first image blocks; Wherein, the target image frame is an image frame to be encoded, and the first image block is an image block to be encoded; or the target image frame is an image frame to be decoded, and the first image block is an image block to be decoded; Wherein, when the prediction value derivation mode corresponding to each first image block in the target image frame is a first derivation mode, the step of performing inter-frame prediction for each first image block based on the target information includes: Acquire first motion information of a first image block and second motion information of a second image block, wherein the first image block and the second image block are adjacent; wherein the first image block is an image block to be encoded and the second image block is an encoded image block; or, the first image block is an image block to be decoded and the second image block is a decoded image block; Based on the position information of the first image block and the first motion information, at least one first pixel region associated with the first image block is determined; Based on the first motion information and the second motion information, determine the predicted value of each pixel in the first pixel region; Based on the predicted value of each pixel, the target predicted value of each pixel in the second pixel region of the first image block is determined; The first export mode is a prediction value export mode determined based on the motion information corresponding to each first image block, the position information corresponding to each first image block, and the motion information corresponding to the adjacent blocks of each first image block. Wherein, the first pixel region includes a first sub-pixel region, which is a portion of the pixel region of the first image block, or the first sub-pixel region is a portion of the pixel region of the second image block, or the first sub-pixel region is a region composed of a portion of the pixel region of the first image block and a portion of the pixel region of the second image block. The step of determining the predicted value of each pixel in the first pixel region based on the first motion information and the second motion information includes: For any first pixel in the first sub-pixel region, the reconstructed value of the first reference pixel corresponding to the first pixel in the first reference frame is determined as the first predicted value of the first pixel. The reconstructed value of the second reference pixel corresponding to the first pixel in the second reference frame is determined as the second predicted value of the first pixel. Wherein, the first reference frame is determined based on the first motion information, and the second reference frame is determined based on the second motion information; The first difference value between the first predicted value and the second predicted value is used to determine the target predicted value for each pixel in the second pixel region.
2. The method according to claim 1, characterized in that, The target information includes the prediction value export mode corresponding to the target image frame; The step of performing inter-frame prediction for each first image patch in the target image frame based on the target information includes at least one of the following: When the prediction value export mode corresponding to the target image frame is the first export mode, the prediction value export mode corresponding to each first image block in the target image frame is determined to be the first export mode, and the first export mode is used to perform inter-frame prediction on each first image block. When the prediction value export mode corresponding to the target image frame is the second export mode, the inter-frame prediction value export method of each first image block in the target image frame is determined to be the second export mode, and the second export mode is used to perform inter-frame prediction on each first image block. The second export mode is a prediction value export mode determined based on the motion information corresponding to each first image block, the preset pixel region corresponding to each first image block, and the motion information corresponding to the adjacent blocks of each first image block. When the prediction value export mode corresponding to the target image frame is the third export mode, the inter-frame prediction value export method of each first image block in the target image frame is determined to be the third export mode, and the third export mode is used to perform inter-frame prediction on each first image block. The third export mode is a prediction value export mode determined based on the motion information corresponding to each first image block.
3. The method according to claim 1, characterized in that, The target information includes the prediction value derivation mode corresponding to each first image block in the target image frame; The step of performing inter-frame prediction for each first image patch in the target image frame based on the target information includes at least one of the following: When the prediction value export mode corresponding to any first image block is the first export mode, the first export mode is used to perform inter-frame prediction on the first image block. When the prediction value export mode corresponding to any first image block is the second export mode, the second export mode is used to perform inter-frame prediction on the first image block. The second export mode is a prediction value export mode determined based on the motion information corresponding to each first image block, the preset pixel region corresponding to each first image block, and the motion information corresponding to the adjacent blocks of each first image block. When the prediction value derivation mode corresponding to any first image block is the third derivation mode, the first image block is predicted inter-frame using the third derivation mode, wherein the third derivation mode is a prediction value derivation mode determined based on the motion information corresponding to each first image block.
4. The method according to claim 1, characterized in that, The target information includes the prediction value export mode corresponding to the target image frame and the prediction value export mode corresponding to each first image block in the target image frame; The step of performing inter-frame prediction for each first image patch in the target image frame based on the target information includes: When the prediction value export mode corresponding to the target image frame is the third export mode, the third export mode is used to perform inter-frame prediction for each first image block; If the prediction value export mode corresponding to the target image frame is not the third export mode, then inter-frame prediction is performed on each first image block based on the prediction value export mode corresponding to each first image block in the target image frame. The third export mode is a prediction value export mode determined based on the motion information corresponding to each first image block.
5. The method according to claim 1, characterized in that, The step of determining the target predicted value of each pixel in the second pixel region of the first image block based on the predicted value of each pixel includes: Calculate the first difference value between the first predicted value and the second predicted value for each first pixel in the first sub-pixel region; If the first difference value is greater than the first preset threshold, the target prediction value of each pixel in the second pixel region is determined based on the position information of each pixel in the second pixel region and the first motion information. When the first difference value is less than or equal to the first preset threshold, the third and fourth predicted values of each pixel in the second pixel region are weighted and summed based on the preset first weight value combination to obtain the target predicted value of each pixel in the second pixel region. The first weight value combination includes at least one weight recombination, the weight recombination includes a first weight value and a second weight value, the first weight value corresponds to the third predicted value, the second weight value corresponds to the fourth predicted value, the third predicted value is determined based on the first motion information, and the fourth predicted value is determined based on the second motion information.
6. The method according to claim 1, characterized in that, The first pixel region includes a second sub-pixel region and a third sub-pixel region, wherein the second sub-pixel region is a portion of the pixel region of the first image block, and the third sub-pixel region is a portion of the pixel region of the second image block; The step of determining the predicted value of each pixel in the first pixel region based on the first motion information and the second motion information includes: For any second pixel in the second sub-pixel region, the reconstructed value of the third reference pixel corresponding to the second pixel in the third reference frame is determined as the fifth predicted value of the second pixel. The reconstructed value of the fourth reference pixel corresponding to the second pixel in the fourth reference frame is determined as the sixth predicted value of the second pixel. For any third pixel in the third sub-pixel region, the reconstructed value of the fifth reference pixel corresponding to the third pixel in the fifth reference frame is determined as the seventh predicted value of the third pixel. The reconstructed value of the sixth reference pixel corresponding to the third pixel in the sixth reference frame is determined as the eighth predicted value of the third pixel. The third and fifth reference frames are determined based on the first motion information, and the fourth and sixth reference frames are determined based on the second motion information.
7. The method according to claim 6, characterized in that, The step of determining the target predicted value of each pixel in the second pixel region of the first image block based on the predicted value of each pixel includes: Based on the fifth, sixth, seventh, and eighth predicted values, a second difference value and a third difference value are determined for each target pixel; the target pixel includes a second pixel and a third pixel. When the second difference value and the third difference value meet the preset conditions, the target prediction value of each pixel in the second pixel region is determined based on the position information of each pixel in the second pixel region and the first motion information; If the second difference value and the third difference value do not meet the preset conditions, the ninth prediction value and the tenth prediction value of each pixel in the second pixel region are weighted and summed based on the preset second weight value combination to obtain the target prediction value of each pixel in the second pixel region. The second weight value combination includes at least one weight recombination, which includes a third weight value and a fourth weight value. The third weight value corresponds to the ninth predicted value, and the fourth weight value corresponds to the tenth predicted value. The ninth predicted value is determined based on the first motion information, and the tenth predicted value is determined based on the second motion information.
8. The method according to claim 7, characterized in that, The step of determining the second difference value and the third difference value based on the fifth predicted value, the sixth predicted value, the seventh predicted value, and the eighth predicted value includes: The difference between the fifth predicted value and the sixth predicted value is determined as the second difference value, and the difference between the seventh predicted value and the eighth predicted value is determined as the third difference value, or; The difference between the fifth predicted value and the seventh predicted value is determined as the second difference value, and the difference between the sixth predicted value and the eighth predicted value is determined as the third difference value.
9. The method according to claim 7, characterized in that, The preset conditions include any one of the following: The second difference value is smaller than the third difference value; The ratio between the second difference value and the third difference value is less than the second preset threshold; The ratio between the third difference value and the second difference value is less than the second preset threshold; The second difference value is greater than the third preset threshold, and the ratio between the second difference value and the third difference value is less than the second preset threshold; The second difference value is greater than the third preset threshold, and the ratio between the third difference value and the second difference value is less than the second preset threshold.
10. The method according to any one of claims 1, 5-9, characterized in that, The first pixel region satisfies at least one of the following: The first pixel region is the encoded or decoded pixel region consisting of M1 rows and N1 columns of pixels adjacent to the top edge of the first image block; The first pixel region is the encoded or decoded pixel region consisting of M2 rows and N2 columns of pixels adjacent to the left of the first image block; The first pixel region is an uncoded or undecoded pixel region consisting of M3 rows and N3 columns of pixels located inside and above the first image block; The first pixel region is an uncoded or undecoded pixel region consisting of M4 rows and N4 columns of pixels located inside the left side of the first image block; The first pixel region is an M5-row, N5-column pixel region consisting of the encoded or decoded pixel region adjacent to the top of the first image block and the unencoded or undecoded pixel region inside the top of the first image block; The first pixel region is a pixel region consisting of an M6-row, N6-column region formed by the coded or decoded pixel region adjacent to the left of the first image block and the uncoded or undecoded pixel region inside the left of the first image block. Where M1, M2, M3, M4, M5, M6, N1, N2, N3, N4, N5, and N6 are all positive integers.
11. An inter-frame prediction method, characterized in that, include: Acquire first motion information of a first image block and second motion information of a second image block, wherein the first image block and the second image block are adjacent; wherein the first image block is an image block to be encoded and the second image block is an encoded image block; or, the first image block is an image block to be decoded and the second image block is a decoded image block; Based on the position information of the first image block and the first motion information, at least one first pixel region associated with the first image block is determined; Based on the first motion information and the second motion information, determine the predicted value of each pixel in the first pixel region; Based on the predicted value of each pixel, the target predicted value of each pixel in the second pixel region of the first image block is determined; Wherein, the first pixel region includes a first sub-pixel region, which is a portion of the pixel region of the first image block, or the first sub-pixel region is a portion of the pixel region of the second image block, or the first sub-pixel region is a region composed of a portion of the pixel region of the first image block and a portion of the pixel region of the second image block. The step of determining the predicted value of each pixel in the first pixel region based on the first motion information and the second motion information includes: For any first pixel in the first sub-pixel region, the reconstructed value of the first reference pixel corresponding to the first pixel in the first reference frame is determined as the first predicted value of the first pixel. The reconstructed value of the second reference pixel corresponding to the first pixel in the second reference frame is determined as the second predicted value of the first pixel. Wherein, the first reference frame is determined based on the first motion information, and the second reference frame is determined based on the second motion information; The first difference value between the first predicted value and the second predicted value is used to determine the target predicted value for each pixel in the second pixel region.
12. The method according to claim 11, characterized in that, The step of determining the target predicted value of each pixel in the second pixel region of the first image block based on the predicted value of each pixel includes: Calculate the first difference value between the first predicted value and the second predicted value for each first pixel in the first sub-pixel region; If the first difference value is greater than the first preset threshold, the target prediction value of each pixel in the second pixel region is determined based on the position information of each pixel in the second pixel region and the first motion information. When the first difference value is less than or equal to the first preset threshold, the third and fourth predicted values of each pixel in the second pixel region are weighted and summed based on the preset first weight value combination to obtain the target predicted value of each pixel in the second pixel region. The first weight value combination includes at least one weight recombination, the weight recombination includes a first weight value and a second weight value, the first weight value corresponds to the third predicted value, the second weight value corresponds to the fourth predicted value, the third predicted value is determined based on the first motion information, and the fourth predicted value is determined based on the second motion information.
13. The method according to claim 11, characterized in that, The first pixel region includes a second sub-pixel region and a third sub-pixel region, wherein the second sub-pixel region is a portion of the pixel region of the first image block, and the third sub-pixel region is a portion of the pixel region of the second image block; The step of determining the predicted value of each pixel in the first pixel region based on the first motion information and the second motion information includes: For any second pixel in the second sub-pixel region, the reconstructed value of the third reference pixel corresponding to the second pixel in the third reference frame is determined as the fifth predicted value of the second pixel. The reconstructed value of the fourth reference pixel corresponding to the second pixel in the fourth reference frame is determined as the sixth predicted value of the second pixel. For any third pixel in the third sub-pixel region, the reconstructed value of the fifth reference pixel corresponding to the third pixel in the fifth reference frame is determined as the seventh predicted value of the third pixel. The reconstructed value of the sixth reference pixel corresponding to the third pixel in the sixth reference frame is determined as the eighth predicted value of the third pixel. The third and fifth reference frames are determined based on the first motion information, and the fourth and sixth reference frames are determined based on the second motion information.
14. The method according to claim 13, characterized in that, The step of determining the target predicted value of each pixel in the second pixel region of the first image block based on the predicted value of each pixel includes: Based on the fifth, sixth, seventh, and eighth predicted values, a second difference value and a third difference value are determined for each target pixel; the target pixel includes a second pixel and a third pixel. When the second difference value and the third difference value meet the preset conditions, the target prediction value of each pixel in the second pixel region is determined based on the position information of each pixel in the second pixel region and the first motion information; If the second difference value and the third difference value do not meet the preset conditions, the ninth prediction value and the tenth prediction value of each pixel in the second pixel region are weighted and summed based on the preset second weight value combination to obtain the target prediction value of each pixel in the second pixel region. The second weight value combination includes at least one weight recombination, which includes a third weight value and a fourth weight value. The third weight value corresponds to the ninth predicted value, and the fourth weight value corresponds to the tenth predicted value. The ninth predicted value is determined based on the first motion information, and the tenth predicted value is determined based on the second motion information.
15. The method according to claim 14, characterized in that, The step of determining the second difference value and the third difference value based on the fifth predicted value, the sixth predicted value, the seventh predicted value, and the eighth predicted value includes: The difference between the fifth predicted value and the sixth predicted value is determined as the second difference value, and the difference between the seventh predicted value and the eighth predicted value is determined as the third difference value, or; The difference between the fifth predicted value and the seventh predicted value is determined as the second difference value, and the difference between the sixth predicted value and the eighth predicted value is determined as the third difference value.
16. The method according to claim 14, characterized in that, The preset conditions include any one of the following: The second difference value is smaller than the third difference value; The ratio between the second difference value and the third difference value is less than the second preset threshold; The ratio between the third difference value and the second difference value is less than the second preset threshold; The second difference value is greater than the third preset threshold, and the ratio between the second difference value and the third difference value is less than the second preset threshold; The second difference value is greater than the third preset threshold, and the ratio between the third difference value and the second difference value is less than the second preset threshold.
17. The method according to any one of claims 11-16, characterized in that, The first pixel region satisfies at least one of the following: The first pixel region is the encoded or decoded pixel region consisting of M1 rows and N1 columns of pixels adjacent to the top edge of the first image block; The first pixel region is the encoded or decoded pixel region consisting of M2 rows and N2 columns of pixels adjacent to the left of the first image block; The first pixel region is an uncoded or undecoded pixel region consisting of M3 rows and N3 columns of pixels located inside and above the first image block; The first pixel region is an uncoded or undecoded pixel region consisting of M4 rows and N4 columns of pixels located inside the left side of the first image block; The first pixel region is an M5-row, N5-column pixel region consisting of the encoded or decoded pixel region adjacent to the top of the first image block and the unencoded or undecoded pixel region inside the top of the first image block; The first pixel region is an M6-row, N6-column pixel region consisting of the encoded or decoded pixel region adjacent to the left of the first image block and the unencoded or undecoded pixel region inside the left of the first image block; Where M1, M2, M3, M4, M5, M6, N1, N2, N3, N4, N5, and N6 are all positive integers.
18. An inter-frame prediction device, characterized in that, include: The acquisition module is used to acquire target information; The target information includes the prediction value export mode corresponding to the target image frame and / or the prediction value export mode corresponding to each first image block in the target image frame; The processing module is used to perform inter-frame prediction on each of the first image blocks based on the target information; Wherein, the target image frame is an image frame to be encoded, and the first image block is an image block to be encoded; or the target image frame is an image frame to be decoded, and the first image block is an image block to be decoded; Specifically, the processing module is used for: Acquire the first motion information of the first image block and the second motion information of the second image block; Based on the position information of the first image block and the first motion information, at least one first pixel region associated with the first image block is determined; wherein, the first image block is an image block to be encoded and the second image block is an encoded image block; or, the first image block is an image block to be decoded and the second image block is a decoded image block; Based on the first motion information and the second motion information, determine the predicted value of each pixel in the first pixel region; Based on the predicted value of each pixel, the target predicted value of each pixel in the second pixel region of the first image block is determined; Wherein, the first pixel region includes a first sub-pixel region, which is a portion of the pixel region of the first image block, or the first sub-pixel region is a portion of the pixel region of the second image block, or the first sub-pixel region is a region composed of a portion of the pixel region of the first image block and a portion of the pixel region of the second image block. The processing module is also specifically used for: For any first pixel in the first sub-pixel region, the reconstructed value of the first reference pixel corresponding to the first pixel in the first reference frame is determined as the first predicted value of the first pixel. The reconstructed value of the second reference pixel corresponding to the first pixel in the second reference frame is determined as the second predicted value of the first pixel. Wherein, the first reference frame is determined based on the first motion information, and the second reference frame is determined based on the second motion information; The first difference value between the first predicted value and the second predicted value is used to determine the target predicted value for each pixel in the second pixel region.
19. An inter-frame prediction device, characterized in that, include: The acquisition module is used to acquire first motion information of a first image block and second motion information of a second image block, wherein the first image block and the second image block are adjacent; wherein the first image block is an image block to be encoded and the second image block is an encoded image block; or, the first image block is an image block to be decoded and the second image block is a decoded image block; The first determining module is configured to determine at least one first pixel region associated with the first image block based on the position information of the first image block and the first motion information. The second determining module is used to determine the predicted value of each pixel in the first pixel region based on the first motion information and the second motion information. The third determining module is used to determine the target predicted value of each pixel in the second pixel region of the first image block based on the predicted value of each pixel. Wherein, the first pixel region includes a first sub-pixel region, which is a portion of the pixel region of the first image block, or the first sub-pixel region is a portion of the pixel region of the second image block, or the first sub-pixel region is a region composed of a portion of the pixel region of the first image block and a portion of the pixel region of the second image block. The second determining module is specifically used for: For any first pixel in the first sub-pixel region, the reconstructed value of the first reference pixel corresponding to the first pixel in the first reference frame is determined as the first predicted value of the first pixel. The reconstructed value of the second reference pixel corresponding to the first pixel in the second reference frame is determined as the second predicted value of the first pixel. Wherein, the first reference frame is determined based on the first motion information, and the second reference frame is determined based on the second motion information; The first difference value between the first predicted value and the second predicted value is used to determine the target predicted value for each pixel in the second pixel region.
20. A terminal, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the inter-frame prediction method as described in any one of claims 1-10, or to implement the steps of the inter-frame prediction method as described in any one of claims 11-17.
21. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the inter-frame prediction method as described in any one of claims 1-10, or implement the steps of the inter-frame prediction method as described in any one of claims 11-17.
Citation Information
Patent Citations
Inter frame prediction method and device for video images and codec
CN109587479A
Overlapped block motion compensation using spatial neighbors
US20210176472A1
Signaling for motion vector refinement
WO2021061023A1