Video coding method and apparatus, video decoding method and apparatus, and device, storage medium and program product
By optimizing motion vectors by selecting target positions within reference frames, the problem of low coding efficiency caused by poor motion vector selection in existing technologies is solved, achieving more efficient video encoding and decoding.
Patent Information
- Application Number
- PCT/CN2025/111658
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-09
- Filing Date
- 2025-07-31
- Publication Date
- 2026-02-12
AI Technical Summary
In video encoding and decoding, existing technologies may not select the optimal motion vector during motion estimation, resulting in low encoding efficiency and affecting the video encoding and decoding effect.
By selecting the target location within the reference frame, the coding cost of the motion vector is ensured to be higher than that of the original motion vector while the pixel distortion is less than a set threshold, thereby optimizing the motion vector of the current block to improve coding efficiency.
It improves the efficiency of video encoding and decoding, reduces pixel distortion through precise motion vector optimization, and enhances the effect of video encoding and decoding.
Smart Images

Figure CN2025111658_12022026_PF_FP_ABST
Abstract
Description
Video coding method, device, apparatus, storage medium and program product
[0001] This application claims priority to the Chinese patent application No. 202411096730.6, filed on August 9, 2024, and entitled "Video coding method, device, computer readable medium and electronic device". TECHNICAL FIELD
[0002] The present application relates to the field of computers, in particular to a video coding method, device, apparatus, storage medium and program product.
[0003] BACKGROUND
[0004] In the field of video coding, inter prediction mode is to use the correlation in the time domain of video to predict the pixels of the current image using the pixels of the adjacent coded image (i.e. reference image), so as to effectively remove the redundancy in the time domain of video. The process of finding the best reference block in the reference image is called motion estimation.
[0005] In the process of motion estimation, a "best" position is selected from multiple possible candidate positions according to the cost function. However, the position of the prediction block selected when the cost function is minimum is not necessarily the best in terms of block compensation (prediction) effect, i.e. the determined motion vector (MV) is not accurate, which will affect the video coding efficiency. SUMMARY
[0006] Embodiments of the present application provide a video coding method, device, apparatus, storage medium and program product, which can optimize the original motion vector of the current block, thereby improving the video coding efficiency.
[0007] In one aspect, the embodiments of the present application provide a video decoding method, comprising:
[0008] decoding a video bitstream to obtain a first motion vector of a current block;
[0009] selecting a target position from multiple candidate positions in a reference frame based on the first motion vector, wherein a coding cost calculated based on a motion vector indicated by the target position is higher than a coding cost calculated based on the first motion vector, and a pixel distortion between the current block and a reference block where the target position is located is less than a set threshold; and
[0010] determining a second motion vector of the current block based on the target position, and performing motion compensation on the current block based on the second motion vector.
[0011] In another aspect, an embodiment of the present application provides a video encoding method, comprising:
[0012] determining a first motion vector of a current block;
[0013] selecting a target position from a plurality of candidate positions in a reference frame based on the first motion vector, wherein a coding cost calculated based on a motion vector indicated by the target position is higher than a coding cost calculated based on the first motion vector, and a pixel distortion between the current block and a reference block in which the target position is located is less than a set threshold; and
[0014] determining a second motion vector of the current block based on the target position, and performing motion compensation on the current block based on the second motion vector.
[0015] In another aspect, an embodiment of the present application provides a video decoding apparatus, comprising:
[0016] a decoding unit configured to decode a video code stream to obtain a first motion vector of a current block;
[0017] a selecting unit configured to select a target position from a plurality of candidate positions in a reference frame based on the first motion vector, wherein a coding cost calculated based on a motion vector indicated by the target position is higher than a coding cost calculated based on the first motion vector, and a pixel distortion between the current block and a reference block in which the target position is located is less than a set threshold; and
[0018] a processing unit configured to determine a second motion vector of the current block based on the target position, and perform motion compensation on the current block based on the second motion vector.
[0019] In another aspect, an embodiment of the present application provides a video encoding apparatus, comprising:
[0020] a determining unit configured to determine a first motion vector of a current block;
[0021] a selecting unit configured to select a target position from a plurality of candidate positions in a reference frame based on the first motion vector, wherein a coding cost calculated based on a motion vector indicated by the target position is higher than a coding cost calculated based on the first motion vector, and a pixel distortion between the current block and a reference block in which the target position is located is less than a set threshold; and
[0022] a processing unit configured to determine a second motion vector of the current block based on the target position, and perform motion compensation on the current block based on the second motion vector.
[0023] In another aspect, an embodiment of the present application provides a computer readable medium having stored thereon a computer program which, when executed by a processor, implements the video decoding method or the video encoding method as described in the above embodiments.
[0024] In another aspect, an embodiment of the present application provides an electronic device, comprising: at least one processor; a storage device configured to store at least one computer program which, when executed by the at least one processor, causes the electronic device to implement the video decoding method or the video encoding method as described in the above embodiments.
[0025] In another aspect, an embodiment of the present application provides a computer program product comprising a computer program stored in a computer readable storage medium. A processor of an electronic device reads and executes the computer program from the computer readable storage medium, so that the electronic device performs the video decoding method or the video encoding method provided in the various optional embodiments described above.
[0026] BRIEF DESCRIPTION OF DRAWINGS
[0027] FIG. 1 shows a schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of the present application can be applied;
[0028] FIG. 2 shows a schematic diagram of the placement of a video encoding device and a video decoding device in a streaming system;
[0029] FIG. 3 shows a basic flowchart of a video encoder;
[0030] FIG. 4 shows a schematic diagram of an inter prediction process;
[0031] FIG. 5 shows a schematic diagram of an inter prediction process;
[0032] FIG. 6 shows a flowchart of a video decoding method according to an embodiment of the present application;
[0033] FIG. 7 shows a flowchart of a video encoding method according to an embodiment of the present application;
[0034] FIG. 8 shows a schematic diagram of a search region according to an embodiment of the present application;
[0035] FIG. 9 shows a block diagram of a video decoding device according to an embodiment of the present application;
[0036] FIG. 10 shows a block diagram of a video encoding device according to an embodiment of the present application;
[0037] FIG. 11 shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present application.
[0038] EMBODIMENTS
[0039] Example implementations are now described with reference to the drawings. Example implementations can, however, be implemented in various forms and should not be construed as limited to the examples presented in this disclosure; rather, these implementations are provided as example embodiments so as to convey the gist of the implementations to those skilled in the art.
[0040] In addition, the features, structures or characteristics described in this disclosure can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are recited in order to provide a thorough understanding of the embodiments of the present disclosure. One skilled in the art, however, will recognize that the embodiments can be practiced without the specific details, that one or more features can be combined into a single feature, a plurality of features two or more features can be divided into multiple features, or the scope of the embodiments is not limited to the features described in this disclosure.
[0041] In the embodiments of the present disclosure, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement at least one module or unit. In addition, each module or unit can be a part of an integral module or unit that includes the functions of the module or unit.
[0042] The block diagrams shown in the drawings are only functional entities, and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in the form of software, or in at least one hardware module or integrated circuit, or in different networks and / or processor devices and / or microcontroller devices.
[0043] The flowcharts shown in the drawings are only exemplary illustrations, and do not necessarily include all contents and operations / steps, nor do they necessarily have to be executed in the order described. For example, some operations / steps can be further divided, and some operations / steps can be combined or partially combined, so the actual execution order can be changed according to the actual situation.
[0044] It should be noted that "multiple" referred to herein means two or more. The association relationship of "and / or" described in association with the associated objects means that there can be three relationships, for example, A and / or B can represent the three cases of A alone, A and B together, and B alone. The character " / " generally represents an "or" relationship between the associated objects before and after it.
[0045] FIG. 1 shows a schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of the present disclosure can be applied.
[0046] As shown in FIG. 1, system architecture 100 includes multiple terminal devices that can communicate with each other via, for example, network 150. For example, system architecture 100 can include a first terminal device 110 and a second terminal device 120 interconnected via network 150. In the example of FIG. 1, first terminal device 110 and second terminal device 120 perform unidirectional transmission of data.
[0047] For example, first terminal device 110 can encode video data (e.g., a stream of video pictures that are captured by terminal device 110) for communication to second terminal device 120 via network 150. The encoded video data can be transmitted in the form of at least one coded video bitstream. Second terminal device 120 can receive the coded video data from network 150, decode the coded video data to recover the video pictures, and display video pictures according to the recovered video data.
[0048] In an embodiment of the present application, system architecture 100 can include a third terminal device 130 and a fourth terminal device 140 that perform bidirectional transmission of encoded video data, which can occur, for example, during a videoconferencing session. For bidirectional transmission of data, each terminal device of the third and fourth terminal devices 130, 140 can code video data (e.g., a stream of video pictures that are captured by the terminal device) for communication to the other terminal device of the third and fourth terminal devices 130, 140 via network 150. Each terminal device of the third and fourth terminal devices 130, 140 also can receive the coded video data transmitted by the other terminal device of the third and fourth terminal devices 130, 140, and can decode the coded video data to recover the video pictures, and can display video pictures on an accessible display according to the recovered video data.
[0049] In the example of FIG. 1, first terminal device 110, second terminal device 120, third terminal device 130, and fourth terminal device 140 can be servers or terminals, although the principles disclosed herein can not be limited to such.
[0050] The server can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, content delivery network (CDN), and basic cloud computing services such as big data and artificial intelligence platform. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart voice interaction device, a smart watch, a smart home appliance, a vehicle-mounted terminal, an aircraft, and the like, but is not limited thereto.
[0051] The network 150 shown in FIG. 1 represents any number of networks that convey coded video data among the first terminal device 110, the second terminal device 120, the third terminal device 130, and the fourth terminal device 140, including, for example, wired and / or wireless communication networks. The communication network 150 can exchange data in circuit- switched and / or packet-switched channels. The network can include telecommunication networks, local area and / or wide area networks, and / or the Internet. For the purposes of the present application, the architecture and topology of the network 150 can be immaterial to the operation of the disclosed subject matter unless explained in the following narratives.
[0052] FIG. 2 shows the placement of video encoding and video decoding devices in a streaming environment, in an embodiment of the present application. The disclosed subject matter can be equally applicable to other video enabled applications including, for example, video conferencing, digital television (TV), storing of compressed video on digital media including CD, DVD, memory stick and the like, and so forth.
[0053] A streaming system can include a capture subsystem 213, which can include a video source 201, such as a digital camera, creating a stream of uncompressed video pictures 202. In an embodiment, the stream of video pictures 202 includes samples as taken by the digital camera. In contrast to encoded video data 204 (or coded video bitstreams 204), the stream of video pictures 202 is depicted as a thick line to emphasize the high data volume of the stream of video pictures 202, which can be processed by an electronic device 220 including a video encoding device 203 coupled to the video source 201. The video encoding device 203 can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter as described in more detail below. In contrast to the stream of video pictures 202, the encoded video data 204 (or coded video bitstream 204) is depicted as a thin line to emphasize the lower data volume of the encoded video data 204 (or coded video bitstream 204), which can be stored on a streaming server 205 for future use. At least one streaming client subsystem, such as client subsystems 206 and 208 in FIG. 2, can access the streaming server 205 to retrieve copies 207 and 209 of the encoded video data 204. A client subsystem 206 can include a video decoding device 210, such as in an electronic device 230. The video decoding device 210 decodes the incoming copy 207 of encoded video data and creates an outgoing stream of video pictures 211, which can be rendered on a display 212, such as a display screen, or other presentation devices. In some streaming systems, the encoded video data 204, 207, and 209 (e.g., video bitstreams) can be encoded according to certain video coding / compression standards.
[0054] It is noted that the electronic device 220 and the electronic device 230 can include other components not shown in the figures. For example, the electronic device 220 can include a video decoding device, and the electronic device 230 can also include a video encoding device.
[0055] In an embodiment of the present application, taking the High Efficiency Video Coding (HEVC) in the international video coding standard, the Versatile Video Coding (VVC), and the Chinese national video coding standard AVS as examples, when an input video frame image is input, the video frame image is divided into a plurality of non-overlapping processing units according to the size of a block, and each processing unit will perform similar compression operations. This processing unit is called a Coding Tree Unit (CTU), or also called a Largest Coding Unit (LCU). The CTU can be further divided into a more detailed division to obtain at least one basic Coding Unit (CU), and the CU is the most basic element in the coding link.
[0056] In another embodiment, this processing unit can also be called a coding tile (i.e., a tile), which is a rectangular area of a multimedia data frame that can be independently decoded and encoded. In the Alliance for Open Media Video 1 (AV1) standard formulated by the Open Media Alliance, the coding tile can be further divided into a more detailed division to obtain at least one Superblock (SB for short), which is the starting point of block division and can be further divided into a plurality of sub-blocks, and then further divided into at least one block. Each block is the most basic element in the coding link. Alternatively, a SB can contain a plurality of blocks.
[0057] The above division method of the video frame image can be called a block partition structure, and some concepts in the coding process are introduced as follows:
[0058] Predictive Coding: Predictive coding includes Intra prediction and Inter prediction, etc. The original video signal is predicted by the selected reconstructed video signal, and the residual video signal is obtained. The encoder needs to determine which prediction coding mode is selected for the current coding unit (or coding block) and inform the decoder. Intra prediction refers to the predicted signal coming from the already coded and reconstructed area within the same image; Inter prediction refers to the predicted signal coming from the already coded image (referred to as reference image) other than the current image.
[0059] Transform & Quantization: The residual video signal is transformed into the transform domain by Discrete Fourier Transform (DFT), Discrete Cosine Transform (DCT), etc. The signal is converted into transform coefficients. The transform coefficients are further subjected to lossy quantization operation, losing some information, so that the quantized signal is conducive to compressed expression. In some video coding standards, more than one transform method can be selected, so the encoder also needs to select one of the transform methods for the current coding unit (or coding block) and inform the decoder. The quantization precision is usually determined by the quantization parameter (QP). A larger QP value indicates that coefficients with a larger value range will be quantized to the same output, so it usually brings larger distortion and lower code rate. Conversely, a smaller QP value indicates that coefficients with a smaller value range will be quantized to the same output, so it usually brings smaller distortion and higher code rate.
[0060] Entropy Coding or Statistical Coding: The quantized transform domain signal will be statistically compressed and coded according to the frequency of each value, and finally output the binary (0 or 1) compressed code stream. At the same time, other information generated by encoding, such as the selected coding mode, motion vector data, etc., also needs to be entropy coded to reduce the code rate. Statistical coding is a lossless coding method that can effectively reduce the code rate required to express the same signal. Common statistical coding methods include variable length coding (VLC) or context-based binary arithmetic coding (CABAC).
[0061] The CABAC process mainly includes three steps: binarization, context modeling and binary arithmetic coding. After the input syntax elements are binarized, the binary data can be encoded by regular coding mode and bypass coding mode. The bypass coding mode does not need to assign a specific probability model for each binary bit, and the input binary bit bin value is directly encoded by a simple bypass encoder to speed up the entire encoding and decoding speed. Generally, different syntax elements are not completely independent, and the same syntax element also has certain memory. Therefore, according to the conditional entropy theory, the use of other coded syntax elements for conditional coding can further improve the coding performance compared with independent coding or non-memory coding. These coded symbol information used as conditions are called contexts. In the regular coding mode, the binary bits of the syntax elements enter the context modeler in order, and the encoder assigns an appropriate probability model for each input binary bit according to the value of the previously coded syntax elements or binary bits. This process is called context modeling. The context index increment (ctxIdxInc) and the context index Start (ctxIdxStart) can be used to locate the context model corresponding to the syntax element. After the bin value and the assigned probability model are sent to the binary arithmetic encoder for coding, the context model needs to be updated according to the bin value, that is, the adaptive process in coding.
[0062] Loop filtering: The signal after transformation and quantization will obtain the reconstructed image through the operation of inverse quantization, inverse transformation and prediction compensation. Due to the influence of quantization, the reconstructed image is different from the original image in some information, that is, the reconstructed image will produce distortion. Therefore, the reconstructed image can be filtered, such as deblocking filter (DB), sample adaptive offset (SAO) or adaptive loop filter (ALF), which can effectively reduce the distortion degree caused by quantization. Since these filtered reconstructed images will be used as references for subsequent encoded images to predict future image signals, the above filtering operation is also called loop filtering, that is, the filtering operation within the coding loop.
[0063] In an embodiment of the present application, FIG. 3 shows a basic flowchart of a video encoder, in which an intra prediction is taken as an example. In the flowchart, an original image signal is subtracted from a predicted image signal to obtain a residual signal, the residual signal is processed by transformation and quantization to obtain quantized coefficients, the quantized coefficients are encoded by entropy coding to obtain a bitstream, and the quantized coefficients are processed by inverse quantization and inverse transformation to obtain a reconstructed residual signal. The predicted image signal is superimposed with the reconstructed residual signal to generate an image signal. The image signal is input to an intra mode decision module and an intra prediction module for intra prediction processing, and is output as a reconstructed image signal by loop filtering. The reconstructed image signal can be used as a reference image for motion estimation and motion compensation prediction of a next frame. Then, a predicted image signal of the next frame is obtained based on the results of motion compensation prediction and intra prediction, and the above process is repeated until the encoding is completed.
[0064] Based on the above encoding process, at the decoding end, for each coding unit (or coding block), after obtaining the compressed code stream (i.e., the bitstream), entropy decoding is performed to obtain various mode information and quantized coefficients. Then, the quantized coefficients are processed by inverse quantization and inverse transformation to obtain a residual signal. On the other hand, according to the known encoding mode information, a prediction signal corresponding to the coding unit (or coding block) can be obtained, and then the residual signal is added to the prediction signal to obtain a reconstructed signal. The reconstructed signal is further processed by loop filtering and the like to generate a final output signal.
[0065] In the field of encoding technology, inter prediction uses the correlation in the time domain of a video to predict pixels of a current image using pixels of a neighboring coded image, so as to effectively remove the temporal redundancy of the video and effectively save bits of the coded residual data.
[0066] As shown in FIG. 4, P represents a current frame, Pr represents a reference frame (or a reference image), B represents a current coding block, and Br represents a reference block of B. The coordinates of B' in the reference frame are the same as the coordinate position of B in the current frame, the coordinates of Br are (xr, yr), and the coordinates of B' are (x, y). The displacement between the current coding block and the reference block thereof is referred to as a motion vector MV, where MV = (xr-x, yr-y). In other words, inter prediction refers to a process of searching for a reference block from a neighboring coded image (i.e., a reference frame) according to a current block to be coded in a current frame, so as to remove the temporal redundancy of a video signal.
[0067] As shown in FIG. 5, for a current block to be encoded in a current frame, a search is performed in a certain range (i.e. a search area formed by a search frame) in a reference frame according to a block matching criterion to obtain a best matching reference block, referred to as a best matching block. Optionally, the block matching criterion commonly used in video encoding includes a minimum mean square error (MSE), a sum of absolute difference (SAD), and the like.
[0068] Considering that adjacent blocks in the time domain or the spatial domain have strong correlation, a MV prediction technique can be used to further reduce the bits required for encoding the MV. In some audio / video standards, such as H.265 / HEVC, inter prediction includes two MV prediction techniques, namely Merge and advanced motion vector prediction (AMVP).
[0069] The Merge mode establishes a MV candidate list for the current block, such as including 5 candidate MVs (and corresponding reference images). The 5 candidate MVs are traversed to select a candidate MV with the minimum rate-distortion cost as the optimal MV. If the codec establishes the MV candidate list in the same way, the encoder only needs to transmit the index of the optimal MV in the MV candidate list.
[0070] It should be noted that the MV prediction technique of HEVC also has a skip mode, which is a special case of the Merge mode. After the optimal MV is found by the Merge mode, if the current block and the reference block are basically the same, the residual data does not need to be transmitted, and only the index of the MV and a skip flag need to be transmitted.
[0071] Similarly, the AMVP mode uses the MV correlation of the adjacent blocks in the spatial domain and the time domain to establish a candidate motion vector predictor (MVP) list for the current block. Unlike the Merge mode, the candidate MVP list of the AMVP mode includes multiple MVPs, i.e. predicted values of the motion vector, referred to as motion vector prediction. The optimal MVP is selected from the candidate MVP list, and the optimal MV obtained by motion search for the current block is differentially encoded, i.e. motion vector difference (MVD) = MV-MVP. The decoding end only needs the sequence number of the MVP in the list and the MVD to calculate the MV of the current block by establishing the same list. The AMVP candidate MV list also includes the spatial and temporal cases.
[0072] In the motion estimation process of the inter prediction technique, a "best" position block needs to be selected from multiple possible candidate position blocks as a reference block of the current block for inter prediction according to a cost function. In an example, the cost function includes pixel distortion (Distortion, D) of block matching and coding efficiency rate (Rate, R) of a motion vector, which can be expressed as J = D + λ x R, where J represents the cost function, and λ represents a parameter for balancing the pixel distortion D and the coding efficiency R at different rates.
[0073] Generally, when the reference block for motion compensation is finally selected, it is not the block position with the minimum D or R cost, but the block position with the minimum cost function J, i.e., the block position with the minimum weighted sum. However, the position block or the reference block selected when the cost function is minimum may not be the best in terms of block compensation (prediction) effect, which leads to inaccurate determined motion vector and affects the video coding efficiency.
[0074] Therefore, the technical scheme of the embodiment of the present application can optimize the original motion vector of the current block to reduce the pixel distortion, and thus can improve the accuracy of the optimized motion vector, which is beneficial to improving the video coding efficiency.
[0075] FIG. 6 shows a flowchart of a video decoding method according to an embodiment of the present application, which can be executed by an electronic device with a computing processing function, such as a terminal device or a server.
[0076] Referring to FIG. 6, the video decoding method includes at least S610 to S630, which are described in detail as follows.
[0077] In S610, a video bitstream is decoded to obtain a first motion vector of a current block.
[0078] In this step, the first motion vector is the original motion vector of the current block.
[0079] In some optional embodiments, the video includes a sequence of video image frames, each of which can be further divided into slices, and each slice can be further divided into a series of LCUs (or CTUs), and each LCU includes a plurality of CUs. The video image frames are encoded in units of blocks. In some new video coding standards, such as in the H.264 standard, there are macroblocks (MBs), which can be further divided into a plurality of prediction blocks that can be used for predictive coding. In the HEVC standard, a plurality of block units are divided in terms of functions by using basic concepts such as coding units (CUs), prediction units (PUs), and transform units (TUs), and a brand-new tree-based structure is used for description. For example, a CU can be divided into smaller CUs according to a quadtree, and the smaller CUs can be further divided, thereby forming a quadtree structure. The current block, the reference block, and the neighboring block in the embodiments of the present application can be CUs, or smaller blocks obtained by dividing the CUs, such as smaller blocks obtained by dividing the CUs.
[0080] Optionally, the current block uses an inter-prediction mode, such as a Merge mode or an AMVP prediction mode.
[0081] In some optional embodiments, if the current block uses the Merge mode, the video bitstream can be decoded to obtain a motion vector index of the current block, and then, based on the motion vector index of the current block, a first motion vector of the current block, i.e., an original motion vector, is determined in a candidate motion vector list (i.e., a Merge list) corresponding to the current block.
[0082] In some optional embodiments, if the current block uses the AMVP mode, the video bitstream can be decoded to obtain an MVP index and a motion vector residual (i.e., an MVD) of the current block, and then, based on the MVP index, an MVP of the current block is determined in a candidate MVP list corresponding to the current block, and further, based on the MVP of the current block and the motion vector residual, an original motion vector of the current block is determined. Optionally, the sum of the MVP and the motion vector residual is the original motion vector of the current block.
[0083] In S620, based on the first motion vector, a target position is selected from a plurality of candidate positions in the reference frame, where a coding cost calculated based on a motion vector indicated by the target position is higher than a coding cost calculated based on the first motion vector, and a pixel distortion between the current block and a reference block in which the target position is located is less than a set threshold.
[0084] In some optional embodiments, the plurality of candidate positions can be positions related to the original motion vector of the current block, such as a plurality of positions determined around the position indicated by the original motion vector of the current block.
[0085] In some optional embodiments, if the current block adopts the Merge mode, the plurality of specified positions can include positions indicated by each candidate motion vector in the candidate motion vector list (i.e., the Merge list). In this case, when selecting the target position, a target motion vector with a coding efficiency higher than that of the original motion vector and a pixel distortion less than a set threshold can be selected from the candidate motion vector list, and the position indicated by the target motion vector is the target position. That is, the coding cost calculated based on the target motion vector is higher than the coding cost calculated based on the first motion vector, and the pixel distortion between the current block and the reference block at the target position indicated by the target motion vector is less than the set threshold.
[0086] In some examples, the coding efficiency can refer to the coding rate, i.e., the data traffic used by the encoder in a unit of time after encoding the video, or the coding cost can refer to the number of bits used to encode the motion vector.
[0087] In some examples, the pixel distortion refers to the difference between the pixels of the current block and the best position block / reference block searched in the reference frame. The difference can be the sum or average of the differences of individual pixels.
[0088] In some optional embodiments, if the index value in the candidate motion vector list is positively correlated with the value of the coding efficiency, a candidate motion vector set can be determined in the candidate motion vector list, where the index value is greater than the motion vector index of the current block. The candidate motion vector set includes candidate motion vectors with a coding efficiency greater than that of the original motion vector of the current block. Then, a candidate motion vector with a pixel distortion less than a set threshold can be selected from the candidate motion vector set as the target motion vector. Alternatively, a candidate motion vector with the smallest pixel distortion can be selected from the candidate motion vector set as the target motion vector.
[0089] In some optional embodiments, if the current block adopts the AMVP mode, the plurality of candidate positions are in a position region determined based on the motion vector residual around the position indicated by the MVP of the current block. In this case, the process of selecting the target position can be to search for a position with a coding efficiency higher than that of the original motion vector and a pixel distortion less than a set threshold in the position region, and then the searched position is the target position.
[0090] In some optional embodiments, if the amplitude of the motion vector residual is positively correlated with the value of the coding efficiency, a circular region centered at the position indicated by the MVP and having a radius of the amplitude of the motion vector residual can be determined, and then, in the region of the predetermined size centered at the position indicated by the MVP, a region outside the circular region can be taken as the target position region.
[0091] Optionally, in the target position region, a position with the minimum pixel distortion can be searched, and the searched position can be taken as the target position.
[0092] In some optional embodiments, if the sum of the absolute values of the horizontal and vertical coordinates of the motion vector residual is positively correlated with the value of the coding efficiency, a circular region centered at the position indicated by the MVP and having a radius of the sum of the absolute values of the horizontal and vertical coordinates of the motion vector residual can be determined, and then, in the region of the predetermined size centered at the position indicated by the MVP, a region outside the circular region can be taken as the target position region.
[0093] Optionally, in the target position region, a position with the minimum pixel distortion can be searched, and the searched position can be taken as the target position.
[0094] In S630, a second motion vector of the current block is determined based on the target position, and motion compensation is performed on the current block based on the second motion vector.
[0095] In some optional embodiments, after the target position is determined, a displacement between the position of the current block in the current image (or the current frame) and the target position can be taken as a new motion vector of the current block, and then, motion compensation can be performed on the current block based on the new motion vector. Since the coding efficiency (such as the coding rate) of the new motion vector is greater than the coding efficiency of the original motion vector of the current block, and the pixel distortion is also less than the set threshold, performing motion compensation on the current block based on the new motion vector can improve the effect of motion compensation, and thus is beneficial to improving the video coding efficiency.
[0096] In some optional embodiments, after the new motion vector of the current block is determined, the new motion vector of the current block can be stored as a temporal motion vector prediction (TMVP). The TMVP technique is a kind of MVP method in the time domain, which uses the motion information of the corresponding position of the current block in the already coded neighboring image (co-located image) to predict the motion vector of the current block, which is helpful to reduce the transmission amount of the motion vector data, and thus improve the coding efficiency.
[0097] In some optional embodiments, after determining the new motion vector of the current block, the original motion vector of the current block can be used as the MVP of the subsequent block; or can not be used as the MVP of the subsequent block, but only used for performing motion compensation on the current block.
[0098] FIG. 6 is an illustration of the technical solution of the embodiments of the present application from the perspective of video decoding. The technical solution of the embodiments of the present application is illustrated again from the perspective of video encoding in combination with FIG. 7.
[0099] FIG. 7 shows a flowchart of a video encoding method according to an embodiment of the present application, which can be executed by an electronic device with computing processing function, such as a terminal device or a server.
[0100] Referring to FIG. 7, the video encoding method at least includes S710 to S730, which are described in detail as follows.
[0101] In S710, a first motion vector of a current block is determined.
[0102] In S720, based on the first motion vector, a target position is selected from a plurality of candidate positions in a reference frame, wherein a coding cost calculated based on a motion vector indicated by the target position is higher than a coding cost calculated based on the first motion vector, and a pixel distortion between the current block and a reference block where the target position is located is less than a set threshold.
[0103] In S730, a second motion vector of the current block is determined based on the target position, and motion compensation is performed on the current block based on the second motion vector.
[0104] It should be noted that the processing process at the video encoding end is similar to the processing process at the video decoding end, and details can be referred to the aforementioned processing process at the decoding end, which will not be described herein.
[0105] In summary, the technical solution of the embodiments of the present application mainly searches in the direction of increasing the coding efficiency (such as coding rate) of the motion vector in the motion estimation process, so as to find a prediction position that can reduce the block matching distortion (i.e. pixel distortion), so as to improve the pixel prediction accuracy. Compared with the prior art, the total sum of the coding bitstream and the pixel distortion after weighting is not minimized, but the two are screened respectively.
[0106] After the decoding end obtains the prediction value and the residual data of the current block, the reconstructed block can be recovered, and then the reconstructed block is compared with the reference blocks at different positions in the reference frame to find the position with the minimum pixel distortion, so as to obtain the new motion vector of the current block.
[0107] In some optional embodiments, the decoder obtains the motion vector of the current block in two ways, depending on whether motion vector residual (MVD) exists.
[0108] If Merge mode is used, the positions in the Merge list can be re-evaluated, and the resulting MV is the final motion vector. In this case, there is no MVD. Alternatively, assuming that the larger the index value of the Merge list, the greater the coding efficiency, the position pointed to by the index after the index of the original MV of the current block (i.e., the coding efficiency is higher than the current index) can be evaluated. Then, compared with the position pointed to by the motion vector corresponding to the current index, the MV corresponding to the position with the smallest pixel distortion is selected as the new motion vector of the current block.
[0109] If the current block uses AMVP mode, the MVP value needs to be added to MVD to determine the motion vector. After calculating the coding efficiency of the current block, a search is performed in the region where the value is greater than the coding efficiency to match the reference block with smaller pixel distortion between it and the current block.
[0110] In one embodiment, the magnitude of MVD is positively correlated with the coding efficiency. For example, the coding efficiency increases with the magnitude of MVD. Increase and grow, MVD x and MVD y These represent the x-coordinate and y-coordinate of the MVD, respectively. As shown in Figure 8, in the reference image, a circular region is defined with the position indicated by the MVP as the center and the MVD amplitude as the radius, as shown by the dashed circle. A region of a predetermined size, centered on the position indicated by the MVP, is defined as the search range based on the MVP position, as shown by the square box. Within this search range, in the area outside the circular region (i.e., the search area), a reference block matching the current block is searched.
[0111] In another embodiment, the sum of the absolute values of the x and y axes of MVD is positively correlated with the coding efficiency. For example, coding efficiency increases with |MVD|. x |+|MVD y |Increases and increases, MVD x and MVD y Let x and y represent the x and y coordinates of MVD, respectively. Then, within the search range based on the MVP position shown in Figure 8 above, within a radius greater than |MVD| centered at the position indicated by the MVP, x |+|MVD y The | area (i.e., the search area) searches for reference blocks that match the current block.
[0112] The matching search process finds the optimal position with the minimum pixel distortion, and then determines the new motion vector of the current block based on the optimal position. The search process can be represented by the following formula: min{D(MVD(i))|MVD(i)εT,R(MVD(i) x ,MVD(i) y )>R(MVD x ,MVD y )}
[0113] wherein R(MVD x ,MVD y ) represents the encoding efficiency (such as the encoding code rate) at the decoded MVD position. D(MVD(i)) represents the pixel distortion between the reference block and the current block at the MVD(i) position, and the smaller the value is, the more similar the two are; and T represents the search region.
[0114] In some optional embodiments, the new motion vector can be used for motion compensation of the current block, and can also be used as the MVP of the future encoding block.
[0115] In some optional embodiments, the new motion vector can not be used for prediction of the subsequent encoding block of the current image, and the original decoded motion vector is still used as the prediction value of the subsequent encoding block.
[0116] In some optional embodiments, the new motion vector can be stored as a temporal MVP (TMVP) to predict the MV of the future image frame.
[0117] It should be noted that the technical solutions of the above-mentioned embodiments of the present application can reduce pixel distortion by optimizing the original motion vector of the current block, thereby improving the accuracy of the optimized motion vector and improving the video coding efficiency. The technical solutions of each of the above-mentioned embodiments can be used alone or in combination. Meanwhile, the technical solutions of the embodiments of the present application can be applied to video encoders or video compression related products.
[0118] The device embodiments of the present application are described below, which can be used to execute the methods described in the above-mentioned embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the above-mentioned method embodiments of the present application.
[0119] FIG. 9 shows a block diagram of a video decoding device according to an embodiment of the present application, which can be arranged in a device with computing processing function, such as a terminal device or a server.
[0120] Referring to FIG. 9, a video decoding apparatus 900 according to an embodiment of the present application includes a decoding unit 902, a selecting unit 904, and a processing unit 906.
[0121] The decoding unit 902 is configured to decode a video bitstream to obtain a first motion vector of a current block.
[0122] The selecting unit 904 is configured to select, based on the first motion vector, a target position from a plurality of candidate positions in a reference frame, wherein a coding efficiency calculated based on a motion vector indicated by the target position is higher than a coding efficiency calculated based on the first motion vector, and a pixel distortion between the current block and a reference block in which the target position is located is less than a set threshold.
[0123] The processing unit 906 is configured to determine a second motion vector of the current block based on the target position, and perform motion compensation on the current block based on the second motion vector.
[0124] In some embodiments of the present application, based on the foregoing scheme, the decoding unit 902 is configured to:
[0125] decode the video bitstream to obtain a motion vector index of the current block;
[0126] determine the first motion vector from a candidate motion vector list corresponding to the current block based on the motion vector index.
[0127] In some embodiments of the present application, based on the foregoing scheme, the plurality of candidate positions include positions indicated by each candidate motion vector in the candidate motion vector list.
[0128] The selecting unit 904 is configured to select, from the candidate motion vector list, a target motion vector, wherein a coding efficiency calculated based on the target motion vector is higher than a coding efficiency calculated based on the first motion vector, and a pixel distortion between the current block and a reference block in which the target position indicated by the target motion vector is located is less than a set threshold.
[0129] In some embodiments of the present application, based on the foregoing scheme, the selecting unit 904 is configured to:
[0130] when an index value in the candidate motion vector list is positively correlated with a value of the coding efficiency, determine, from the candidate motion vector list, a candidate motion vector set whose index value is greater than the motion vector index;
[0131] select, from the candidate motion vector set, a candidate motion vector whose pixel distortion is less than a set threshold as the target motion vector.
[0132] In some embodiments of the present application, based on the foregoing scheme, the selection unit 904 is configured to:
[0133] select, from the set of candidate motion vectors, a candidate motion vector with the minimum pixel distortion as the target motion vector.
[0134] In some embodiments of the present application, based on the foregoing scheme, the decoding unit 902 is configured to:
[0135] decode the video bitstream to obtain a motion vector prediction index and a motion vector residual of the current block;
[0136] based on the motion vector prediction index, determine a motion vector prediction of the current block in a candidate motion vector prediction list corresponding to the current block;
[0137] determine the first motion vector based on the motion vector prediction and the motion vector residual.
[0138] In some embodiments of the present application, based on the foregoing scheme, the selection unit 904 is configured to:
[0139] determine a position region based on the motion vector residual, with the position indicated by the motion vector prediction as the center;
[0140] search for the target position in the position region.
[0141] In some embodiments of the present application, based on the foregoing scheme, the selection unit 904 is configured to:
[0142] when the amplitude of the motion vector residual is positively correlated with the value of the coding efficiency, determine a circular region with the position indicated by the motion vector prediction as the center and with the amplitude of the motion vector residual as the radius;
[0143] in a region of a predetermined size centered at the position indicated by the motion vector prediction, regard a region outside the circular region as the position region.
[0144] In some embodiments of the present application, based on the foregoing scheme, the selection unit 904 is configured to:
[0145] when the sum of the absolute values of the horizontal and vertical coordinates of the motion vector residual is positively correlated with the value of the coding efficiency, determine a circular region with the position indicated by the motion vector prediction as the center and with the sum of the absolute values of the horizontal and vertical coordinates of the motion vector residual as the radius;
[0146] in a region of a predetermined size centered at the position indicated by the motion vector prediction, regard a region outside the circular region as the position region.
[0147] In some embodiments of the present application, based on the foregoing scheme, the selection unit 904 is configured to:
[0148] In the position region, a position with minimum pixel distortion is taken as the target position.
[0149] In some embodiments of the present application, based on the foregoing scheme, the processing unit 906 is further configured to:
[0150] The second motion vector is stored as a motion vector prediction in the time domain.
[0151] In some embodiments of the present application, based on the foregoing scheme, the processing unit 906 is further configured to:
[0152] The first motion vector is taken as a motion vector prediction of a subsequent block of the current block.
[0153] FIG. 10 shows a block diagram of a video encoding apparatus according to an embodiment of the present application, which can be arranged in a device with computing processing function, such as a terminal device or a server.
[0154] Referring to FIG. 10, a video encoding apparatus 1000 according to an embodiment of the present application includes a determination unit 1002, a selection unit 1004, and a processing unit 1006.
[0155] The determination unit 1002 is configured to determine a first motion vector of a current block.
[0156] The selection unit 1004 is configured to select, based on the first motion vector, a target position from a plurality of candidate positions in a reference frame, wherein a coding cost calculated based on a motion vector indicated by the target position is higher than a coding cost calculated based on the first motion vector, and a pixel distortion between the current block and a reference block where the target position is located is less than a set threshold; and
[0157] The processing unit 1006 is configured to determine a second motion vector of the current block based on the target position, and perform motion compensation on the current block based on the second motion vector.
[0158] FIG. 11 shows a structural schematic diagram of an electronic device suitable for implementing embodiments of the present application, which can be the video encoding apparatus or the video decoding apparatus in the foregoing embodiments.
[0159] It should be noted that the electronic device 1100 shown in FIG. 11 is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present application.
[0160] As shown in FIG. 11, the computer system 1100 can include a central processing unit (CPU) 1101 which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1102 or loaded into a random access memory (RAM) 1103 from a storage section 1108, such as the methods described in the above embodiments. Various programs and data required for the operation of the system are also stored in the RAM 1103. The CPU 1101, the ROM 1102, and the RAM 1103 are connected to each other through a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0161] The following components can be connected to the I / O interface 1105: an input section 1106 including input devices such as a keyboard and a mouse; an output section 1107 including output devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1108 including a hard disk; and a communication section 1109 including a network interface card such as a LAN (Local Area Network) card, a modem, and the like. The communication section 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to the I / O interface 1105 as necessary. A removable recording medium 1111 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is attached to the drive 1110 as necessary, so that a computer program read therefrom is installed into the storage section 1108 as necessary.
[0162] In particular, according to embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present application include a computer program product comprising a computer program for performing the methods illustrated by the flowcharts carried on a computer readable medium. In such embodiments, the computer program can be downloaded and installed from a network via the communication section 1109 and / or installed from the removable recording medium 1111. When the computer program is executed by the central processing unit (CPU) 1101, various functions defined in the system of the present application are performed.
[0163] It should be noted that the computer-readable medium in the embodiments of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having at least one conductive wire, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (Compact Disc Read-Only Memory, CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present application, the computer-readable storage medium can be any tangible medium containing or storing a computer program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present application, the computer-readable signal medium can include a data signal carrying a computer-readable computer program in a baseband or as a part of a carrier wave. Such a propagated data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or apparatus. The computer program contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to wireless, wired, or the like, or any suitable combination thereof.
[0164] The flowcharts and block diagrams in the drawings illustrate the possible implementation architectures, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In the flowcharts or block diagrams, each block can represent a module, a program segment or a part of code containing at least one executable instruction for implementing a specified logic function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in different orders than that shown in the drawings. For example, two blocks represented in succession can actually be executed substantially in parallel, and sometimes in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams or flowcharts, and the combination of blocks in the block diagrams or flowcharts, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer programs.
[0165] The units described in the embodiments of the present application can be implemented by software, or can be implemented by hardware, and the units described can also be arranged in a processor. In some cases, the names of the units do not constitute a limitation on the units themselves.
[0166] As another aspect, the present application also provides a computer readable medium, which can be included in the electronic device described in the above embodiments, or can exist independently without being assembled into the electronic device. The computer readable medium carries one or more computer programs, which, when executed by the electronic device, enable the electronic device to implement the method described in the above embodiments.
[0167] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into several modules or units.
[0168] From the above description of the embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.) or a network, and includes several instructions to enable an electronic device to perform the method according to the embodiments of the present application.
[0169] For example, the electronic device can be a video decoding apparatus, which can perform the video decoding method shown in FIG. 6; for another example, the electronic device can be a video encoding apparatus, which can perform the video encoding method shown in FIG. 7.
[0170] Other embodiments of the present application will be apparent to those skilled in the art from consideration of the specification and practice of the embodiments disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the application following, in general, the principles of the application and including such departures from the present disclosure as come within known or customary practice in the art to which the application pertains.
[0171] It is to be understood that the application is not limited to the precise construction already described above and shown in the drawings, and that various modifications and changes can be made by those skilled in the art without departing from the scope of the application. The scope of the application should only be limited by the claims appended hereto.
Claims
1. A method of video decoding performed by an electronic device, comprising: decoding a video bitstream to obtain a first motion vector of a current block; selecting, based on the first motion vector, a target position from a plurality of candidate positions in a reference frame, wherein a coding efficiency calculated based on a motion vector indicated by the target position is higher than a coding efficiency calculated based on the first motion vector, and a pixel distortion between the current block and a reference block in which the target position is located is less than a set threshold; and determining a second motion vector of the current block based on the target position, and performing motion compensation on the current block based on the second motion vector. The decoding of the video bitstream to obtain the first motion vector of the current block comprises:
2. The video decoding method of claim 1, wherein, decoding the video bitstream to obtain a motion vector index of the current block; determining the first motion vector from a candidate motion vector list corresponding to the current block based on the motion vector index. The plurality of candidate positions comprises positions indicated by each candidate motion vector in the candidate motion vector list.
3. The video decoding method of claim 2, wherein, The selecting, based on the first motion vector, the target position from the plurality of candidate positions in the reference frame comprises: selecting a target motion vector from the candidate motion vector list, wherein the coding efficiency calculated based on the target motion vector is higher than the coding efficiency calculated based on the first motion vector, and the pixel distortion between the current block and a reference block in which the target position indicated by the target motion vector is located is less than the set threshold. The selecting the target motion vector from the candidate motion vector list comprises:
4. The video decoding method of claim 3, wherein, when an index value in the candidate motion vector list is positively correlated with a value of the coding efficiency, determining a candidate motion vector set in which an index value is greater than the motion vector index from the candidate motion vector list; selecting, from the candidate motion vector set, a candidate motion vector in which the pixel distortion is less than the set threshold as the target motion vector. The selecting, from the candidate motion vector set, the candidate motion vector in which the pixel distortion is less than the set threshold as the target motion vector comprises:
5. The video decoding method of claim 4, wherein, selecting, from the candidate motion vector set, a candidate motion vector in which the pixel distortion is the smallest as the target motion vector. The decoding of the video bitstream to obtain the first motion vector of the current block comprises:
6. The video decoding method of claim 1, wherein, decoding the video bitstream to obtain a motion vector prediction index and a motion vector residual of the current block; determining a motion vector prediction of the current block from a candidate motion vector prediction list corresponding to the current block based on the motion vector prediction index; determining the first motion vector based on the motion vector prediction and the motion vector residual. The selecting, based on the first motion vector, the target position from the plurality of candidate positions in the reference frame comprises:
7. The video decoding method of claim 6, wherein, determining a position region based on the motion vector residual with a position indicated by the motion vector prediction as a center; searching the target position in the position region. The determining the position region based on the motion vector residual with the position indicated by the motion vector prediction as the center comprises:
8. The video decoding method of claim 7, wherein, determine a circular region centered at the position indicated by the motion vector prediction and having a radius of a magnitude of the motion vector residual when the magnitude of the motion vector residual is positively correlated with a value of the coding efficiency; determine a region outside the circular region as the position region in a region of a predetermined size centered at the position indicated by the motion vector prediction.
9. The video decoding method of claim 7, wherein, The determining the position region based on the motion vector residual and centered at the position indicated by the motion vector prediction includes: determine a circular region centered at the position indicated by the motion vector prediction and having a radius of a sum of absolute values of horizontal and vertical coordinates of the motion vector residual when the sum of absolute values of horizontal and vertical coordinates of the motion vector residual is positively correlated with a value of the coding efficiency; determine a region outside the circular region as the position region in a region of a predetermined size centered at the position indicated by the motion vector prediction.
10. The video decoding method of claim 8 or 9, wherein, The searching the target position in the position region includes: determine a position with a minimum pixel distortion in the position region as the target position.
11. The video decoding method of any of claims 1-9, further comprising: storing the second motion vector as a motion vector prediction in a temporal domain.
12. The video decoding method of any of claims 1-9, further comprising: using the first motion vector as a motion vector prediction for a subsequent block of the current block.
13. A video encoding method, wherein, including: determining a first motion vector of a current block; selecting a target position from a plurality of candidate positions in a reference frame based on the first motion vector, wherein a coding cost calculated based on a motion vector indicated by the target position is higher than a coding cost calculated based on the first motion vector, and a pixel distortion between the current block and a reference block in which the target position is located is less than a threshold value; and determining a second motion vector of the current block based on the target position, and performing motion compensation on the current block based on the second motion vector.
14. A video decoding apparatus, comprising: a decoding unit configured to decode a video bitstream to obtain a first motion vector of a current block; a selecting unit configured to select a target position from a plurality of candidate positions in a reference frame based on the first motion vector, wherein a coding cost calculated based on a motion vector indicated by the target position is higher than a coding cost calculated based on the first motion vector, and a pixel distortion between the current block and a reference block in which the target position is located is less than a threshold value; and a processing unit configured to determine a second motion vector of the current block based on the target position, and perform motion compensation on the current block based on the second motion vector. including: a determining unit configured to determine a first motion vector of a current block; 15. A video encoding apparatus, wherein, The selecting unit is configured to select a target position from a plurality of candidate positions in a reference frame based on the first motion vector, wherein a coding cost calculated based on a motion vector indicated by the target position is higher than a coding cost calculated based on the first motion vector, and a pixel distortion between the current block and a reference block in which the target position is located is less than a set threshold. And, The processing unit is configured to determine a second motion vector of the current block based on the target position, and perform motion compensation on the current block based on the second motion vector. 16.A computer readable medium having stored thereon a computer program, the computer program, when executed by a processor, implements the video decoding method of any one of claims 1 to 12, or the video encoding method of claim 13. 17.An electronic device comprising: at least one processor; a memory configured to store at least one computer program, the at least one computer program, when executed by the at least one processor, causes the electronic device to implement the video decoding method of any one of claims 1 to 12, or the video encoding method of claim 13. 18.A computer program product comprising a computer program stored in a computer readable storage medium, the computer program, when executed by a processor of an electronic device, causes the electronic device to perform the video decoding method of any one of claims 1 to 12, or the video encoding method of claim 13.
Citation Information
Patent Citations
Video coding and decoding method and device, computer readable medium and electronic equipment
CN116805968A
Video coding mode decision method and device, storage medium and electronic equipment
CN118138753A
Area division coder for moving picture
JP2000050279A
A method for video encoding using motion estimation and an apparatus thereof
KR1020170044599A
Video encoding method and device
WO2021163862A1