Video coding method and device, video decoding method and device, electronic equipment and storage medium
By constructing a candidate list of multiple hypothetical motion vectors and selecting the optimal motion vector for encoding and decoding, the problem of low encoding and decoding accuracy in existing technologies is solved, achieving higher prediction accuracy and encoding compression rate.
Patent Information
- Application Number
- CN202511179846.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-11-14
AI Technical Summary
The problem of low encoding and decoding accuracy in existing video encoding and decoding technologies.
By fusing any pair of motion vectors in the motion vector candidate list, a multi-hypothesis motion vector candidate list is constructed. The optimal motion vector is then selected for encoding and decoding, increasing the number of candidate motion vectors and improving accuracy by utilizing temporal pixel correlation.
It improves the accuracy of motion vectors, enhances prediction accuracy and encoding/decoding accuracy, and increases the encoding compression ratio without increasing the bitrate.
Smart Images

Figure CN120956903A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of video processing, and in particular to a video encoding method and apparatus, a video decoding method and apparatus, an electronic device, and a storage medium. Background Technology
[0002] Various electronic devices (such as digital televisions, laptops or desktop computers, tablets, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video conferencing equipment, video streaming devices, etc.) support digital video. Electronic devices send and receive, or otherwise transmit, digital video data via communication networks, and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited storage resources of storage devices, video data can be compressed using one or more video codec standards before being transmitted or stored. For example, video codec standards include Universal Video Codec (VVC), High Efficiency Video Codec (HEVC / H.265), and Advanced Video Codec (AVC / H.264). Video codecs typically employ prediction methods that utilize the inherent redundancy in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video codecs aim to compress video data to a lower bitrate while avoiding or minimizing degradation in video quality. Summary of the Invention
[0003] This disclosure provides a video encoding method and apparatus, a video decoding method and apparatus, an electronic device, and a storage medium to at least solve the problem of low encoding and decoding accuracy in related technologies.
[0004] According to a first aspect of the present disclosure, a video encoding method is provided, applied at an encoding end, comprising: fusing any two motion vectors in a motion vector candidate list of a current block to obtain multiple hypothetical motion vectors, and constructing a multiple hypothetical motion vector candidate list; selecting the optimal motion vector from the multiple hypothetical motion vector candidate list and the motion vector candidate list; and encoding the current block based on the optimal motion vector.
[0005] According to a second aspect of the present disclosure, a video decoding method is provided, applied at a decoding end, comprising: in response to determining that a multi-hypothesis fusion mode is enabled for the current block, fusing any two motion vectors in the motion vector candidate list of the current block to obtain a multi-hypothesis motion vector, and constructing a multi-hypothesis motion vector candidate list; selecting the optimal motion vector from the multi-hypothesis motion vector candidate list; and decoding the current block based on the optimal motion vector.
[0006] According to a third aspect of the present disclosure, a video encoding apparatus is provided, comprising: a first acquisition unit configured to fuse any two motion vectors in a motion vector candidate list of a current block to obtain multiple hypothetical motion vectors, and construct a multiple hypothetical motion vector candidate list; a first selection unit configured to select the optimal motion vector from the multiple hypothetical motion vector candidate list and the motion vector candidate list; and an encoding unit configured to encode the current block based on the optimal motion vector.
[0007] According to a fourth aspect of the present disclosure, a video decoding apparatus is provided, comprising: a second acquisition unit configured to, in response to determining that a multi-hypothesis fusion mode is enabled for a current block, fuse any two motion vectors in a motion vector candidate list for the current block to obtain a multi-hypothesis motion vector, and construct a multi-hypothesis motion vector candidate list; a second selection unit configured to select an optimal motion vector from the multi-hypothesis motion vector candidate list; and a decoding unit configured to decode the current block based on the optimal motion vector.
[0008] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and at least one memory storing computer-executable instructions, wherein the computer-executable instructions, when executed by the at least one processor, cause the at least one processor to perform the video encoding method and / or video decoding method as described above.
[0009] According to a sixth aspect of the present disclosure, a computer-readable storage medium is provided that, when instructions in the computer-readable storage medium are executed by at least one processor, causes the at least one processor to execute a bitstream generated by the video encoding method and / or a bitstream decoded by the video decoding method as described above.
[0010] According to a seventh aspect of the present disclosure, a computer program product is provided having instructions for storing a bitstream, the bitstream including: video data generated by the video encoding method described above and / or video data decoded by the video decoding method.
[0011] According to an eighth aspect of the present disclosure, a method for storing a bitstream is provided, comprising: generating a bitstream according to the video encoding method described above; and storing the bitstream.
[0012] According to an eighth aspect of the present disclosure, a method for transmitting a bitstream is provided, comprising: generating a bitstream according to the video encoding method described above; and transmitting the bitstream.
[0013] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects: According to the video coding method and apparatus, video decoding method and apparatus, electronic device, and storage medium disclosed herein, the number of candidate motion vectors used to determine the optimal motion vector is increased, and the newly added candidate motion vectors (i.e., multiple hypothetical motion vectors) have different reference frames than the existing motion vectors; that is, the newly added candidate motion vectors have two reference frames. This allows for the utilization of temporal pixel correlation to improve the accuracy of the candidate motion vectors, thereby improving the accuracy of the optimal motion vector, and thus improving prediction accuracy and encoding / decoding accuracy. Therefore, this disclosure solves the problem of low encoding / decoding accuracy in related technologies.
[0014] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0016] Figure 1 This is a block diagram illustrating an exemplary system for encoding and decoding video blocks according to exemplary embodiments of the present disclosure; Figure 2 This is a block diagram of a video encoder illustrated according to exemplary embodiments of the present disclosure; Figure 3 This is a block diagram of a video decoder illustrated according to exemplary embodiments of the present disclosure; Figure 4 This is a schematic diagram illustrating bidirectional Merge prediction according to exemplary embodiments of the present disclosure; Figure 5 This is a flowchart illustrating a video encoding method according to exemplary embodiments of the present disclosure; Figure 6 This is a schematic diagram illustrating a method for determining template matching cost for multiple hypothetical motion vectors according to exemplary embodiments of the present disclosure; Figure 7 This is a schematic diagram illustrating the fusion process for obtaining multiple hypothetical motion vectors according to exemplary embodiments of the present disclosure; Figure 8 This is a schematic diagram illustrating a modified motion vector according to an exemplary embodiment of the present disclosure; Figure 9 This is a flowchart illustrating a video decoding method according to exemplary embodiments of the present disclosure; Figure 10 This is a block diagram illustrating a video encoding apparatus according to exemplary embodiments of the present disclosure; Figure 11This is a block diagram illustrating a video decoding apparatus according to exemplary embodiments of the present disclosure; Figure 12 This is a diagram illustrating a computing environment coupled to a user interface according to exemplary embodiments of the present disclosure. Detailed Implementation
[0017] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0018] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following examples do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0019] It should be noted that the phrase "at least one of several items" in this disclosure refers to three parallel cases: "any one of the several items", "a combination of any number of the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel cases: (1) including A; (2) including B; (3) including A and B. Another example is "performing at least one of step one and step two", which means the following three parallel cases: (1) performing step one; (2) performing step two; (3) performing both step one and step two.
[0020] Figure 1 This is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel, according to some embodiments of the present disclosure. Figure 1 As shown, system 10 includes a source device 12 that generates and encodes video data that will later be decoded by a target device 14. The source device 12 and target device 14 can include any electronic device from a wide variety of electronic devices, including cloud servers, server computers, desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some embodiments, the source device 12 and target device 14 are equipped with wireless communication capabilities.
[0021] In some implementations, the target device 14 may receive the encoded video data to be decoded via link 16. Link 16 may include any type of communication medium or device capable of moving the encoded video data from the source device 12 to the target device 14.
[0022] In some other implementations, the encoded video data can be sent from the output interface 22 to the storage device 32. Subsequently, the target device 14 can access the encoded video data in the storage device 32 via the input interface 28.
[0023] like Figure 1 As shown, source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The video source 18 may include sources or combinations of such sources, such as: a video capture device (e.g., a camera), a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video.
[0024] The captured, pre-captured, or computer-generated video can be encoded by the video encoder 20. The encoded video data can be sent directly to the target device 14 via the output interface 22 of the source device 12. Alternatively, the encoded video data can be stored on the storage device 32 for later access by the target device 14 or other devices for decoding and / or playback.
[0025] Target device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 may include a receiver and / or a modem, and receives encoded video data via link 16. The encoded video data transmitted via link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 when decoding the video data. Such syntax elements may be included within encoded video data transmitted on a communication medium, stored on a storage medium, or stored on a file server.
[0026] Video encoder 20 and video decoder 30 can operate according to proprietary or industry standards (e.g., VVC, HEVC, MPEG-4 Part 10, AVC) or extensions of such standards. It should be understood that this application is not limited to any particular video encoding / decoding standard and can be applied to other video encoding / decoding standards. It is generally understood that the video encoder 20 of source device 12 can be configured to encode video data according to any of these current or future standards. Similarly, it is also generally understood that the video decoder 30 of target device 14 can be configured to decode video data according to any of these current or future standards.
[0027] The video encoder 20 and video decoder 30 can be implemented as any circuit of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic devices, software, hardware, firmware, or any combination thereof. When partially implemented in software, the electronic device may store instructions for software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in hardware to perform the video encoding / decoding operations disclosed in this disclosure. Each of the video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, and either encoder or decoder may be integrated as part of a combined encoder / decoder (CODEC) in the respective device.
[0028] Figure 2 This is a block diagram illustrating an exemplary video encoder 20 according to some embodiments described in this application. The video encoder 20 can perform intra-frame predictive coding and inter-frame predictive coding on video blocks within a video frame. Intra-frame predictive coding relies on spatial prediction to reduce or remove spatial redundancy in the video data within a given video frame or picture. Inter-frame predictive coding relies on temporal prediction to reduce or remove temporal redundancy in the video data within neighboring video frames or pictures of a video sequence. It should be noted that in the field of video encoding and decoding, the term "frame" can be used as a synonym for the terms "image" or "picture".
[0029] like Figure 2As shown, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a segmentation unit 45, an intra-frame prediction processing unit 46, and an intra-frame block copy (BC) unit 48. In some embodiments, the video encoder 20 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62 for video block reconstruction. A loop filter 63, such as a deblocking filter, can be located between the adder 62 and the DPB 64 to filter block boundaries to remove block artifacts from the reconstructed video. In addition to the deblocking filter, another loop filter (e.g., a sample adaptive offset (SAO) filter, a cross-component sample adaptive offset (CCSAO) filter, and / or an adaptive loop filter (ALF)) can be used to filter the output of the adder 62. In some examples, the loop filter can be omitted, and the decoded video block can be directly provided to the DPB 64 by the adder 62. The video encoder 20 can take the form of a fixed or programmable hardware unit, or it can be distributed among one or more of the fixed or programmable hardware units described.
[0030] The video data storage device 40 can store video data encoded by the components of the video encoder 20. For example, it can store data from... Figure 1 The video source 18 shown obtains video data from the video data storage 40. The DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) used by the video encoder 20 (e.g., in intra-frame or inter-frame predictive coding mode) when encoding the video data.
[0031] like Figure 2 As shown, after receiving video data, the segmentation unit 45 within the prediction processing unit 41 segments the video data into video blocks. This segmentation may also include segmenting the video frame into strips, tiles (e.g., a collection of video blocks) or other larger coding units (CUs) according to a predefined splitting structure (e.g., a quadtree (QT) structure) associated with the video data. It should be noted that the term "block" or "video block" as used herein can be a portion of a frame or image, particularly a rectangular (square or non-square) portion. Referring to, for example, HEVC and VVC, a block or video block can be or corresponds to a coding tree unit (CTU), CU, prediction unit (PU), or transform unit (TU) and / or can be or corresponds to a corresponding block (e.g., coding tree block (CTB), coding block (CB), prediction block (PB), or transform block (TB)) and / or sub-block.
[0032] The prediction processing unit 41 can select one of several feasible predictive coding modes for the current video block based on error results (e.g., coding rate and distortion level), such as one or more inter-frame predictive coding modes among multiple intra-frame predictive coding modes. The prediction processing unit 41 can provide the resulting intra-frame or inter-frame predictive coding block to adder 50 to generate a residual block, and to adder 62 to reconstruct the coding block for subsequent use as part of a reference frame. The prediction processing unit 41 also provides syntax elements (e.g., motion vectors, intra-frame mode indicators, segmentation information, and other such syntax information) to entropy coding unit 56.
[0033] To select a suitable intra-predictive coding mode for the current video block, the intra-predictive processing unit 46 within the prediction processing unit 41 can perform intra-predictive coding of the current video block in relation to one or more neighboring blocks in the same frame as the current block to be encoded to provide spatial prediction. The motion estimation unit 42 and motion compensation unit 44 within the prediction processing unit 41 perform inter-predictive coding of the current video block in relation to one or more prediction blocks in one or more reference frames to provide temporal prediction. The video encoder 20 can perform multiple coding passes, for example, to select a suitable coding mode for each block of video data.
[0034] In some implementations, motion estimation unit 42 determines an inter-frame prediction mode for the current video frame by generating motion vectors based on a predetermined pattern within the video frame sequence. The motion vectors indicate the displacement of a video block within the current video frame relative to a prediction block within a reference video frame. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimate the motion of video blocks. For example, the motion vectors may indicate the displacement of a video block within the current video frame or picture relative to a prediction block within a reference frame associated with the current block being encoded in the current frame. The predetermined pattern may designate video frames in the sequence as P-frames or B-frames. Intra-frame BC unit 48 may determine vectors (e.g., block vectors) for intra-frame BC coding in a similar manner to how motion estimation unit 42 determines motion vectors for inter-frame prediction, or the block vectors may be determined using motion estimation unit 42.
[0035] Regardless of whether the predicted block comes from the same frame predicted intra-frame or from different frames predicted inter-frame, the video encoder 20 can form a residual video block by subtracting the pixel values of the predicted block from the pixel values of the current video block being encoded. The pixel difference forming the residual video block can include both luma component difference and chroma component difference.
[0036] Intra-prediction processing unit 46 can encode the current block using various intra-prediction modes, for example, during individual encoding passes, and intra-prediction processing unit 46 (or, in some examples, mode selection unit) can select a suitable intra-prediction mode from the tested intra-prediction modes for use. Intra-prediction processing unit 46 can provide information indicating the intra-prediction mode selected for the block to entropy coding unit 56. Entropy coding unit 56 can encode the information indicating the selected intra-prediction mode into the bitstream.
[0037] After prediction processing unit 41 determines the prediction block for the current video block via inter-frame prediction or intra-frame prediction, adder 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and provided to transform processing unit 52. Transform processing unit 52 uses a transform (e.g., discrete cosine transform (DCT) or a conceptually similar transform) to transform the residual video data into residual transform coefficients.
[0038] The transform processing unit 52 can send the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process can also reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting the quantization parameters. In some examples, the quantization unit 54 can subsequently perform a scan on the matrix including the quantized transform coefficients. Alternatively, the entropy coding unit 56 can perform the scan.
[0039] After quantization, the entropy coding unit 56 entropy codes the quantized transform coefficients into a video bitstream using, for example, context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probabilistic interval segmented entropy (PIPE) coding, or another entropy coding method or technique. The encoded bitstream can then be sent to, for example,... Figure 1 The video decoder 30 shown, or archived in, for example Figure 1 The data is stored in storage device 32 for later transmission to or retrieval by video decoder 30. Entropy coding unit 56 can also entropy code the motion vectors and other syntax elements used for the current video frame being encoded.
[0040] The inverse quantization unit 58 and the inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct residual video blocks in the pixel domain for generating reference blocks to predict other video blocks. As noted above, the motion compensation unit 44 can generate motion-compensated prediction blocks from one or more reference blocks of frames stored in the DPB 64. The motion compensation unit 44 can also apply one or more interpolation filters to the prediction blocks to compute sub-integer pixel values for use in motion estimation.
[0041] Adder 62 adds the reconstructed residual block to the motion-compensated prediction block generated by motion compensation unit 44 to generate a reference block for storage in DPB 64. The reference block can then be used as a prediction block by intra-frame BC unit 48, motion estimation unit 42, and motion compensation unit 44 for inter-frame prediction of another video block in subsequent video frames.
[0042] Figure 3 This is a block diagram illustrating an exemplary video decoder 30 according to some embodiments of this application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra-frame prediction unit 84, and an intra-frame BC unit 85. The video decoder 30 can perform operations in conjunction with the above. Figure 2 The decoding process described for the video encoder 20 is essentially the inverse of the encoding process. For example, the motion compensation unit 82 can generate prediction data based on the motion vectors received from the entropy decoding unit 80, while the intra-frame prediction unit 84 can generate prediction data based on the intra-frame prediction mode indicator received from the entropy decoding unit 80.
[0043] In some examples, embodiments of this disclosure may be distributed across one or more units of the video decoder 30. For example, the intra-frame prediction (BC) unit 85 may perform embodiments of this application individually or in combination with other units of the video decoder 30, such as the motion compensation unit 82, the intra-frame prediction unit 84, and the entropy decoding unit 80. In some examples, the video decoder 30 may not include the intra-frame prediction (BC) unit 85, and the functionality of the intra-frame prediction (BC) unit 85 may be performed by other components of the prediction processing unit 81, such as the motion compensation unit 82.
[0044] The video data storage device 79 can store video data, such as encoded video bitstreams, that will be decoded by other components of the video decoder 30. The video data stored in the video data storage device 79 can be obtained, for example, from the storage device 32, from a local video source (e.g., a camera), via wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or hard disk).
[0045] During the decoding process, the video decoder 30 receives a encoded video bitstream representing video blocks and associated syntax elements of encoded video frames. The entropy decoding unit 80 of the video decoder 30 performs entropy decoding on the bitstream to generate quantization coefficients, motion vectors or intra-prediction mode indicators, and other syntax elements. The entropy decoding unit 80 then forwards the motion vectors or intra-prediction mode indicators, and other syntax elements to the prediction processing unit 81.
[0046] When a video frame is encoded as an intra-predictive coded (I) frame or used as an intra-coded prediction block in other types of frames, the intra-predictive unit 84 of the prediction processing unit 81 can generate prediction data for the video block of the current video frame based on the intra-predictive mode transmitted by the signal and reference data from the previous decoded block of the current frame.
[0047] When a video frame is encoded as an inter-frame predictive coded (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks for the current video frame based on motion vectors and other syntax elements received from the entropy decoding unit 80. Each of the prediction blocks can be generated from a reference frame within a reference frame list. The video decoder 30 can construct the reference frame list, i.e., list 0 and list 1, based on the reference frames stored in the DPB 92 using a default construction technique.
[0048] In some examples, when a video block is encoded according to the intra-frame BC mode described herein, the intra-frame BC unit 85 of the prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from the entropy decoding unit 80. The prediction block can be located within a reconstructed region of the same image as the current video block, as defined by the video encoder 20.
[0049] The motion compensation unit 82 and / or the intra-frame BC unit 85 determine the prediction information for the video block of the current video frame by parsing motion vectors and other syntax elements, and then use the prediction information to generate a prediction block for the current video block being decoded.
[0050] The motion compensation unit 82 can also perform interpolation using interpolation filters, such as those used by the video encoder 20 during encoding of video blocks, to calculate interpolated values for sub-integer pixels of the reference block. In this case, the motion compensation unit 82 can determine the interpolation filters used by the video encoder 20 based on the received syntax elements and use these interpolation filters to generate the prediction block.
[0051] The dequantization unit 86 dequantizes the quantized transform coefficients provided in the bitstream and entropy-decoded by the entropy decoding unit 80 using the same quantization parameters calculated by the video encoder 20 for each video block in the video frame to determine the degree of quantization. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT, inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients in order to reconstruct the residual block in the pixel domain.
[0052] After the motion compensation unit 82 or the intra-frame BC unit 85 generates a prediction block for the current video block based on vectors and other syntax elements, the adder 90 reconstructs the decoded video block for the current video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82 and the intra-frame BC unit 85. A loop filter 91 (e.g., a deblocking filter, SAO filter, CCSAO filter, and / or ALF) may be located between the adder 90 and the DPB 92 for further processing of the decoded video block. In some examples, the loop filter 91 may be omitted, and the decoded video block may be directly provided to the DPB 92 by the adder 90. The decoded video block in a given frame is then stored in the DPB 92, which stores reference frames for subsequent motion compensation of the next video block. The DPB 92 or a separate memory device may also store the decoded video for later presentation on a display device (e.g., ...). Figure 1 On the display device 34).
[0053] Currently, VVC (Versatile Video Coding) / H.266 is the latest generation video codec standard released by the international standardization organization JPEG in 2021. Its compression rate is reduced by about 40% compared to the HEVC (High Efficiency Video Coding) / H.265 standard, making it the most advanced video codec standard at present. The Enhanced Compression Model (ECM) is the code development platform for the next generation standard under development after the VVC standard. Template Matching (TM) is an important technology adopted in the prediction module of ECM. It is mainly used in inter-frame prediction to refine motion vectors or sort prediction candidates. Its purpose is to use the correlation between adjacent blocks in the spatial domain to hide the coding syntax overhead.
[0054] Merge is a motion vector (MV) prediction technique in coding standards. It uses the MVs of neighboring blocks in the temporal or spatial domains to predict the MV of the current block. In the VVC standard, Merge mode creates a candidate list of motion vectors for the current coding unit (CU). This list contains information about five candidate motion vectors, each including MV, prediction direction, reference frame index, and motion vector precision. By iterating through these five candidate motion vectors, the candidate motion vector with the lowest rate-distortion cost is selected as the candidate motion vector for the current CU. The bitstream does not need to transmit the motion vector information itself (such as the reference frame index, motion vector precision index, and prediction direction); only the index of the optimal candidate motion vector in the candidate list needs to be transmitted, thus saving coding overhead.
[0055] Merge mode can consider motion vectors in up to two directions, i.e., bidirectional merge prediction, such as... Figure 4 As shown, the predicted pixel value in bidirectional Merge mode is obtained by weighting the motion compensation values of forward MV0 and backward MV1. Bidirectional prediction is a multi-hypothesis prediction compared to unidirectional prediction, meaning that the prediction effect when the current block references multiple reference frames is better than the prediction effect when referencing a single reference frame. However, VVC Merge only supports bidirectional prediction at most. But the content of consecutive frames in a video is consistent, so theoretically, increasing the number of reference frames in Merge mode can achieve better prediction accuracy. In other words, there is room for performance optimization in the existing Merge technology.
[0056] To address the aforementioned issues, this disclosure proposes a video encoding / decoding method that increases the number of candidate motion vectors used to determine the optimal motion vector. Furthermore, the newly added candidate motion vectors (i.e., multi-hypothesis motion vectors) have different reference frames than existing motion vectors; that is, each newly added candidate motion vector has two reference frames. This allows for the utilization of temporal pixel correlation to improve the accuracy of candidate motion vectors, thereby improving the accuracy of the optimal motion vector, and consequently, improving prediction accuracy and encoding / decoding accuracy. Moreover, this disclosure uses template matching cost to filter the fused multi-hypothesis motion vectors, minimizing the loss incurred by adding multi-hypothesis motion vectors to the multi-hypothesis motion vector list and significantly increasing the probability of multi-hypothesis motion vectors being selected. The template matching method ensures that the number of motion vectors in the multi-hypothesis motion vector list is the same as in the ordinary motion vector list, allowing direct use of the merge index in the ordinary merge mode without increasing the range of the original merge index. This enables multi-hypothesis merge prediction without explicitly increasing the merge candidate index, meaning no additional bitrate is consumed during encoding, thus improving both encoding accuracy and compression ratio. Furthermore, this disclosure sets a fusion mode identifier for each current block, so that when encoding the current block in subsequent iterations, only the identifier at the CU level can be encoded, avoiding the addition of other syntax information, thereby reducing the encoding compression ratio.
[0057] Hereinafter, video encoding methods and apparatus, video decoding methods and apparatus, electronic devices, and storage media according to exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.
[0058] Figure 5 This is a flowchart illustrating a video encoding method according to exemplary embodiments of the present disclosure, such as... Figure 5 As shown, the video encoding method includes the following steps: In step S501, any two motion vectors in the current block's motion vector candidate list are fused to obtain multiple hypothetical motion vectors, and a multiple hypothetical motion vector candidate list is constructed.
[0059] Specifically, before step S501, the motion vector candidate list for the current block can be constructed in the order of spatially adjacent candidates, temporal candidates, spatially non-adjacent candidates, historical candidates, and chained candidates; or the motion vector candidate list for the current block can be constructed in the order of spatially adjacent candidates, temporal candidates, and historical candidates. This disclosure does not limit the scope of the proposed method.
[0060] It should be noted that spatial neighbor candidates refer to motion information (such as motion vectors, reference frame indices, etc.) obtained from spatially adjacent coded blocks of the current block (CU). These neighboring blocks are directly adjacent to the current block in spatial location and are the most prioritized type of merge candidate. Temporal candidates refer to motion information obtained from blocks in the reference frame that have a temporal correlation with the current block's location (usually called "collocated blocks"). Spatial non-neighbor candidates refer to motion information from blocks that are not directly adjacent to the current block in space but may have a motion correlation (such as blocks in the same region that are slightly farther away). Historical candidates refer to candidates extracted from motion information cached in encoded historical blocks. These historical blocks may not have a direct spatial or temporal relationship with the current block, but their motion patterns may repeat in the current frame (such as periodically moving objects). Chained candidates refer to new candidates obtained by performing secondary derivation on the motion information of existing candidates. For example, two existing candidate motion vectors can be combined (such as by addition or averaging) to generate a new motion vector as a candidate, or the possible motion direction can be further predicted based on the motion trajectory of existing candidates.
[0061] According to an exemplary embodiment of this disclosure, the construction of the multi-hypothesis motion vector candidate list in step S501 can be implemented as follows: For each multi-hypothesis motion vector obtained by fusion, the following processing is performed: determining a first template matching cost and two second template matching costs, wherein the first template matching cost is the template matching cost calculated for the current multi-hypothesis motion vector, and the two second template matching costs are the template matching costs calculated for the two motion vectors used to fuse the current multi-hypothesis motion vector; in response to the first template matching cost being less than each of the second template matching costs, the current multi-hypothesis motion vector is added to the multi-hypothesis motion vector candidate list. Through this embodiment, the multi-hypothesis motion vectors obtained by fusion are filtered based on the template matching cost, so that the loss caused by the multi-hypothesis motion vectors added to the multi-hypothesis motion vector list is small, greatly increasing the probability of the multi-hypothesis motion vectors being selected in the multi-hypothesis motion vector list. Moreover, the template matching method can ensure that the number of motion vectors in the multi-hypothesis motion vector list is the same as that in the ordinary motion vector list, so that the index in the ordinary Merge mode can be directly used without increasing the range of the original Merge index, so that no additional bitrate is consumed during encoding, and the encoding compression ratio can be improved while improving the encoding accuracy.
[0062] As an example, after obtaining a multi-hypothesis motion vector, first calculate the template matching cost (e.g., TM cost) of the multi-hypothesis motion vector and the template matching costs of the two motion vectors that are fused to obtain the multi-hypothesis motion vector; then, compare the TM cost of the multi-hypothesis motion vector with the TM costs of the two corresponding motion vectors. If the TM cost of the multi-hypothesis motion vector is less than the TM costs of the two corresponding motion vectors, then set the MHMV flag of the multi-hypothesis motion vector to 1 (default is 0) and add it to the candidate list of multi-hypothesis motion vectors; otherwise, skip to continue verifying the next multi-hypothesis motion vector.
[0063] The following sections describe the process of determining the first template matching cost and the second template matching cost: According to an exemplary embodiment of this disclosure, two second template matching costs can be determined as follows: For each of the two motion vectors, based on the position of the template of the current block and the current motion vector, a predicted template corresponding to the current motion vector is determined, wherein the position of the predicted template is the position of the template of the current block on the corresponding reference frame under the direction of the current motion vector; the rate-distortion cost of the pixels of the predicted template and the template of the current block is determined as the template matching cost of the current motion vector. This embodiment allows for convenient and rapid determination of the template matching cost of multiple hypothetical motion vectors.
[0064] As an example, the template for the current block can be selected as a reconstructed block with a width of 4 on the left and a height of 4 on the top in the current frame; this disclosure does not limit this. After obtaining the template for the current block, the corresponding position of the template in the reference frame can be obtained based on the corresponding MV, that is, the position of the predicted template can be obtained; then, the sum of absolute differences (Sum of Absolute Difference, abbreviated as sad) between the pixels at the template position and the pixels at the predicted template position is calculated as the template matching cost corresponding to the MV.
[0065] It should be noted that after obtaining the position of the prediction template, this position can also be used as the center point of the search. The TM cost of each search position can be calculated within a certain range, and the position with the minimum TM cost can be selected as the final position of the prediction template. This disclosure does not limit this.
[0066] According to an exemplary embodiment of this disclosure, the first template matching cost can be determined as follows: For each of the two motion vectors, based on the position of the template of the current block and the current motion vector, a predicted template corresponding to the current motion vector is determined, wherein the position of the predicted template is the position of the template of the current block on the corresponding reference frame under the direction of the current motion vector; the pixels of the predicted template corresponding to each motion vector are weighted to obtain weighted pixels; the rate-distortion cost of the weighted pixels and the pixels of the template of the current block is determined as the first template matching cost of the current multi-hypothesis motion vector. Through this embodiment, the template matching cost of multi-hypothesis motion vectors can be determined conveniently and quickly.
[0067] As an example, Figure 6 This demonstrates a method for determining the template matching cost for multiple hypothesis motion vectors, such as... Figure 6 As shown, firstly, the two MVs (MV0 and MV1) that have been fused to obtain the multi-hypothesis motion vector are in the current frame (i.e., Figure 6 The reference frame is pointed to based on the template of the current block in the middle (cur pic). Figure 6 Prediction templates on (ref pic) (i.e.) Figure 6 The pixels of the interpolation templates of MV0 and MV1 are weighted to obtain the weighted pixels (i.e., the interpolation templates of MV0 and MV1). Figure 6 (Multiple hypothesis interpolation template); then, calculate the rate-distortion cost of the weighted processed pixel and the pixel of the template of the current block, and determine the rate-distortion cost as the template matching cost of the multiple hypothesis motion vector.
[0068] It should be noted that, Figure 6 For ease of illustration, one-way MV is used for the two MVs. In actual application, the prediction direction of these two MVs is not restricted.
[0069] According to exemplary embodiments of this disclosure, the pixels of the prediction template corresponding to each motion vector are weighted to obtain weighted pixels. This can include, but is not limited to: performing average weighting on the pixels of the prediction template corresponding to each motion vector to obtain weighted pixels; or, weighting the pixels of the prediction template corresponding to each motion vector based on the template matching cost of each motion vector to obtain weighted pixels. In this embodiment, average weighting is relatively simple, while loss-based weighting ensures that the weighted pixels are closer to the actual situation. These two methods allow for flexible acquisition of the required weighted pixels.
[0070] According to an exemplary embodiment of this disclosure, before determining the first template matching cost and two second template matching costs, in response to the fact that the reference frames corresponding to the two motion vectors are the same, the current multi-hypothesis motion vector is skipped and the next multi-hypothesis motion vector is obtained; in response to the fact that the reference frames corresponding to the two motion vectors are the same and the difference information between the two motion vectors is not greater than a preset threshold, the current multi-hypothesis motion vector is skipped and the next multi-hypothesis motion vector is obtained. Considering that there is only one block closest to the current block on a reference frame, in order to avoid obtaining two coded blocks on the same reference frame, since one of them will inevitably be inaccurate, before determining the two template matching costs, it can be ensured that the reference frames corresponding to the two motion vectors are different, so that the multi-hypothesis motion vectors added to the multi-hypothesis motion vector list are relatively accurate; furthermore, in order to avoid deleting too many multi-hypothesis motion vectors, the current multi-hypothesis motion vector can be skipped only when the reference frames corresponding to the two motion vectors are the same and the difference information between the two motion vectors is greater than a preset threshold. This can avoid obtaining two blocks that are close to each other on the same reference frame, that is, avoid the situation where one is wrong and the other will inevitably be wrong.
[0071] As an example, after obtaining a multi-hypothesis motion vector, it can be determined whether the multi-hypothesis motion vector needs to be filtered out, and then the template matching cost of the multi-hypothesis motion vector and the corresponding two motion vectors can be calculated. Specifically, the filtering process can be as follows: if the reference frames of the two motion vectors obtained by fusing the multi-hypothesis motion vector are the same, the multi-hypothesis motion vector can be skipped and the next multi-hypothesis motion vector can be obtained. In this way, inaccurate multi-hypothesis motion vectors can be filtered out. Alternatively, if the reference frames of the two motion vectors obtained by fusing the multi-hypothesis motion vector are the same and the difference between the two motion vectors is not greater than a preset threshold, the multi-hypothesis motion vector can be skipped and the next multi-hypothesis motion vector can be obtained. In this way, relatively inaccurate multi-hypothesis motion vectors can be filtered out, and too many multi-hypothesis motion vectors can be deleted.
[0072] According to an exemplary embodiment of this disclosure, fusing any pair of motion vectors in the motion vector candidate list of the current block to obtain multiple hypothetical motion vectors may include: concatenating any pair of motion vectors in the motion vector candidate list of the current block to obtain multiple hypothetical motion vectors; or, weighting any pair of motion vectors in the motion vector candidate list of the current block to obtain multiple hypothetical motion vectors. These two methods allow for flexible fusing as needed to obtain the desired multiple hypothetical motion vectors.
[0073] As an example, a multi-hypothesis motion vector candidate list can be derived from a basic motion vector candidate list, specifically... Figure 7 This demonstrates the fusion process for obtaining multiple hypothetical motion vectors, such as... Figure 7As shown, for motion vectors in the basic motion vector candidate list (such as...) Figure 7 The motion vectors (cand0, cand1, cand2, ...) are combined in pairs, and then concatenated or weighted to obtain the desired multi-hypothesis motion vector, which can be denoted as MHMV. Figure 7 MH0, MH1, MH2, etc.
[0074] According to an exemplary embodiment of this disclosure, before fusing any pairwise motion vectors in the motion vector candidate list of the current block to obtain multiple hypothetical motion vectors, the motion vectors in the motion vector candidate list can be corrected based on the template matching cost. This embodiment can improve the accuracy of pairwise motion vectors, ensuring that the pairwise motion vectors used to obtain multiple hypothetical motion vectors are optimal, thereby improving the accuracy of the multiple hypothetical motion vectors.
[0075] To better illustrate this embodiment, let's briefly introduce TM. TM is a technique that utilizes the correlation between spatially adjacent pixels to hide the coding syntax, thereby improving compression ratio. Taking Template-Matching based Merge mode as an example, firstly, a TM Merge candidate list is constructed, similar to the construction method of a regular Merge candidate list. Then, for each candidate MV in the TM Merge candidate list, the TM cost is used to correct it, resulting in a corrected MV. The corrected MV is then rate-distortion optimized (RDO) with other Merge modes. Whether a candidate MV is selected as the optimal MV can be recorded using a CU-level flag and written into the bitstream. After the decoder identifies the TM Merge mode, it performs the same MV candidate correction process as the encoder during the MV construction stage.
[0076] As an example, the correction process can be as follows: Figure 8 As shown, the template can be a reconstructed block with a left width of 4 and a top height of 4 on the current frame (cur pic) of the current block (cur block). Then, according to the corresponding MV, the corresponding position of the template of the current block on the reference frame (Ref0) can be obtained. Using this position as the center point of the search, the TM cost of each search position is calculated within a certain range. Finally, the MV with the smallest TM cost is selected as the corrected MV.
[0077] return Figure 5 In step S502, the optimal motion vector is selected from the multi-hypothesis motion vector candidate list and the motion vector candidate list.
[0078] As an example, an RDO can be performed on each multi-hypothesis motion vector in the multi-hypothesis motion vector candidate list and each motion vector in the motion vector candidate list to obtain the optimal motion vector.
[0079] According to an exemplary embodiment of this disclosure, selecting the optimal motion vector from a multi-hypothesis motion vector candidate list and a motion vector candidate list may include: for each multi-hypothesis motion vector in the multi-hypothesis motion vector candidate list, performing the following processing: when performing motion compensation on the current multi-hypothesis motion vector, performing motion compensation processing on the two motion vectors fused to obtain the current multi-hypothesis motion vector respectively, to obtain a compensation block corresponding to each motion vector; performing weighted processing on the compensation blocks corresponding to each motion vector to obtain a multi-hypothesis interpolation block corresponding to the current multi-hypothesis motion vector. Through this embodiment, by performing motion compensation on the two motion vectors corresponding to the multi-hypothesis motion vector respectively, and then performing weighted processing on the two compensation blocks obtained, it can be ensured that the obtained multi-hypothesis interpolation block is more accurate, thereby enabling more accurate encoding and improving the encoding compression ratio.
[0080] As an example, in the RDO process, for each multi-hypothesis motion vector, when motion compensation is required for the current multi-hypothesis motion vector, motion compensation can be performed on the two corresponding motion vectors of the current multi-hypothesis motion vector respectively, and then the two compensation blocks obtained by motion compensation are weighted to obtain the multi-hypothesis interpolation block corresponding to the current multi-hypothesis motion vector.
[0081] return Figure 5 In step S503, the current block is encoded based on the optimal motion vector.
[0082] According to an exemplary embodiment of this disclosure, after encoding the current block based on the optimal motion vector, a fusion mode identifier can be signaled in the bitstream. The fusion mode identifier is set to indicate that a multi-hypothesis fusion mode is enabled for the current block in response to the optimal motion vector being from a multi-hypothesis motion vector candidate list; conversely, the fusion mode identifier is set to indicate that the multi-hypothesis fusion mode is disabled for the current block in response to the optimal motion vector being from a motion vector candidate list. Through this embodiment, a fusion mode identifier is set for each current block, so that when encoding the current block subsequently, only the identifier at the CU level can be encoded, avoiding the addition of other syntax information and thus reducing the encoding compression ratio.
[0083] As an example, for each coding block, a fusion mode flag can be set to indicate whether multi-hypothesis fusion mode is enabled for the current block. For instance, if the optimal motion vector comes from a multi-hypothesis motion vector candidate list, the fusion mode flag can be set to 1 to indicate that multi-hypothesis fusion mode is enabled for the current block; if the optimal motion vector comes from a basic motion vector candidate list, the fusion mode flag can be set to 0 to indicate that multi-hypothesis fusion mode is disabled for the current block. Subsequent encoding of the current block only requires encoding the CU-level flag, without adding any other syntax information.
[0084] Figure 9 This is a flowchart illustrating a video decoding method according to exemplary embodiments of the present disclosure, such as... Figure 9 As shown, the video decoding method includes the following steps: In step S901, in response to determining that the multi-hypothesis fusion mode is enabled for the current block, any two motion vectors in the motion vector candidate list of the current block are fused to obtain multi-hypothesis motion vectors, and a multi-hypothesis motion vector candidate list is constructed.
[0085] Specifically, before fusion, the motion vector candidate list for the current block can be constructed in the order of spatially adjacent candidates, temporal candidates, spatially non-adjacent candidates, historical candidates, and chained candidates; or the motion vector candidate list for the current block can be constructed in the order of spatially adjacent candidates, temporal candidates, and historical candidates. This disclosure does not limit the scope of the fusion.
[0086] According to an exemplary embodiment of this disclosure, enabling a multi-hypothesis fusion mode for the current block can be determined by: obtaining a fusion mode identifier from the bitstream; and determining that a multi-hypothesis fusion mode is enabled for the current block in response to the fusion mode identifier indicating that a multi-hypothesis fusion mode is enabled for the current block. Through this embodiment, a fusion mode identifier is set for each current block, so that when encoding the current block, only the identifier at the CU level can be encoded, avoiding the addition of other syntax information and thus reducing the encoding compression ratio.
[0087] As an example, after the decoder receives the bitstream sent by the encoder, it can parse the bitstream. If the parsed prediction mode of the current block is Merge mode, the Merge Flag and corresponding index of the current block can be further parsed. If the Merge Flag is 1, it means that the multi-hypothesis fusion mode is enabled for the current block. In this case, any two motion vectors in the motion vector candidate list of the current block can be fused to obtain multi-hypothesis motion vectors, and a multi-hypothesis motion vector candidate list can be constructed. The specific process is the same as that of the encoder, and will not be discussed further here. If the Merge Flag is 0, it means that the multi-hypothesis fusion mode is disabled for the current block. In this case, the multi-hypothesis fusion mode can be skipped, and other fusion modes can be parsed.
[0088] According to an exemplary embodiment of this disclosure, before fusing any pairwise motion vectors in the motion vector candidate list of the current block to obtain multiple hypothetical motion vectors, the size of the multiple hypothetical motion vector candidate list can be determined; in response to the size of the multiple hypothetical motion vector candidate list being greater than the value of the fusion mode index corresponding to the optimal motion vector, the fusion operation is stopped. In this embodiment, considering that the decoding end only needs the multiple hypothetical motion vector corresponding to the fusion mode index, when constructing the multiple hypothetical motion vector candidate list, if it is detected that the multiple hypothetical motion vector corresponding to the fusion mode index has already been added to the list, subsequent multiple hypothetical motion vectors do not need to be acquired, thereby reducing the computational load.
[0089] As an example, since the decoding end only needs the multi-hypothesis motion vectors corresponding to the fusion mode index, when constructing the multi-hypothesis motion vector candidate list, if the multi-hypothesis motion vector corresponding to the fusion mode index has already been added to the multi-hypothesis motion vector candidate list, that is, if the size of the multi-hypothesis motion vector candidate list is greater than the value of the fusion mode index corresponding to the optimal motion vector, then subsequent multi-hypothesis motion vectors do not need to be acquired. In other words, the pairwise fusion of the remaining motion vectors in the motion vector candidate list can be stopped, thereby reducing the amount of computation.
[0090] In step S902, the optimal motion vector is selected from the list of candidate motion vectors based on multiple hypotheses.
[0091] Specifically, the optimal motion vector can be determined from a list of candidate motion vectors based on the merge index parsed from the bitstream and used for subsequent decoding.
[0092] In step S903, the current block is decoded based on the optimal motion vector.
[0093] As an example, after obtaining the optimal motion vector, motion compensation can be performed based on the optimal motion vector to obtain the predicted block for the current block. For instance, two motion vectors corresponding to the optimal motion vector can be determined, and motion compensation can be performed based on these two motion vectors respectively; this disclosure does not limit the scope of the invention.
[0094] According to an exemplary embodiment of this disclosure, decoding the current block based on the optimal motion vector may include: performing motion compensation processing on each pair of motion vectors corresponding to the optimal motion vector to obtain a compensation block (i.e., a prediction block for the current block) corresponding to each motion vector; performing weighted processing on the compensation blocks corresponding to each motion vector to obtain a multi-hypothesis interpolation block; and decoding the current block based on the multi-hypothesis interpolation block. Through this embodiment, by performing motion compensation on each of the two motion vectors corresponding to the multi-hypothesis motion vector, and then performing weighted processing on the two obtained compensation blocks, it is possible to ensure that the obtained multi-hypothesis interpolation block is more accurate, thereby enabling more accurate decoding.
[0095] In summary, to address the limitation that existing VVC Merge techniques can only support bidirectional prediction, this disclosure proposes a template-matching-based multi-hypothesis Merge technique to enhance bidirectional prediction in VVC Merge mode. This technique adds additional multi-hypothesis candidates (i.e., newly added multi-hypothesis motion vectors). These new candidates have different reference frames than existing candidates, allowing for further improvement in prediction accuracy through temporal pixel correlation. Furthermore, this disclosure employs a template-matching-based multi-hypothesis MV candidate construction method. Considering that directly adding multiple hypothesis candidates inevitably leads to syntax overhead, this method uses template matching to derive multi-hypothesis MV candidates, significantly increasing the probability of selection and effectively improving video compression rate.
[0096] To verify the feasibility of this disclosure, this disclosure was tested on 100 sequences from a commonly used sequence set under standard video codec and random access configuration with CRF values of 24, 26, 28, and 30. The test results are as follows: BD-Rate SSIM611 is -0.80%, BD-Rate PSNR611 is -0.82%, BD-Rate VMAF611 is -0.85%, and the encoding time increased by 4%.
[0097] Figure 10 This is a block diagram illustrating a video encoding apparatus according to exemplary embodiments of the present disclosure. (Refer to...) Figure 10 The device includes a first acquisition unit 1000, a first selection unit 1002, and an encoding unit 1004.
[0098] The first acquisition unit 1000 is configured to fuse any two motion vectors in the motion vector candidate list of the current block to obtain multiple hypothetical motion vectors and construct a multiple hypothetical motion vector candidate list; the first selection unit 1002 is configured to select the optimal motion vector from the multiple hypothetical motion vector candidate list and the motion vector candidate list; the encoding unit 1004 is configured to encode the current block based on the optimal motion vector.
[0099] According to an exemplary embodiment of the present disclosure, the encoding unit 1004 is further configured to, after encoding the current block based on the optimal motion vector, send a fusion mode identifier in the bit stream by signaling, wherein, in response to the optimal motion vector being from a multi-hypothesis motion vector candidate list, the fusion mode identifier is set to indicate that a multi-hypothesis fusion mode is enabled for the current block; and in response to the optimal motion vector being from a motion vector candidate list, the fusion mode identifier is set to indicate that a multi-hypothesis fusion mode is disabled for the current block.
[0100] According to an exemplary embodiment of this disclosure, the first acquisition unit 1000 is further configured to perform the following processing for each multi-hypothesis motion vector obtained by fusion: determining a first template matching cost and two second template matching costs, wherein the first template matching cost is a template matching cost calculated for the current multi-hypothesis motion vector, and the two second template matching costs are template matching costs calculated for the two motion vectors used to fuse the current multi-hypothesis motion vector; and adding the current multi-hypothesis motion vector to the multi-hypothesis motion vector candidate list in response to the first template matching cost being less than each of the second template matching costs.
[0101] According to an exemplary embodiment of this disclosure, the first acquisition unit 1000 is further configured to, for each of the two motion vectors, determine a prediction template corresponding to the current motion vector based on the position of the template of the current block and the current motion vector, wherein the position of the prediction template is the position of the template of the current block on the corresponding reference frame under the direction of the current motion vector; perform weighted processing on the pixels of the prediction template corresponding to each motion vector to obtain weighted pixels; and determine the rate-distortion cost of the weighted pixels and the pixels of the template of the current block as the first template matching cost.
[0102] According to an exemplary embodiment of this disclosure, the first acquisition unit 1000 is further configured to perform average weighting processing on the pixels of the prediction template corresponding to each motion vector to obtain weighted pixels; or, based on the template matching cost of each motion vector, perform weighting processing on the pixels of the prediction template corresponding to each motion vector to obtain weighted pixels.
[0103] According to an exemplary embodiment of the present disclosure, the first acquisition unit 1000 is further configured to, before determining the first template matching cost and the two second template matching costs, skip the current multi-hypothesis motion vector and acquire the next multi-hypothesis motion vector in response to the fact that the reference frames corresponding to the two motion vectors are the same; or, skip the current multi-hypothesis motion vector and acquire the next multi-hypothesis motion vector in response to the fact that the reference frames corresponding to the two motion vectors are the same and the difference information between the two motion vectors is not greater than a preset threshold.
[0104] According to an exemplary embodiment of the present disclosure, the first selection unit 1002 is further configured to perform the following processing for each multi-hypothesis motion vector in the multi-hypothesis motion vector candidate list: when performing motion compensation on the current multi-hypothesis motion vector, perform motion compensation processing on the two motion vectors that are fused to obtain the current multi-hypothesis motion vector respectively to obtain a compensation block corresponding to each motion vector; and perform weighted processing on the compensation block corresponding to each motion vector to obtain a multi-hypothesis interpolation block corresponding to the current multi-hypothesis motion vector.
[0105] According to an exemplary embodiment of this disclosure, the first acquisition unit 1000 is further configured to concatenate any two motion vectors in the motion vector candidate list of the current block to obtain multiple hypothetical motion vectors; or, to weight any two motion vectors in the motion vector candidate list of the current block to obtain multiple hypothetical motion vectors.
[0106] According to an exemplary embodiment of this disclosure, the first acquisition unit 1000 is further configured to modify the motion vectors in the motion vector candidate list based on template matching cost before fusing any pair of motion vectors in the motion vector candidate list of the current block to obtain multiple hypothetical motion vectors.
[0107] Figure 11 This is a block diagram illustrating a video decoding apparatus according to exemplary embodiments of the present disclosure. (Refer to...) Figure 11 The device includes a second acquisition unit 1100, a second selection unit 1102, and a decoding unit 1104.
[0108] The second acquisition unit 1100 is configured to, in response to determining that a multi-hypothesis fusion mode is enabled for the current block, fuse any two motion vectors in the motion vector candidate list of the current block to obtain a multi-hypothesis motion vector and construct a multi-hypothesis motion vector candidate list; the second selection unit 1102 is configured to select the optimal motion vector from the multi-hypothesis motion vector candidate list; and the decoding unit 1104 is configured to decode the current block based on the optimal motion vector.
[0109] According to an exemplary embodiment of this disclosure, the second acquisition unit 1100 is further configured to determine the size of the multi-hypothesis motion vector candidate list before fusing any pairwise motion vectors in the motion vector candidate list of the current block to obtain a multi-hypothesis motion vector; and to stop the fusion operation in response to the size of the multi-hypothesis motion vector candidate list being greater than the value of the fusion mode index corresponding to the optimal motion vector.
[0110] According to an exemplary embodiment of this disclosure, the multi-hypothesis fusion mode for the current block is determined by: obtaining a fusion mode identifier from the bitstream; and determining that the multi-hypothesis fusion mode is enabled for the current block in response to the fusion mode identifier indicating that the multi-hypothesis fusion mode is enabled for the current block.
[0111] According to an exemplary embodiment of this disclosure, the decoding unit 1104 is further configured to perform motion compensation processing on each pair of motion vectors corresponding to the optimal motion vector to obtain a compensation block corresponding to each motion vector; perform weighted processing on the compensation block corresponding to each motion vector to obtain a multi-hypothesis interpolation block; and decode the current block based on the multi-hypothesis interpolation block.
[0112] Figure 12 A computing environment 1210 coupled to a user interface 1250 is shown. The computing environment 1210 may be part of a data processing server. The computing environment 1210 includes a processor 1220, memory 1230, and input / output (I / O) interface 1240.
[0113] Processor 1220 typically controls the overall operation of computing environment 1210, such as operations associated with display, data acquisition, data communication, and image processing. Processor 1220 may include one or more processors for executing instructions to perform all or some of the steps in the methods described above. Furthermore, processor 1220 may include one or more modules that facilitate interaction between processor 1220 and other components. The processor may be a central processing unit (CPU), microprocessor, microcontroller, graphics processing unit (GPU), etc.
[0114] Memory 1230 is configured to store various types of data to support the operation of computing environment 1210. Memory 1230 may include predefined software 1232. Examples of such data include instructions for any application or method operating on computing environment 1210, video datasets, image data, etc. Memory 1230 can be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0115] I / O interface 1240 provides an interface between processor 1220 and peripheral interface modules (such as keyboard, click wheel, buttons, etc.). Buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. I / O interface 1240 can be coupled to encoders and decoders.
[0116] In an embodiment, a non-transitory computer-readable storage medium is also provided, including, for example, a plurality of programs in a memory 1230 and / or a bitstream generated by the video encoding method described above or a bitstream to be decoded by the video decoding method described above. The plurality of programs can be executed by a processor 1220 in a computing environment 1210 to perform the methods described above. In one example, the plurality of programs can be executed by a processor 1220 in a computing environment 1210 to (e.g., from...) Figure 2 The video encoder 20 in the computing environment 1210 receives a bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.), and can also be executed by the processor 1220 in the computing environment 1210 to perform the above-described decoding method based on the received bitstream or data stream. In another example, the plurality of programs can be executed by the processor 1220 in the computing environment 1210 to perform the above-described encoding method to encode video information (e.g., video blocks representing video frames, and / or one or more associated syntax elements, etc.) into a bitstream or data stream, and can also be executed by the processor 1220 in the computing environment 1210 to (e.g., to...) Figure 3 The video decoder 30 in the middle sends the bitstream or data stream. Alternatively, a non-transitory computer-readable storage medium may store data generated by the encoder (e.g., Figure 2 The video encoder 20 in the video encoder uses, for example, the encoding method described above to generate the video for the decoder (e.g., Figure 3 The video decoder 30 in the video decoder uses a bitstream or data stream that includes encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.) when decoding video data. Non-transitory computer-readable storage media can be, for example, ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage devices, etc.
[0117] In an embodiment, the computing environment 1210 may be implemented by one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components for performing the methods described above.
[0118] According to embodiments of the present disclosure, an electronic device may be provided, the electronic device including at least one memory and at least one processor, wherein the at least one memory stores a set of computer-executable instructions, and when the set of computer-executable instructions is executed by the at least one processor, a video encoding method and / or a video decoding method according to embodiments of the present disclosure are performed.
[0119] As an example, an electronic device can be a PC, tablet, personal digital assistant, smartphone, or other device capable of executing the aforementioned set of instructions. Here, the electronic device is not necessarily a single device; it can be any collection of devices or circuits capable of executing the aforementioned instructions (or instruction sets) individually or in combination. The electronic device can also be part of an integrated control system or system manager, or can be configured to interconnect locally or remotely (e.g., via wireless transmission) through an interface.
[0120] In addition, electronic devices may include video displays (such as liquid crystal displays) and user interaction interfaces (such as keyboards, mice, touch input devices, etc.). All components of the electronic device may be interconnected via buses and / or networks.
[0121] According to embodiments of this disclosure, a computer program product having instructions for storing a bitstream, the bitstream including video data generated by the video encoding method described above and / or video data decoded by the video decoding method described above, is also provided. In embodiments, a computer program product including, for example, a plurality of programs in a memory 1230 is also provided, the plurality of programs being executable by a processor 1220 in a computing environment 1210 to perform the methods described above. For example, the computer program product may include a non-transitory computer-readable storage medium.
[0122] In one embodiment, a method for generating a bitstream is provided, the method comprising generating a bitstream according to the video encoding method described above; and storing the bitstream.
[0123] In one embodiment, a method for generating a bitstream is provided, the method comprising: generating a bitstream according to the video encoding method described above; and sending the bitstream.
[0124] Unless otherwise specifically stated, the order of steps in the method according to this disclosure is intended to be illustrative only, and the steps of the method according to this disclosure are not limited to the specific order described above, but may be changed according to actual circumstances. Furthermore, at least one step in the method according to this disclosure may be adjusted, combined, or omitted as needed.
[0125] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.
[0126] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A video encoding method, characterized in that, Applied to the encoding end, including: Merge any pair of motion vectors in the candidate list of motion vectors for the current block to obtain multiple hypothesis motion vectors, and construct a candidate list of multiple hypothesis motion vectors; Select the optimal motion vector from the multi-hypothesis motion vector candidate list and the motion vector candidate list; The current block is encoded based on the optimal motion vector.
2. The method as described in claim 1, characterized in that, After encoding the current block based on the optimal motion vector, the method further includes: The fusion mode identifier is sent via signal in the bitstream. Wherein, in response to the optimal motion vector being derived from the multi-hypothesis motion vector candidate list, the fusion mode identifier is set to indicate that the multi-hypothesis fusion mode is enabled for the current block; In response to the optimal motion vector being selected from the motion vector candidate list, the fusion mode identifier is set to indicate that the multi-hypothesis fusion mode is turned off for the current block.
3. The method as described in claim 1, characterized in that, The construction of the multi-hypothesis motion vector candidate list includes: For each multi-hypothesis motion vector obtained by fusion, the following processing is performed: Determine a first template matching cost and two second template matching costs, wherein the first template matching cost is the template matching cost calculated for the current multi-hypothesis motion vector, and the two second template matching costs are the template matching costs calculated for each of the two motion vectors used to fuse the current multi-hypothesis motion vector; In response to the first template matching cost being less than the cost of each second template matching, the current multi-hypothesis motion vector is added to the multi-hypothesis motion vector candidate list.
4. The method as described in claim 3, characterized in that, The determination of the first template matching cost includes: For each of the two motion vectors, based on the position of the template of the current block and the current motion vector, the prediction template corresponding to the current motion vector is determined, wherein the position of the prediction template is the position of the template of the current block on the corresponding reference frame under the direction of the current motion vector; The pixels of the prediction template corresponding to each motion vector are weighted to obtain weighted pixels; The rate-distortion cost of the weighted processed pixel and the pixel of the template of the current block is determined as the first template matching cost.
5. The method as described in claim 4, characterized in that, The step of weighting the pixels of the prediction template corresponding to each motion vector to obtain weighted pixels includes: The pixels of the prediction template corresponding to each motion vector are averaged and weighted to obtain the weighted pixels; Alternatively, based on the template matching cost of each motion vector, the pixels of the predicted template corresponding to each motion vector are weighted to obtain weighted pixels.
6. The method as described in claim 3, characterized in that, Before determining the first template matching cost and the two second template matching costs, the following is also included: In response to the fact that the reference frames corresponding to the two motion vectors are the same, the current multi-hypothesis motion vector is skipped and the next multi-hypothesis motion vector is obtained; Alternatively, in response to the fact that the reference frames corresponding to the two motion vectors are the same and the difference information between the two motion vectors is not greater than a preset threshold, the current multi-hypothesis motion vector is skipped and the next multi-hypothesis motion vector is obtained.
7. The method as described in claim 1, characterized in that, The step of selecting the optimal motion vector from the multi-hypothesis motion vector candidate list and the motion vector candidate list includes: For each multi-hypothesis motion vector in the candidate list of multi-hypothesis motion vectors, the following processing is performed: When performing motion compensation on the current multi-hypothesis motion vector, the two motion vectors that are fused to obtain the current multi-hypothesis motion vector are respectively subjected to motion compensation processing to obtain the compensation block corresponding to each motion vector; The compensation block corresponding to each motion vector is weighted to obtain the multi-hypothesis interpolation block corresponding to the current multi-hypothesis motion vector.
8. The method as described in claim 1, characterized in that, The process of fusing any pair of motion vectors from the candidate list of motion vectors for the current block to obtain multiple hypothetical motion vectors includes: Concatenate any two motion vectors from the candidate list of motion vectors for the current block to obtain multiple hypothetical motion vectors; Alternatively, weighting can be performed on any pair of motion vectors in the candidate list of motion vectors for the current block to obtain multiple hypothetical motion vectors.
9. A video decoding method, characterized in that, Applied to the decoding end, including: In response to determining that the multi-hypothesis fusion mode is enabled for the current block, any two motion vectors in the motion vector candidate list of the current block are fused to obtain multi-hypothesis motion vectors, and a multi-hypothesis motion vector candidate list is constructed. Select the optimal motion vector from the list of multiple hypothetical motion vector candidates; The current block is decoded based on the optimal motion vector.
10. The method as described in claim 9, characterized in that, Before fusing any pairwise motion vectors from the candidate list of motion vectors for the current block to obtain multiple hypothetical motion vectors, the following steps are also included: Determine the size of the candidate list of multiple hypothesis motion vectors; The fusion operation is stopped in response to the fact that the size of the candidate list of multiple hypothetical motion vectors is greater than the value of the fusion mode index corresponding to the optimal motion vector.
11. The method as described in claim 9, characterized in that, The multi-hypothesis fusion mode for the current block is determined in the following way: Obtain the fusion mode identifier from the bitstream; In response to the fusion mode identifier indicating that a multi-hypothesis fusion mode is enabled for the current block, it is determined that a multi-hypothesis fusion mode is enabled for the current block.
12. A video encoding device, characterized in that, include: The first acquisition unit is configured to fuse any two motion vectors in the motion vector candidate list of the current block to obtain multiple hypothetical motion vectors and construct a multiple hypothetical motion vector candidate list; The first selection unit is configured to select the optimal motion vector from the multi-hypothesis motion vector candidate list and the motion vector candidate list; The encoding unit is configured to encode the current block based on the optimal motion vector.
13. A video decoding device, characterized in that, include: The second acquisition unit is configured to, in response to determining that a multi-hypothesis fusion mode is enabled for the current block, fuse any two pairs of motion vectors in the motion vector candidate list of the current block to obtain multi-hypothesis motion vectors, and construct a multi-hypothesis motion vector candidate list; The second selection unit is configured to select the optimal motion vector from the multi-hypothetical motion vector candidate list; The decoding unit is configured to decode the current block based on the optimal motion vector.
14. An electronic device, characterized in that, include: At least one processor; At least one memory that stores computer-executable instructions. Wherein, when the computer-executable instructions are executed by the at least one processor, they cause the at least one processor to perform the video encoding method as described in any one of claims 1 to 8 and / or the video decoding method as described in any one of claims 9 to 11.
15. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by at least one processor, the at least one processor causes the processor to execute a bitstream generated by the video encoding method of any one of claims 1 to 8 and / or a bitstream decoded by the video decoding method of any one of claims 9 to 11.