Video encoding and decoding method and device, electronic equipment, storage medium, bit stream storage method and program product

By dynamically adjusting the encoding order of MMVD syntax elements in CTU and optimizing the offset step range, the high bit rate consumption problem of the MMVD algorithm is solved, improving encoding efficiency and quality.

CN121967678APending Publication Date: 2026-05-01BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
Filing Date
2026-01-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Among existing video encoding and decoding technologies, the MMVD algorithm has high bitrate consumption and high encoding complexity, which affects encoding efficiency.

Method used

By dynamically adjusting the encoding order of MMVD syntax elements of CTU in the current frame, utilizing the CTU selection status of the reference frame adjacent to the current frame, optimizing the offset step range of MMVD, adjusting the merged prediction mode syntax tree structure, and pre-encoding the MMVD enable flag, unnecessary encoding overhead is reduced.

Benefits of technology

It effectively reduces the bitrate consumption of MMVD, improves encoding efficiency and quality, and optimizes the encoding performance of MMVD.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967678A_ABST
    Figure CN121967678A_ABST
Patent Text Reader

Abstract

The invention provides a video coding and decoding method and device, electronic equipment, a storage medium, a bit stream storage method and a program product, and the video coding method comprises the steps: obtaining a first statistical value for a current coding tree unit (CTU) in a current frame, the first statistical value is a statistical value of a block of an MMVD mode in a selected merge prediction mode in an associated CTU related to the current CTU, and the associated CTU comprises a co-located CTU in a reference frame of the current frame and an adjacent CTU of the co-located CTU, wherein the co-located CTU and the current CTU are co-located; in response to the first statistical value meeting a preset threshold condition, adjusting the merge prediction mode syntax tree structure to firstly encode an MMVD enabling flag, the preset threshold condition being that the first statistical value exceeds a threshold related to the preset threshold condition; and in response to the fact that the current block in the current CTU is in the merge prediction mode, coding the current block based on the adjusted merge prediction mode syntax tree structure.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding and decoding methods and apparatuses, electronic devices, storage media, methods for storing bit streams, and software products Technical Field

[0001] This disclosure relates to the field of video encoding / decoding and compression technology. More specifically, this disclosure relates to video encoding / decoding methods and apparatus, electronic devices, storage media, methods for storing bitstreams, and program products. Background Technology

[0002] Various electronic devices support digital video. These devices transmit and receive, or otherwise transfer, digital video data via communication networks, and / or store digital video data on storage devices. Because communication networks have limited bandwidth capacity and storage devices have limited storage resources, video data can be compressed using one or more video codec standards before transmission or storage to generate coded video data using a lower bit rate, while avoiding or minimizing video quality degradation. Summary of the Invention

[0003] Examples of this disclosure provide video encoding / decoding methods and apparatus, electronic devices, storage media, methods for storing bit streams, and program products.

[0004] According to a first aspect of this disclosure, a video coding method is provided, comprising: obtaining a first statistical value for a current coding tree unit (CTU) in a current frame, wherein the first statistical value is a statistical value of a block in MMVD mode of a selected merge prediction mode in an associated CTU related to the current CTU, the associated CTU including a co-located CTU in a reference frame of the current frame and adjacent CTUs of the co-located CTU; in response to the first statistical value satisfying a preset threshold condition, adjusting the merge prediction mode syntax tree structure to first encode an MMVD enable flag, the preset threshold condition being that the first statistical value exceeds a threshold related to the preset threshold condition; and in response to the current block in the current CTU being in merge prediction mode, encoding the current block based on the adjusted merge prediction mode syntax tree structure.

[0005] According to an exemplary embodiment of this disclosure, the threshold related to the preset threshold condition is determined based on a second statistical value for the current CTU, wherein the second statistical value is the statistical value of the block in the associated CTU that selects other prediction modes besides the MMVD mode in the merged prediction mode.

[0006] According to an exemplary embodiment of this disclosure, the preset threshold condition includes at least one of the following conditions: the area of ​​the block with the selected MMVD mode in the associated CTU is greater than a predetermined multiple of the largest area of ​​the block with the selected prediction mode in each of the other prediction modes in the associated CTU; the number of blocks with the selected MMVD mode in the associated CTU is greater than a predetermined multiple of the maximum number of blocks with the selected prediction mode in each of the other prediction modes in the associated CTU.

[0007] According to an exemplary embodiment of this disclosure, the reference frame is the nearest neighbor reference frame of the current frame.

[0008] According to an exemplary embodiment of this disclosure, the method further includes: selecting a target offset step set from a plurality of preset offset step set for the current CTU, as an offset step set for the current CTU; wherein each of the plurality of preset offset step set is a subset of the offset step set set set for the MMVD mode, and each preset offset step set is a different subset of each other.

[0009] According to an exemplary embodiment of this disclosure, the plurality of preset offset step size sets include a first preset offset step size set and a second preset offset step size set, wherein the second preset offset step size set includes a larger range of offset step size candidates than the first preset offset step size set.

[0010] According to an exemplary embodiment of this disclosure, the first preset offset step size set is {1 / 4, 1 / 2, 1, 2, 4}, and the second preset offset step size set is {1 / 4, 1 / 2, 1, 2, 4, 8, 16}.

[0011] According to an exemplary embodiment of this disclosure, selecting a target offset step set from a plurality of preset offset step set includes: selecting a second preset offset step set as the target offset step set in response to the current CTU satisfying a preset offset step condition; and selecting a first preset offset step set as the target offset step set in response to the current CTU not satisfying the preset offset step condition; wherein the preset offset step condition includes: the average offset step size of the blocks of the selected MMVD mode in the associated CTU is greater than a preset threshold.

[0012] According to a second aspect of this disclosure, a video decoding method is provided, comprising: in response to a current block in a current coding tree unit (CTU) of a current frame being in merge prediction mode, decoding the current block based on a merge prediction mode syntax tree structure; wherein, when a first statistical value of the current CTU satisfies a preset threshold condition, the merge prediction mode syntax tree structure is adjusted to first decode an MMVD enable flag, the preset threshold condition being that the first statistical value exceeds a threshold related to the preset threshold condition; wherein the first statistical value is a statistical value of a block in the merge prediction mode selected in the MMVD mode among associated CTUs associated with the current CTU, the associated CTUs including co-located CTUs in a reference frame of the current frame and adjacent CTUs of the co-located CTUs.

[0013] According to an exemplary embodiment of this disclosure, the threshold related to the preset threshold condition is determined based on a second statistical value for the current CTU, wherein the second statistical value is the statistical value of the block in the associated CTU that selects other prediction modes besides the MMVD mode in the merged prediction mode.

[0014] According to an exemplary embodiment of this disclosure, the preset threshold condition includes at least one of the following conditions: the area of ​​the block with the selected MMVD mode in the associated CTU is greater than a predetermined multiple of the largest area of ​​the block with the selected prediction mode in each of the other prediction modes in the associated CTU; the number of blocks with the selected MMVD mode in the associated CTU is greater than a predetermined multiple of the maximum number of blocks with the selected prediction mode in each of the other prediction modes in the associated CTU.

[0015] According to an exemplary embodiment of this disclosure, the reference frame is the nearest neighbor reference frame of the current frame.

[0016] According to an exemplary embodiment of this disclosure, the method further includes: selecting a target offset step set from a plurality of preset offset step set for the current CTU, as an offset step set for the current CTU; wherein each of the plurality of preset offset step set is a subset of the offset step set set set for the MMVD mode, and each preset offset step set is a different subset of each other.

[0017] According to an exemplary embodiment of this disclosure, the plurality of preset offset step size sets include a first preset offset step size set and a second preset offset step size set, wherein the second preset offset step size set includes a larger range of offset step size candidates than the first preset offset step size set.

[0018] According to an exemplary embodiment of this disclosure, the first preset offset step size set is {1 / 4, 1 / 2, 1, 2, 4}, and the second preset offset step size set is {1 / 4, 1 / 2, 1, 2, 4, 8, 16}.

[0019] According to an exemplary embodiment of this disclosure, the method further includes: receiving an index indicating the target offset step set at the CTU level; wherein selecting the target offset step set from a plurality of preset offset step sets includes: selecting the target offset step set from the plurality of preset offset step sets according to the received index.

[0020] According to a third aspect of this disclosure, a video encoding apparatus is provided, comprising: a statistics module configured to: acquire a first statistical value for a current coding tree unit (CTU) in a current frame, wherein the first statistical value is a statistical value of a block in the selected merge prediction mode (MMVD) mode among associated CTUs related to the current CTU, the associated CTUs including co-located CTUs in a reference frame of the current frame and adjacent CTUs of the co-located CTUs; an adjustment module configured to: adjust the merge prediction mode syntax tree structure to first encode an MMVD enable flag in response to the first statistical value satisfying a preset threshold condition, wherein the preset threshold condition is that the first statistical value exceeds a threshold related to the preset threshold condition; and an encoding module configured to: encode the current block based on the adjusted merge prediction mode syntax tree structure in response to the current block in the current CTU being in merge prediction mode.

[0021] According to an exemplary embodiment of this disclosure, the threshold related to the preset threshold condition is determined based on a second statistical value for the current CTU, wherein the second statistical value is the statistical value of the block in the associated CTU that selects other prediction modes besides the MMVD mode in the merged prediction mode.

[0022] According to an exemplary embodiment of this disclosure, the preset threshold condition includes at least one of the following conditions: the area of ​​the block with the selected MMVD mode in the associated CTU is greater than a predetermined multiple of the largest area of ​​the block with the selected prediction mode in each of the other prediction modes in the associated CTU; the number of blocks with the selected MMVD mode in the associated CTU is greater than a predetermined multiple of the maximum number of blocks with the selected prediction mode in each of the other prediction modes in the associated CTU.

[0023] According to an exemplary embodiment of this disclosure, the reference frame is the nearest neighbor reference frame of the current frame.

[0024] According to an exemplary embodiment of the present disclosure, the video encoding apparatus further includes: an offset step set selection module, configured to select a target offset step set from a plurality of preset offset step sets for the current CTU, as an offset step set for the current CTU; wherein each of the plurality of preset offset step sets is a subset of the offset step set set set for the MMVD mode, and each preset offset step set is a different subset of each other.

[0025] According to an exemplary embodiment of this disclosure, the plurality of preset offset step size sets include a first preset offset step size set and a second preset offset step size set, wherein the second preset offset step size set includes a larger range of offset step size candidates than the first preset offset step size set.

[0026] According to an exemplary embodiment of this disclosure, the first preset offset step size set is {1 / 4, 1 / 2, 1, 2, 4}, and the second preset offset step size set is {1 / 4, 1 / 2, 1, 2, 4, 8, 16}.

[0027] According to an exemplary embodiment of this disclosure, the offset step size set selection module is configured to: select the second preset offset step size set as the target offset step size set in response to the current CTU satisfying the preset offset step size condition; and select the first preset offset step size set as the target offset step size set in response to the current CTU not satisfying the preset offset step size condition; wherein the preset offset step size condition includes: the average offset step size of the blocks of the selected MMVD mode in the associated CTU is greater than a preset threshold.

[0028] According to a fourth aspect of this disclosure, a video decoding apparatus is provided, comprising: a decoding module configured to decode the current block based on a merge prediction mode syntax tree structure in response to a current block in a current coding tree unit (CTU) of a current frame being in merge prediction mode; wherein, when a first statistical value of the current CTU satisfies a preset threshold condition, the merge prediction mode syntax tree structure is adjusted to first decode an MMVD enable flag, the preset threshold condition being a threshold related to the first statistical value; wherein the first statistical value is a statistical value of a block in the merge prediction mode selected in the MMVD mode among associated CTUs associated with the current CTU, the associated CTUs including co-located CTUs in a reference frame of the current frame and adjacent CTUs of the co-located CTUs.

[0029] According to an exemplary embodiment of this disclosure, the threshold related to the preset threshold condition is determined based on a second statistical value for the current CTU, wherein the second statistical value is the statistical value of the block in the associated CTU that selects other prediction modes besides the MMVD mode in the merged prediction mode.

[0030] According to an exemplary embodiment of this disclosure, the preset threshold condition includes at least one of the following conditions: the area of ​​the block with the selected MMVD mode in the associated CTU is greater than a predetermined multiple of the largest area of ​​the block with the selected prediction mode in each of the other prediction modes in the associated CTU; the number of blocks with the selected MMVD mode in the associated CTU is greater than a predetermined multiple of the maximum number of blocks with the selected prediction mode in each of the other prediction modes in the associated CTU.

[0031] According to an exemplary embodiment of this disclosure, the reference frame is the nearest neighbor reference frame of the current frame.

[0032] According to an exemplary embodiment of the present disclosure, the video decoding apparatus further includes: an offset step set selection module, configured to select a target offset step set from a plurality of preset offset step sets for the current CTU, as an offset step set for the current CTU; wherein each of the plurality of preset offset step sets is a subset of the offset step set set set for the MMVD mode, and each preset offset step set is a different subset of each other.

[0033] According to an exemplary embodiment of this disclosure, the plurality of preset offset step size sets include a first preset offset step size set and a second preset offset step size set, wherein the second preset offset step size set includes a larger range of offset step size candidates than the first preset offset step size set.

[0034] According to an exemplary embodiment of this disclosure, the first preset offset step size set is {1 / 4, 1 / 2, 1, 2, 4}, and the second preset offset step size set is {1 / 4, 1 / 2, 1, 2, 4, 8, 16}.

[0035] According to an exemplary embodiment of the present disclosure, the video decoding apparatus further includes: an index receiving module configured to receive an index indicating the target offset step set at the CTU level; the offset step set selection module is configured to select the target offset step set from the plurality of preset offset step sets according to the received index.

[0036] According to a fifth aspect of this disclosure, an electronic device includes: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement a video encoding method or a video decoding method according to this disclosure.

[0037] According to a sixth aspect of this disclosure, a non-transitory computer-readable storage medium is provided for storing computer-executable instructions, which, when executed by one or more computer processors, cause the one or more computer processors to perform a video encoding method or a video decoding method according to this disclosure.

[0038] According to a seventh aspect of this disclosure, a non-transitory computer-readable storage medium stores instructions and a bit stream, wherein, when executed by a computing device having one or more processors, the instructions cause the one or more processors to perform a video encoding method according to this disclosure to generate the bit stream.

[0039] According to an eighth aspect of this disclosure, a method for storing a bitstream includes: generating a bitstream by performing a video encoding method according to this disclosure; and storing the bitstream.

[0040] According to a ninth aspect of this disclosure, a computer program product includes a computer program that, when executed by a processor, implements a video encoding method according to this disclosure, or implements a video decoding method according to this disclosure.

[0041] The technical solutions provided by the embodiments of this disclosure bring at least the following beneficial effects: Since the selection of MMVD by CTU between adjacent video frames in the temporal domain is highly correlated, in this disclosure, the encoding order of the syntax elements of MMVD of CTU in the current frame can be dynamically adjusted based on the selection of MMVD by CTU in the reference frame adjacent to the current frame. That is, the encoding position of the MMVD syntax elements of CTU in the current frame in the merged prediction mode syntax tree structure can be dynamically adjusted, thereby effectively reducing the bit rate consumption of MMVD and improving the encoding efficiency of MMVD.

[0042] Furthermore, according to this disclosure, the selection status of MMVD by the CTU of the reference frame adjacent to the current frame can be used to dynamically optimize the offset step range of the MMVD corresponding to the CU contained in the current CTU to be encoded in the current frame, thereby improving the encoding performance of MMVD while ensuring encoding quality.

[0043] It will be understood that the above general description and the following detailed description are merely examples and do not limit this disclosure. Attached Figure Description

[0044] The accompanying drawings, which are incorporated in and form part of this specification, illustrate examples according to this disclosure and, together with this description, serve to explain the principles of this disclosure.

[0045] Figure 1 is a schematic diagram showing the merged prediction mode syntax tree structure of CU in related technologies.

[0046] Figure 2 is a schematic diagram showing the MMVD offset direction in the related technology.

[0047] Figure 3 is a flowchart illustrating a video encoding method according to an exemplary embodiment of the present disclosure.

[0048] Figure 4 is a schematic diagram illustrating an associated CTU related to the current CTU according to an exemplary embodiment of the present disclosure.

[0049] Figure 5 is a schematic diagram illustrating an optimized merged prediction pattern syntax tree structure according to an exemplary embodiment of the present disclosure.

[0050] Figure 6 is a flowchart illustrating a video decoding method according to an exemplary embodiment of the present disclosure.

[0051] Figure 7 is a block diagram illustrating a video encoding apparatus according to an exemplary embodiment of the present disclosure.

[0052] Figure 8 is a block diagram illustrating a video decoding apparatus according to an exemplary embodiment of the present disclosure.

[0053] Figure 9 is a schematic diagram illustrating a computing environment coupled with a user interface. Detailed Implementation

[0054] Referring now to the detailed description, examples of which are illustrated in the accompanying drawings. Numerous non-limiting details are set forth in the following detailed description to aid in understanding the subject matter presented herein. However, various alternatives may be used without departing from the scope of the claims, and the subject matter may be practiced without these specific details. For example, the subject matter presented herein can be implemented on many types of electronic devices with digital video capabilities.

[0055] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this disclosure are used to distinguish objects and are not used to describe any specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in sequences other than those shown in the drawings or described in this disclosure.

[0056] In this disclosure, the video encoder and video decoder can operate according to proprietary or industry standards (e.g., Universal Video Codec (VVC), Joint Explore Test Model (JEM), High Efficiency Video Codec (HEVC / H.265), Advanced Video Codec (AVC / H.264), Moving Picture Experts Group (MPEG) codec) or extensions of such standards (e.g., encoding and decoding video data). It should be understood that this application is not limited to any particular video coding / decoding standard and may be applicable to other current and future video coding / decoding standards.

[0057] The video encoder and video decoder can be implemented as any circuit of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic devices, software, hardware, firmware, or any combination thereof. When partially implemented in software, the electronic device may store instructions for the software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in the hardware to perform the video encoding / decoding operations disclosed in this disclosure. Each of the video encoder and video decoder may be included in one or more encoders or decoders, and either encoder or decoder may be integrated as part of a combined encoder / decoder (CODEC) in the respective device.

[0058] The VVC merge prediction mode includes five algorithms: Merge with Motion Vector Difference (MMVD), regular merge prediction, Combined Inter and Intra Prediction (CIIP), affine merge prediction, and Gemetric Partitioning Mode (GPM) prediction. For each coding unit (CU) in a video frame, if a merge prediction mode is selected, it means that the CU must have selected one of the above five algorithms. Figure 1 is a schematic diagram showing the syntax tree structure of the merge prediction mode for CUs in related technologies.

[0059] Referring to Figure 1, `merge_subblock_flag` corresponds to the affine merge prediction algorithm; `regular_merge_flag` corresponds to both the regular merge prediction algorithm and the MMVD prediction algorithm. These two algorithms share an enable flag, therefore, an MMVD enable switch needs to be encoded to determine which algorithm is used. `mmvd_enable_flag` and `ciip_flag` are enable flags for each algorithm. If the enable flag is 1, it indicates that the corresponding algorithm is selected; otherwise, it indicates that the corresponding algorithm is not selected. It should be noted that the GPM enable switch is not encoded at this point. This is because if `ciip_flag` is 0, only GPM remains out of the five algorithms, so there is no need to encode the GPM enable switch; GPM is selected by default. The leaf nodes in the syntax tree represent the syntax information that needs to be encoded for the corresponding algorithm.

[0060] Therefore, as shown in Figure 1, the execution logic of the entire syntax tree is as follows: First, `merge_subblock_flag` is encoded to determine whether the affine merge prediction algorithm is selected. If so, the syntax information of the specific prediction mode corresponding to the algorithm (represented here as Subblock Merge) can be encoded, and the process can end directly; otherwise, `regular_merge_flag` can be encoded, and it can be determined whether the regular merge prediction algorithm or the MMVD prediction algorithm is selected. If so, `mmvd_enable_flag` can be encoded, and it can be determined whether the MMVD prediction algorithm is selected. If so, the syntax elements of MMVD can be encoded, and the process ends; otherwise, the syntax elements of the regular merge prediction algorithm can be encoded, and the process ends. If `regular_merge_flag` is 0, `ciip_flag` can be encoded, and it can be determined whether the CIIP prediction algorithm is selected. If so, the syntax elements of CIIP can be encoded; otherwise, the syntax elements of GPM can be encoded, and the process finally ends.

[0061] Therefore, from the execution logic of the entire syntax tree, it can be seen that the affine merge prediction algorithm is closer to the root node. At this point, encoding the affine merge prediction algorithm requires the fewest enable flags, i.e., only `merge_subblock_flag` needs to be encoded. The remaining four algorithms, however, require encoding three enable flags before encoding their own motion information. For example, for the MMVD prediction algorithm, three enable flags—`merge_subblock_flag`, `regular_merge_flag`, and `mmvd_enable_flag`—must be encoded before encoding the motion information of the MMVD prediction algorithm itself. The other three algorithms are similar to the MMVD prediction algorithm. The more enable flags encoded, the higher the bitrate required. Therefore, encoding too many enable flags upfront further increases the bitrate overhead of MMVD.

[0062] In video coding, the MMVD algorithm is a prediction algorithm in the merge prediction mode of the VVC coding standard. This algorithm can effectively improve the accuracy of the motion vector (MV) of the CU, thereby improving the inter-frame prediction efficiency of the CU. MMVD constructs MMVD candidates using the first two candidate motion information from the merge candidate list, and performs offsets in both horizontal and vertical directions. There are a total of 8 candidate offset steps in the horizontal and vertical directions. Finally, through rate-distortion optimization, the optimal candidate motion information index, offset direction, and offset step size are selected from 64 candidates. These values ​​all need to be written into the bitstream.

[0063] Furthermore, MMVD can include unidirectional MMVD and bidirectional MMVD. Unidirectional MMVD refers to the prediction direction in the candidate motion information being unidirectional, meaning the current coding block is predicted unidirectionally. In this case, only the forward reference frame (L0 reference) or the backward reference frame (L1 reference) needs to be referenced. Bidirectional MMVD refers to the prediction direction in the candidate motion information being bidirectional, meaning the current coding block is predicted bidirectionally. In this case, both the forward reference frame and the backward reference frame need to be referenced simultaneously.

[0064] MMVD adjustments to the motion vector (MV) can include offset direction and offset step size. Here, we will illustrate this using unidirectional MMVD as an example. MMVD specifies that after extracting the initial candidate MV, the starting point should be the position pointed to by the extracted initial candidate MV in the reference frame. Different motion vectors can be formed based on four offset directions (e.g., positive and negative x-axis, positive and negative y-axis) and eight offset step sizes. Figure 2 is a schematic diagram illustrating the offset direction of MMVD in related technologies. Referring to Figure 2, the origin of the dashed line is the starting point of the initial candidate MV in the reference frame, and the initial candidate MV will be offset in the horizontal direction (positive or negative horizontal direction) or the vertical direction (positive or negative vertical direction) according to the offset step size, thereby obtaining a new MV.

[0065] As mentioned earlier, MMVD can specify 8 offset steps. For example, these 8 offset steps can be {1 / 4, 1 / 2, 1, 2, 4, 8, 16, 32}. Below, we will first use unidirectional MMVD as an example to illustrate the implementation process of adjusting the motion vector based on the offset steps.

[0066] Assuming the initial candidate MV is {4, 6} and the offset step size is 2, the offsets MVoffset in the four directions can be: positive x-axis {2, 0}; negative x-axis {-2, 0}; positive y-axis {0, 2}; negative y-axis {0, -2}. Furthermore, the adjusted motion vector can be expressed as: MVfinal = MV + MVoffset. For example, for the positive x-axis direction, the adjusted motion vector can be: MVfinal = {4, 6} + {2, 0} = {6, 6}.

[0067] Next, let's take bidirectional MMVD as an example to illustrate the implementation process of adjusting motion vectors based on offset step size. Since bidirectional MMVD has two MVs, two offsets, MV0 offset and MV1 offset, need to be derived simultaneously. Furthermore, the derivation process of MV0 offset is the same as that of unidirectional MMVD; while MV1 offset is determined by three factors: the Picture Order Count (POC) of the current frame, the POC of the preceding reference frame (POC0), and the POC of the following reference frame (POC1). Specifically: if (POC0 - POC) If (POC1 - POC) >= 0, then MV1 offset = MV0 offset; otherwise, if (POC0 - POC) >= 0, then MV1 offset = MV0 offset. If (POC1 – POC) < 0, then MV1 offset = MV0 offset (-1). For example, assuming the forward initial candidate motion vector MV0 is {4,6}, the backward initial candidate motion vector MV1 is {5,5}, POC=2, POC0=1, POC1=3, and the offset step size is 2, then the offset step sizes in the four directions can be as follows: positive x-axis direction: MV0offset={2,0}, MV1 offset={-2,0}; negative x-axis direction: MV0offset={-2,0}, MV1 offset={2,0}; positive y-axis direction: MV0offset={0,2}, MV1 offset={0,-2}; negative y-axis direction: MV0offset={0,-2}, MV1 offset={0,2}.

[0068] Taking the positive x-axis as an example, the adjusted motion vector can be expressed as: MV0 final={4, 6} + {2,0} = {6, 6}; MV1 final={5, 5} + {-2,0} = {3, 5}.

[0069] In addition, the encoding method of MMVD is as follows: each encoding block that selects MMVD needs to write three syntax elements into the bitstream. These three syntax elements are: encoding candidate list index (MMVD_cand_flag), offset step index (MMVD_step_idx), and offset direction index (MMVD_direction_idx). Among them, MMVD_cand_flag only needs 1 bit to represent its value, as shown in Table 1.

[0070] Table 1. Meaning of MMVD_cand_flag syntax elements. MMVD_step_idx uses truncated unary code encoding. Depending on the selected offset step size, it requires 1 to 7 bits to represent its value, as shown in Table 2.

[0071] Table 2 shows the meanings of the MMVD_step_idx syntax elements. If the 0th step size is selected, it can be represented by 1 bit "0"; if the 2nd step size is selected, it can be represented by 3 bits "110"; if the 5th step size is selected, it can be represented by 6 bits "111110", and so on.

[0072] MMVD_direction_idx requires 2 bits to represent its value, as shown in Table 3:

[0073] As described above, Table 3 defines the meaning of the MMVD_direction_idx syntax elements. MMVD needs to select the optimal candidate motion information index from 64 candidates (combined with motion information candidates, offset direction, and offset step size). These values, including offset direction and offset step size, must be written into the bitstream. Therefore, compared to other inter-frame prediction algorithms, MMVD has higher encoding complexity and consumes a larger bitrate due to the larger amount of motion information it needs to encode.

[0074] To address the aforementioned issues, this disclosure provides a video encoding / decoding method and apparatus, electronic device, storage medium, method for storing bitstreams, and program product. Since the selection of MMVD by the CTU between adjacent video frames in the temporal domain is highly correlated, this disclosure allows for the dynamic adjustment of the encoding order of the MMVD syntax elements of the CTU in the current frame based on the selection of MMVD by the CTU in a reference frame adjacent to the current frame. This means the encoding position of the MMVD syntax elements of the CTU in the current frame within the merged prediction mode syntax tree structure can be dynamically adjusted, thereby effectively reducing the bitrate consumption of MMVD and improving its encoding efficiency.

[0075] Furthermore, according to this disclosure, the selection status of MMVD by the CTU of the reference frame adjacent to the current frame can be used to dynamically optimize the offset step range of the MMVD corresponding to the CU contained in the current CTU to be encoded in the current frame, thereby improving the encoding performance of MMVD while ensuring encoding quality.

[0076] Figure 3 is a flowchart illustrating a video encoding method according to an exemplary embodiment of the present disclosure.

[0077] Referring to Figure 3, in step 301, a first statistical value for the current coding tree unit (CTU) in the current frame can be obtained. This first statistical value can be the statistical value of the block in the selected merge prediction mode of the associated CTUs related to the current CTU, specifically the MMVD mode. Here, the associated CTUs can include the co-located CTUs in the reference frame of the current frame, as well as the adjacent CTUs of the co-located CTUs. In the case of unidirectional prediction, the reference frame includes the forward reference frame; in the case of bidirectional prediction, the reference frame includes both the forward and backward reference frames. A co-located CTU can refer to a CTU in the reference frame whose spatial position in the reference frame is the same as that of a CTU in the current frame. For example, assuming the current CTU to be encoded is the first CTU located in the upper left corner of the current frame, then the co-located CTU in the reference frame also refers to the first CTU located in the upper left corner of the reference frame.

[0078] According to an exemplary embodiment of this disclosure, the aforementioned reference frame can be the nearest neighbor reference frame of the current frame. Here, the nearest neighbor reference frame can refer to the frame with the smallest difference between the corresponding POC and the POC of the current frame among all reference frames of the current frame. That is, in the case of one-way prediction, the reference frame can include the forward nearest neighbor reference frame of the current frame; in the case of two-way prediction, the reference frame can include the forward nearest neighbor reference frame and the backward nearest neighbor reference frame of the current frame.

[0079] As described above, the optimizations of this disclosure can be at the CTU level. Figure 4 is a schematic diagram illustrating associated CTUs related to the current CTU according to an exemplary embodiment of this disclosure. Referring to Figure 4, associated CTUs related to the current CTU may include co-located CTUs in a reference frame of the current frame and neighboring CTUs of the co-located CTUs. For the forward prediction direction, the co-located CTUs are... Furthermore, the forward co-position CTU There are a total of 4 adjacent CTUs, which are located in the forward co-position CTU. The top, bottom, left, and right positions: , , , Similarly, for the backward prediction direction, the co-position CTU is... Furthermore, the backward co-located CTU There are a total of 4 adjacent CTUs, which are located in the backward co-position CTU. The top, bottom, left, and right positions: , , , .

[0080] It should be noted that in the case of one-way prediction, only one nearest neighbor reference frame can be used. For example, only the forward nearest neighbor reference frame can be used, and correspondingly, only the co-located CTU and the four adjacent CTUs contained in the forward nearest neighbor reference frame will be used. In the case of two-way prediction, both the forward nearest neighbor reference frame and the backward nearest neighbor reference frame will be used. That is, the co-located CTU and the four adjacent CTUs contained in the forward nearest neighbor reference frame will be used simultaneously, as will the co-located CTU and the four adjacent CTUs contained in the backward nearest neighbor reference frame.

[0081] Furthermore, for non-boundary CTUs in the current frame, each corresponding CTU will have four adjacent CTUs; while for boundary CTUs in the current frame, each corresponding CTU will have only two or three adjacent CTUs. For example, for CTUs in the top, bottom, left, and right corners of the current frame, each corresponding CTU will have only two adjacent CTUs; while for boundary CTUs not in the four corners of the current frame, each corresponding CTU will have only three adjacent CTUs.

[0082] In step 302, in response to the first statistical value satisfying a preset threshold condition, the merged prediction mode syntax tree structure can be adjusted to first encode the MMVD enable flag, wherein the preset threshold condition can be that the first statistical value exceeds a threshold related to the preset threshold condition.

[0083] In step 303, in response to the current block in the current CTU being in merge prediction mode, the current block can be encoded based on the adjusted merge prediction mode syntax tree structure. That is, when there are many blocks in the associated CTU that select MMVD mode, or when the proportion of such blocks is relatively large, the probability of the current block in the current CTU selecting MMVD mode is also relatively high. Therefore, the encoding position of the MMVD enable flag can be moved forward to the root node of the merge prediction mode syntax tree structure to save the number of encoding syntax elements in the merge prediction mode syntax tree structure, thereby reducing the number of encoding bits and improving encoding efficiency.

[0084] According to an exemplary embodiment of this disclosure, the threshold related to the preset threshold condition can be determined based on a second statistical value for the current CTU, wherein the second statistical value can be the statistical value of the block of other prediction modes besides the MMVD mode in the selected merge prediction modes of the associated CTU.

[0085] For example, for forward co-location CTU CTU with forward co-position Four adjacent CTUs: , , , ; and, backward co-position CTU , and the backward co-located CTU Four adjacent CTUs: , , , With a total of 10 CTUs, we can determine the status of each of the five algorithms included in the merge prediction mode selected from these 10 CTUs: MMVD prediction algorithm, conventional merge prediction algorithm, CIIP prediction algorithm, affine merge prediction algorithm, and GPM prediction algorithm.

[0086] For example, we can determine the number of Custodians (CUs) selected from each of the five prediction algorithms mentioned above among these 10 CTUs, as well as the total area of ​​the CUs. Furthermore, the number of CUs can be denoted as: , , as well as The area of ​​CU can be denoted as: , , as well as .

[0087] According to an exemplary embodiment of this disclosure, the aforementioned preset threshold condition may include: the area of ​​the block with the selected MMVD mode in the associated CTU is greater than a predetermined multiple of the area of ​​the largest area among the blocks with selected other prediction modes in the associated CTU. (1) Where threshold2 is a positive real number greater than 1. For example, threshold2 can be, but is not limited to, 1.25.

[0088] The aforementioned preset threshold condition may also include: the number of blocks in the associated CTU with selected MMVD modes is greater than a predetermined multiple of the maximum number of blocks in the associated CTU with selected other prediction modes for each prediction mode. >threshold3 max ( , , (2) Where threshold3 is a positive real number greater than 1. For example, threshold3 can be, but is not limited to: 1.3.

[0089] It should be noted that if at least one of formulas (1) and (2) is satisfied, it indicates that there are many CUs in the co-located CTU and adjacent CTUs of the reference frame that have selected the MMVD prediction algorithm. Therefore, it can be determined that the probability of the CUs in the current CTU to be encoded in the current frame selecting the MMVD prediction algorithm is also relatively high. Of course, the preset threshold condition can be, in addition to the threshold condition that limits the number of blocks or the block area mentioned above, it can also be a threshold condition used to limit other statistical values ​​about blocks, as long as it is a statistical value that can characterize the number or proportion of blocks that select each prediction algorithm in the CTU.

[0090] Figure 5 is a schematic diagram illustrating an optimized merged prediction pattern syntax tree structure according to an exemplary embodiment of the present disclosure.

[0091] Referring to Figure 5, in this disclosure, since the encoding position of the MMVD prediction algorithm is advanced to the root (i.e., the MMVD enable flag is encoded first), only `mmvd_enable_flag=1` needs to be encoded first before encoding the syntax elements of the MMVD prediction algorithm itself. Compared to related technologies (refer to Figure 1), this disclosure can save the encoding of the syntax elements `(merge_subblock_flag=0)` and `(regular_merge_flag=1)`, which saves 2 bits of bitrate and effectively reduces the bitrate consumption of MMVD. Furthermore, since the MMVD syntax element optimization of this disclosure is at the CTU level, if it is determined that the encoding position of the MMVD prediction algorithm is advanced for the current CTU to be encoded in the current frame, all CUs contained in the current CTU to be encoded can enjoy the advantage of the advanced MMVD encoding position, that is, the bitrate consumption of MMVD can be reduced for each CU contained in the current CTU to be encoded. Furthermore, the MMVD syntax element optimization disclosed herein is not limited to the CTU level; the MMVD syntax element optimization algorithm can be applied to any other level as needed, such as the stripe level, etc. Additionally, the current block in this disclosure can refer to, but is not limited to, the current CU, or a block at other levels, such as a sub-CU, luma block, chroma block, etc.

[0092] As mentioned earlier, related technologies suffer from high encoding complexity in determining the offset step set for the MMVD mode, leading to low compression efficiency. Therefore, this disclosure proposes a CTU-level dynamic adjustment scheme for the offset step range in the MMVD mode. This scheme utilizes the selection of MMVDs by the CTUs in adjacent reference frames to dynamically optimize the offset step range of the MMVDs corresponding to the CUs contained in the current CTU to be encoded in the current frame. This improves MMVD encoding performance while maintaining encoding quality.

[0093] According to an exemplary embodiment of this disclosure, for the current CTU, a target offset step set can also be selected from a plurality of preset offset step set sets as the offset step set for the current CTU. Each of the plurality of preset offset step set sets can be a subset of the offset step set set set for MMVD mode, and each preset offset step set can be a different subset of each other.

[0094] According to an exemplary embodiment of this disclosure, the aforementioned plurality of preset offset step size sets may include a first preset offset step size set and a second preset offset step size set. The second preset offset step size set may include a larger range of offset step size candidates than the first preset offset step size set.

[0095] According to an exemplary embodiment of this disclosure, in response to the current CTU satisfying the preset offset step size condition, a second preset offset step size set can be selected as the target offset step size set; in response to the current CTU not satisfying the preset offset step size condition, a first preset offset step size set can be selected as the target offset step size set.

[0096] According to an exemplary embodiment of this disclosure, the preset offset step size condition may include: the average offset step size of the blocks of the selected MMVD mode in the associated CTU related to the current CTU is greater than a preset threshold. For example, the preset offset step size condition may be expressed as the following two formulas: (3) (4) Among them, It can be a custom real number greater than or equal to 1, for example, This can be, but is not limited to, 2. That is, when the offset step size of the block selecting the MMVD mode in the associated CTU is large, the probability of the block selecting the MMVD mode in the current CTU choosing a larger offset step size is also higher. Therefore, a larger set of offset step sizes can be selected. When the offset step size of the block selecting the MMVD mode in the associated CTU is small, the probability of the block selecting the MMVD mode in the current CTU choosing a smaller offset step size is also higher. Therefore, a smaller set of offset step sizes can be selected. Thus, by utilizing this high correlation between the selection of MMVD by CTUs between adjacent video frames in the temporal domain, the offset step size range of the MMVD corresponding to the CU contained in the current CTU in the current frame can be dynamically selected, thereby improving coding performance.

[0097] According to an exemplary embodiment of this disclosure, the first preset offset step size set S1 can be {1 / 4, 1 / 2, 1, 2, 4}, and the second preset offset step size set S2 can be {1 / 4, 1 / 2, 1, 2, 4, 8, 16}. It can be seen that the second preset offset step size set S2 includes elements from the first preset offset step size set S1, namely 1 / 4, 1 / 2, 1, 2, 4, as well as elements 8 and 16. Therefore, the second preset offset step size set S2 has a larger range of candidate offset step sizes than the first preset offset step size set S1. Of course, this disclosure does not limit the number and range of preset offset step size sets; more or fewer preset offset step size sets can be set as needed, or preset offset step size sets with other ranges can be set.

[0098] In addition, the encoding method for the first preset offset step size set S1 can be shown in Table 4:

[0099] Table 4. Meaning of MMVD_step_idx syntax elements. The encoding method for the second preset offset step set S2 can be shown in Table 5:

[0100] Table 5 shows the meaning of the MMVD_step_idx syntax elements. Figure 6 is a flowchart illustrating a video decoding method according to an exemplary embodiment of the present disclosure.

[0101] Referring to Figure 6, in step 601, in response to the current block in the current coding tree unit (CTU) of the current frame being in merge prediction mode, the current block can be decoded based on the merge prediction mode syntax tree structure. Specifically, if the first statistical value of the current CTU satisfies a preset threshold condition, the merge prediction mode syntax tree structure can be adjusted to first decode the MMVD enable flag. The preset threshold condition can be that the first statistical value exceeds a threshold related to the preset threshold condition. The first statistical value can be the statistical value of the block in the selected merge prediction mode of the MMVD mode in the associated CTUs related to the current CTU. The associated CTUs can include the co-located CTUs in the reference frame of the current frame and the adjacent CTUs of the co-located CTUs.

[0102] According to an exemplary embodiment of this disclosure, the threshold related to the preset threshold condition can be determined based on a second statistical value for the current CTU, wherein the second statistical value can be the statistical value of the block of other prediction modes besides the MMVD mode in the selected merge prediction modes of the associated CTU.

[0103] According to an exemplary embodiment of the present disclosure, the preset threshold condition may include at least one of the following conditions: the area of ​​the block with the selected MMVD mode in the associated CTU is greater than a predetermined multiple of the largest area of ​​the block with the selected other prediction modes in the associated CTU; the number of blocks with the selected MMVD mode in the associated CTU is greater than a predetermined multiple of the maximum number of blocks with the selected other prediction modes in the associated CTU.

[0104] According to an exemplary embodiment of this disclosure, the aforementioned reference frame may be the nearest neighbor reference frame of the current frame.

[0105] According to an exemplary embodiment of this disclosure, for the current CTU, a target offset step set can also be selected from a plurality of preset offset step set sets as the offset step set for the current CTU, wherein each of the plurality of preset offset step set sets can be a subset of the offset step set set for the MMVD mode, and each preset offset step set can be a different subset of each other.

[0106] According to an exemplary embodiment of this disclosure, the aforementioned plurality of preset offset step size sets may include a first preset offset step size set and a second preset offset step size set, wherein the second preset offset step size set may include a larger range of offset step size candidates than the first preset offset step size set.

[0107] According to an exemplary embodiment of this disclosure, the aforementioned first preset offset step size set can be {1 / 4, 1 / 2, 1,2, 4}, and the aforementioned second preset offset step size set can be {1 / 4, 1 / 2, 1, 2, 4, 8, 16}.

[0108] According to an exemplary embodiment of this disclosure, an index indicating the target offset step set can also be received at the CTU level. Then, the target offset step set can be selected from multiple preset offset step sets based on the received index. It should be noted that, in this disclosure, besides selecting the target offset step set by receiving an index at the CTU level, other methods can also be used to select the target offset step set. For example, the decoding end can select the target offset step set in the same way as the encoding end; that is, the target offset step set can be selected from multiple preset offset step sets based on the selection of MMVD mode blocks in the associated CTU.

[0109] Figure 7 is a block diagram illustrating a video encoding apparatus 700 according to an exemplary embodiment of the present disclosure.

[0110] Referring to Figure 7, the video encoding device 700 may include a statistics module 701, an adjustment module 702, and an encoding module 703.

[0111] The statistics module 701 is configured to: obtain a first statistical value for the current coding tree unit (CTU) in the current frame, wherein the first statistical value is the statistical value of the block in the selected merge prediction mode of the associated CTU related to the current CTU, and the associated CTU includes the co-located CTU in the reference frame of the current frame and the adjacent CTU of the co-located CTU; the adjustment module 702 is configured to: adjust the merge prediction mode syntax tree structure to first encode the MMVD enable flag in response to the first statistical value satisfying a preset threshold condition, wherein the preset threshold condition is that the first statistical value exceeds a threshold related to the preset threshold condition; the encoding module 703 is configured to: encode the current block based on the adjusted merge prediction mode syntax tree structure in response to the current block in the current CTU being in merge prediction mode.

[0112] According to an exemplary embodiment of this disclosure, the threshold related to the preset threshold condition is determined based on a second statistical value for the current CTU, wherein the second statistical value is the statistical value of the block of other prediction modes besides the MMVD mode in the selected merge prediction modes in the associated CTU.

[0113] According to an exemplary embodiment of the present disclosure, the preset threshold condition includes at least one of the following conditions: the area of ​​the block with the selected MMVD mode in the associated CTU is greater than a predetermined multiple of the largest area of ​​the block with the selected other prediction modes in the associated CTU; the number of blocks with the selected MMVD mode in the associated CTU is greater than a predetermined multiple of the maximum number of blocks with the selected other prediction modes in the associated CTU.

[0114] According to an exemplary embodiment of this disclosure, the reference frame is the nearest neighbor reference frame of the current frame.

[0115] According to an exemplary embodiment of the present disclosure, the video encoding apparatus 700 further includes: an offset step set selection module, configured to select a target offset step set from a plurality of preset offset step sets for the current CTU, as an offset step set for the current CTU; wherein each of the plurality of preset offset step sets is a subset of the offset step set set set for the MMVD mode, and each preset offset step set is a different subset of each other.

[0116] According to an exemplary embodiment of the present disclosure, the plurality of preset offset step size sets include a first preset offset step size set and a second preset offset step size set, wherein the second preset offset step size set includes a larger range of offset step size candidates than the first preset offset step size set.

[0117] According to an exemplary embodiment of this disclosure, the first preset offset step size set is {1 / 4, 1 / 2, 1, 2, 4}, and the second preset offset step size set is {1 / 4, 1 / 2, 1, 2, 4, 8, 16}.

[0118] According to an exemplary embodiment of this disclosure, the offset step size set selection module is configured to: select a second preset offset step size set as the target offset step size set in response to the current CTU meeting the preset offset step size condition; and select a first preset offset step size set as the target offset step size set in response to the current CTU not meeting the preset offset step size condition; wherein the preset offset step size condition includes: the average offset step size of the blocks of the selected MMVD mode in the associated CTU is greater than a preset threshold.

[0119] Figure 8 is a block diagram illustrating a video decoding apparatus 800 according to an exemplary embodiment of the present disclosure.

[0120] Referring to Figure 8, the video decoding device 800 may include a decoding module 801.

[0121] The decoding module 801 is configured to decode the current block based on the merge prediction mode syntax tree structure in response to the current block in the current coding tree unit (CTU) of the current frame being in merge prediction mode. Specifically, if the first statistical value of the current CTU satisfies a preset threshold condition, the merge prediction mode syntax tree structure is adjusted to first decode the MMVD enable flag. The preset threshold condition is that the first statistical value exceeds a threshold related to the preset threshold condition. The first statistical value is the statistical value of the block in the selected merge prediction mode of the MMVD mode in the associated CTUs related to the current CTU. The associated CTUs include the co-located CTUs in the reference frame of the current frame and the adjacent CTUs of the co-located CTUs.

[0122] According to an exemplary embodiment of this disclosure, the threshold related to the preset threshold condition is determined based on a second statistical value for the current CTU, wherein the second statistical value is the statistical value of the block of other prediction modes besides the MMVD mode in the selected merge prediction modes in the associated CTU.

[0123] According to an exemplary embodiment of the present disclosure, the preset threshold condition includes at least one of the following conditions: the area of ​​the block with the selected MMVD mode in the associated CTU is greater than a predetermined multiple of the largest area of ​​the block with the selected other prediction modes in the associated CTU; the number of blocks with the selected MMVD mode in the associated CTU is greater than a predetermined multiple of the maximum number of blocks with the selected other prediction modes in the associated CTU.

[0124] According to an exemplary embodiment of this disclosure, the reference frame is the nearest neighbor reference frame of the current frame.

[0125] According to an exemplary embodiment of the present disclosure, the video decoding apparatus 800 further includes: an offset step set selection module, configured to select a target offset step set from a plurality of preset offset step sets for the current CTU, as an offset step set for the current CTU; wherein each of the plurality of preset offset step sets is a subset of the offset step set set set for the MMVD mode, and each preset offset step set is a different subset of each other.

[0126] According to an exemplary embodiment of the present disclosure, the plurality of preset offset step size sets include a first preset offset step size set and a second preset offset step size set, wherein the second preset offset step size set includes a larger range of offset step size candidates than the first preset offset step size set.

[0127] According to an exemplary embodiment of this disclosure, the first preset offset step size set is {1 / 4, 1 / 2, 1, 2, 4}, and the second preset offset step size set is {1 / 4, 1 / 2, 1, 2, 4, 8, 16}.

[0128] According to an exemplary embodiment of the present disclosure, the video decoding apparatus 800 further includes: an index receiving module configured to receive an index indicating a target offset step set at the CTU level; the offset step set selection module is configured to select a target offset step set from a plurality of preset offset step sets according to the received index.

[0129] Figure 9 illustrates a computing environment 910 coupled to a user interface 950. The computing environment 910 may be part of a data processing server. The computing environment 910 includes a processor 920, a memory 930, and an input / output (I / O) interface 940.

[0130] Processor 920 typically controls the overall operation of computing environment 910, such as operations associated with display, data acquisition, data communication, and image processing. Processor 920 may include one or more processors for executing instructions to perform all or some of the steps in the methods described above. Furthermore, processor 920 may include one or more modules that facilitate interaction between processor 920 and other components. The processor may be a central processing unit (CPU), microprocessor, microcontroller, graphics processing unit (GPU), etc.

[0131] Memory 930 is configured to store various types of data to support the operation of computing environment 910. Memory 930 may include predefined software 932. Examples of such data include instructions for any application or method operating on computing environment 910, video datasets, image data, etc. Memory 930 can be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0132] I / O interface 940 provides an interface between processor 920 and peripheral interface modules (such as keyboard, click wheel, buttons, etc.). Buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. I / O interface 940 can be coupled to encoders and decoders.

[0133] In an embodiment, a non-transitory computer-readable storage medium is also provided, including, for example, a memory 930, a plurality of programs, and / or a bitstream generated by the above-described encoding method or a bitstream to be decoded by the above-described decoding method. The plurality of programs can be executed by a processor 920 in a computing environment 910 to perform the above-described methods. In one example, the plurality of programs can be executed by the processor 920 in the computing environment 910 to receive a bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.), and can also be executed by the processor 920 in the computing environment 910 to perform the above-described decoding method based on the received bitstream or data stream. In another example, the plurality of programs can be executed by the processor 920 in the computing environment 910 to perform the above-described encoding method to encode video information (e.g., video blocks representing video frames, and / or one or more associated syntax elements, etc.) into a bitstream or data stream, and can also be executed by the processor 920 in the computing environment 910 to transmit the bitstream or data stream. Alternatively, the non-transitory computer-readable storage medium may store a bitstream or data stream, generated by an encoder using, for example, the encoding methods described above, for use by a decoder when decoding video data. This bitstream includes encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.). The non-transitory computer-readable storage medium may be, for example, ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.

[0134] In one embodiment, a bitstream generated by the above-described encoding method or a bitstream to be decoded by the above-described decoding method is provided. In another embodiment, a bitstream comprising encoded video information generated by the above-described encoding method or encoded video information to be decoded by the above-described decoding method is provided.

[0135] In one embodiment, a computing device is also provided, comprising: one or more processors (e.g., processor 920); and a non-transitory computer-readable storage medium or memory 930 therein storing a plurality of programs executable by the one or more processors, wherein the one or more processors are configured to perform the methods described above when executing the plurality of programs.

[0136] In one embodiment, a computer program product having instructions for storing or transmitting a bitstream is also provided, the bitstream including encoded video information generated by the encoding method described above or encoded video information to be decoded by the decoding method described above. In another embodiment, a computer program product including, for example, a plurality of programs in a memory 930 is also provided, the plurality of programs being executable by a processor 920 in a computing environment 910 to perform the video encoding method or video decoding method described above. For example, the computer program product may include a non-transitory computer-readable storage medium.

[0137] In an embodiment, the computing environment 910 may be implemented by one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components for performing the methods described above.

[0138] In one embodiment, a method for storing a bitstream is also provided, comprising: storing the bitstream on a digital storage medium, wherein the bitstream includes encoded video information generated by the above-described encoding method or encoded video information to be decoded by the above-described decoding method.

[0139] In one embodiment, a method for transmitting a bitstream generated by the encoder described above is also provided. In another embodiment, a method for receiving a bitstream decoded by the decoder described above is also provided.

[0140] According to the video encoding and decoding methods, apparatus, electronic devices, storage media, methods for storing bit streams, and program products provided in this disclosure, since the selection of MMVD by the CTU between adjacent video frames in the temporal domain is highly correlated, in this disclosure, the encoding order of the syntax elements of the MMVD of the CTU in the current frame can be dynamically adjusted based on the selection of MMVD by the CTU in the reference frame adjacent to the current frame. That is, the encoding position of the MMVD syntax elements of the CTU in the current frame in the merged prediction mode syntax tree structure can be dynamically adjusted, thereby effectively reducing the bit rate consumption of MMVD and improving the encoding efficiency of MMVD.

[0141] Furthermore, according to this disclosure, the selection status of MMVD by the CTU of the reference frame adjacent to the current frame can be used to dynamically optimize the offset step range of the MMVD corresponding to the CU contained in the current CTU to be encoded in the current frame, thereby improving the encoding performance of MMVD while ensuring encoding quality.

[0142] The description in this disclosure has been presented for illustrative purposes and is not intended to be exhaustive or limited to this disclosure. Many modifications, variations, and alternative embodiments will be apparent to those skilled in the art from the teachings presented in the foregoing description and the associated drawings.

[0143] Unless otherwise specifically stated, the order of steps in the method according to this disclosure is intended to be illustrative only, and the steps of the method according to this disclosure are not limited to the specific order described above, but may be changed according to actual circumstances. Furthermore, at least one step in the method according to this disclosure may be adjusted, combined, or omitted as needed.

[0144] The examples chosen and described are intended to explain the principles of this disclosure and to enable others skilled in the art to understand the various embodiments of this disclosure, and preferably to utilize the basic principles and various embodiments with various modifications suitable for the intended particular purpose. Therefore, it will be understood that the scope of this disclosure is not limited to the specific examples of the disclosed embodiments, and that modifications and other embodiments are intended to be included within the scope of this disclosure.

Claims

1. A video encoding method, characterized in that, include: Obtain a first statistical value for the current coding tree unit (CTU) in the current frame, wherein the first statistical value is the statistical value of the block of the MMVD mode in the selected merge prediction mode among the associated CTUs related to the current CTU, and the associated CTUs include the co-located CTUs in the reference frame of the current frame that are in the same position as the current CTU and the adjacent CTUs of the co-located CTUs. In response to the first statistical value meeting a preset threshold condition, the merge prediction mode syntax tree structure is adjusted to first encode the MMVD enable flag, wherein the preset threshold condition is that the first statistical value exceeds a threshold related to the preset threshold condition; in response to the current block in the current CTU being in merge prediction mode, the current block is encoded based on the adjusted merge prediction mode syntax tree structure.

2. The video encoding method as described in claim 1, characterized in that, The threshold related to the preset threshold condition is determined based on a second statistical value for the current CTU, wherein the second statistical value is the statistical value of the block of the selected prediction mode other than the MMVD mode in the merged prediction mode of the associated CTU.

3. The video encoding method as described in claim 2, characterized in that, The preset threshold condition includes at least one of the following conditions: the area of ​​the block with the selected MMVD mode in the associated CTU is greater than a predetermined multiple of the largest area of ​​the block with the selected prediction mode in each of the other prediction modes in the associated CTU; the number of blocks with the selected MMVD mode in the associated CTU is greater than a predetermined multiple of the maximum number of blocks with the selected prediction mode in each of the other prediction modes in the associated CTU.

4. The video encoding method as described in claim 1, characterized in that, The reference frame is the nearest neighbor reference frame of the current frame.

5. The video encoding method as described in claim 1, characterized in that, Also includes: For the current CTU, a target offset step set is selected from multiple preset offset step set sets as the offset step set for the current CTU; wherein each preset offset step set in the multiple preset offset step set is a subset of the offset step set set for MMVD mode, and each preset offset step set is a different subset of each other.

6. The video encoding method as described in claim 5, characterized in that, The plurality of preset offset step size sets include a first preset offset step size set and a second preset offset step size set, wherein the second preset offset step size set includes a larger range of offset step size candidates than the first preset offset step size set.

7. The video encoding method as described in claim 6, characterized in that, The first preset offset step size set is {1 / 4, 1 / 2, 1, 2, 4}, and the second preset offset step size set is {1 / 4, 1 / 2, 1, 2, 4, 8, 16}.

8. The video encoding method as described in claim 6, characterized in that, The step of selecting a target offset step set from multiple preset offset step set includes: in response to the current CTU satisfying the preset offset step condition, selecting the second preset offset step set as the target offset step set; in response to the current CTU not satisfying the preset offset step condition, selecting the first preset offset step set as the target offset step set; wherein, the preset offset step condition includes: the average offset step size of the blocks in the selected MMVD mode in the associated CTU is greater than a preset threshold.

9. A video decoding method, characterized in that, include: In response to the current block in the current coding tree unit (CTU) of the current frame being in merge prediction mode, the current block is decoded based on the merge prediction mode syntax tree structure. Wherein, if the first statistical value of the current CTU satisfies a preset threshold condition, the merge prediction mode syntax tree structure is adjusted to first decode the MMVD enable flag. The preset threshold condition is when the first statistical value exceeds a threshold related to the preset threshold condition. The first statistical value is the statistical value of the block in the merge prediction mode selected in the associated CTUs related to the current CTU. The associated CTUs include the co-located CTUs in the reference frame of the current frame and the adjacent CTUs of the co-located CTUs.

10. A video encoding device, characterized in that, include: The statistics module is configured to: obtain a first statistical value for the current coding tree unit (CTU) in the current frame, wherein the first statistical value is the statistical value of the block in the selected merge prediction mode of the associated CTU in the merge prediction mode, and the associated CTU includes the co-located CTU in the reference frame of the current frame and the adjacent CTU of the co-located CTU; the adjustment module is configured to: adjust the merge prediction mode syntax tree structure to first encode the MMVD enable flag in response to the first statistical value satisfying a preset threshold condition, wherein the preset threshold condition is that the first statistical value exceeds a threshold related to the preset threshold condition; the encoding module is configured to: encode the current block based on the adjusted merge prediction mode syntax tree structure in response to the current block in the current CTU being in the merge prediction mode.

11. A video decoding device, characterized in that, include: The decoding module is configured to decode the current block based on the merge prediction mode syntax tree structure in response to the current block in the current coding tree unit (CTU) of the current frame being in merge prediction mode. Specifically, if a first statistical value of the current CTU satisfies a preset threshold condition, the merge prediction mode syntax tree structure is adjusted to first decode the MMVD enable flag. The preset threshold condition is when the first statistical value exceeds a threshold related to the preset threshold condition. The first statistical value is the statistical value of the block in the merge prediction mode selected in the associated CTUs related to the current CTU. The associated CTUs include the co-located CTUs in the reference frame of the current frame and the adjacent CTUs of the co-located CTUs.

12. An electronic device, characterized in that, include: processor; A memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the video encoding method as described in any one of claims 1 to 8, or to implement the video decoding method as described in claim 9.

13. A non-transitory computer-readable storage medium storing instructions and a bit stream, wherein, When executed by a computing device having one or more processors, the instructions cause the one or more processors to perform the video encoding method according to any one of claims 1 to 8 to generate the bitstream.

14. A method for storing a bit stream, comprising: The video encoding method according to any one of claims 1 to 8 is used to generate a bitstream; Store the bit stream.

15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the video encoding method as described in any one of claims 1 to 8, or implements the video decoding method as described in claim 9.