Video encoding method, video decoding method, and related apparatus

CN122824894APending Publication Date: 2026-09-25ZHONGXING INTELLIGENT SYST TECH CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611001089.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]然而,在如快速运动、物体遮挡等复杂图像场景下,多存在帧间画面差异较大的问题,如果直接从历史重建帧中选取真实参考帧作为重建依据,则可能难以准确表征当前帧的图像内容,导致预测失准,进而引入较大的残差数据量,这不仅会增大码率开销,还严重影响当前帧的重建质量

Benefits of technology

[0019]本申请实施例中,基于各历史重建帧,确定当前帧中各图像块各自对应的多个候选参考帧;其中,多个候选参考帧包括:从各历史重建帧中选取的真实参考帧、基于图像生成模型对历史重建帧特征提取,生成的虚拟参考帧、基于真实参考帧和/或虚拟参考帧确定的融合参考帧。上述设置不仅充分利用了历史重建帧所提供的内容新型,还通过图像生成模型创建出能够补充或增强历史重建帧的虚拟参考帧,并通过对真实参考帧和/或虚拟参考帧进行融合处理,生成融合参考帧,这样可在保障数据可靠性的同时,有效扩充候选参考帧的内容丰富度,进而提高重建质量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824894A_ABST
    Figure CN122824894A_ABST
Patent Text Reader

Abstract

The application provides a video encoding method, a video decoding method and related devices, which are used for improving the reconstruction quality of a current frame and reducing the code rate overhead. The method comprises the following steps: determining a plurality of candidate reference frames corresponding to each image block in the current frame based on each historical reconstructed frame; wherein the plurality of candidate reference frames comprise: a real reference frame selected from the historical reconstructed frames, a virtual reference frame generated based on image generation model and historical reconstructed frame feature extraction, and a fusion reference frame determined based on the real reference frame and / or the virtual reference frame; for each image block, selecting a target reference frame used for reconstructing the image block from the plurality of candidate reference frames based on the multi-dimensional content features of the image block; and adding the frame type identification of the target reference frame corresponding to each image block into a code stream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video encoding and decoding technology, and in particular to a video encoding method, a video decoding method, and related apparatus. Background Technology

[0002] To support fast start-up and smooth playback of video streams, video encoding and decoding tasks typically do not involve independent and complete encoding of each video frame. Instead, reference frames are selected from previously decoded historical reconstructed frames, and the image content of the current frame is reconstructed based on these reference frames, thereby reducing bitrate overhead.

[0003] However, in complex image scenarios such as fast motion and object occlusion, there are often large differences between frames. If the real reference frame is directly selected from the historical reconstruction frames as the basis for reconstruction, it may be difficult to accurately represent the image content of the current frame, resulting in inaccurate prediction and introducing a large amount of residual data. This will not only increase the bit rate overhead, but also seriously affect the reconstruction quality of the current frame.

[0004] Therefore, improving the reconstruction quality of the current frame and reducing bitrate overhead are urgent problems that need to be solved. Summary of the Invention

[0005] This application provides a video encoding method, a video decoding method, and related apparatus for improving the reconstruction quality of the current frame and reducing bit rate overhead.

[0006] In a first aspect, embodiments of this application provide a video encoding method, the method comprising: Based on each historical reconstructed frame, multiple candidate reference frames are determined for each image block in the current frame; wherein, the multiple candidate reference frames include: real reference frames selected from each historical reconstructed frame, virtual reference frames generated by extracting features from historical reconstructed frames based on an image generation model, and fused reference frames determined based on the real reference frames and / or the virtual reference frames. For each image block, a target reference frame for reconstructing the image block is selected from the plurality of candidate reference frames based on the multidimensional content features of the image block; Add a frame type identifier for the target reference frame corresponding to each image block to the bitstream.

[0007] Secondly, embodiments of this application provide a video decoding method, the method comprising: The frame type identifier of each image block in the current frame is obtained by decoding from the bitstream; wherein, the frame type identifier is the identifier of the target reference frame; the target reference frame is selected from multiple candidate reference frames, the multiple candidate reference frames include: real reference frames selected from each historical reconstructed frame, virtual reference frames generated by extracting features from historical reconstructed frames based on the image generation model, and fused reference frames determined based on the real reference frames and / or the virtual reference frames; For each image block, the target reference frame indicated by the frame type identifier is obtained, and the image block is reconstructed based on the target reference frame to obtain the reconstructed block corresponding to the image block; The reconstructed frame of the current frame is determined based on the reconstructed block corresponding to each of the image blocks.

[0008] In some embodiments, the target reference frame is determined in the following manner: For each candidate reference frame, motion estimation is performed on the image block using the candidate reference frame to obtain motion vector data of the image block, and the rate-distortion cost between the image block and the candidate reference frame is determined based on the motion vector data and the multidimensional content features of the image block. The candidate reference frame with the lowest rate-distortion cost is used as the target reference frame for the image block.

[0009] In some embodiments, determining the rate-distortion cost between the image patch and the candidate reference frame based on the motion vector data and the multidimensional content features of the image patch includes: Based on the motion vector data, the prediction block corresponding to the image block in the candidate reference frame is determined; Based on the pixel difference between the image block and the prediction block, residual data is determined; Based on the multidimensional content features and the residual data, the rate-distortion cost between the image patch and the candidate reference frame is determined.

[0010] In some embodiments, the method further includes: If the target reference frame is the fusion reference frame, then the fusion identifier is decoded from the bitstream; The fusion reference frame is generated using the generation method indicated by the fusion identifier. The fusion reference frame is determined in the following way: If the difference between the rate-distortion cost between the image patch and the real reference frame and the rate-distortion cost between the image patch and the virtual reference frame is less than a preset threshold, then the trained image fusion model is used to perform feature fusion on the real reference frame and the virtual reference frame to obtain the fused reference frame. If the difference is not less than the preset threshold, then image enhancement processing is performed on the specified reference frame to obtain the fused reference frame; the specified reference frame is the one with the smaller rate-distortion cost between the real reference frame and the virtual reference frame.

[0011] In some embodiments, the true reference frame is determined in the following manner: Extract the first image features of the image block; Based on the position data of the image block in the current frame, determine the second image features of each historical reconstructed frame at the same position; Determine the feature similarity between the first image feature and the corresponding second image feature of each of the historical reconstructed frames; Based on the feature similarity, the real reference frame is selected from each of the historical reconstructed frames.

[0012] In some embodiments, the image generation model is obtained through multiple rounds of iterative training based on a set of historical video frames; wherein each iteration is as follows: Select N consecutive historical video frames from the set of historical video frames; where N≥3; From the N historical video frames, select one frame as the training label and use the remaining historical video frames (excluding the first frame) as training samples. Based on the current model parameters, feature extraction is performed on the training samples to generate virtual video frames for this round; The iteration loss for this round is determined based on the feature differences between the virtual video frame and the sample label; The current model parameters are adjusted based on the loss from this iteration.

[0013] Thirdly, embodiments of this application provide a video encoding apparatus, the apparatus comprising: The reference frame acquisition unit is configured to: determine multiple candidate reference frames corresponding to each image block in the current frame based on each historical reconstruction frame; wherein, the multiple candidate reference frames include: real reference frames selected from each historical reconstruction frame, virtual reference frames generated based on feature extraction of historical reconstruction frames by an image generation model, and fused reference frames determined based on the real reference frames and / or the virtual reference frames; The reference frame confirmation unit is configured to: for each image block, select a target reference frame for reconstructing the image block from the plurality of candidate reference frames based on the multidimensional content features of the image block; The encoding unit is configured to add a frame type identifier of the target reference frame corresponding to each image block to the bitstream.

[0014] Fourthly, embodiments of this application provide a video decoding apparatus, the apparatus comprising: The decoding unit is configured to: decode from the bitstream to obtain the frame type identifier of each image block in the current frame; wherein the frame type identifier is the identifier of the target reference frame; the target reference frame is selected from multiple candidate reference frames, the multiple candidate reference frames including: real reference frames selected from each historical reconstructed frame, virtual reference frames generated based on feature extraction of historical reconstructed frames based on an image generation model, and fused reference frames determined based on the real reference frames and / or the virtual reference frames; The reconstruction block unit is configured to: for each image block, obtain a target reference frame indicated by the frame type identifier, and reconstruct the image block based on the target reference frame to obtain a reconstruction block corresponding to the image block; The reconstruction frame unit is configured to determine the reconstruction frame of the current frame based on the reconstruction block corresponding to each of the image blocks.

[0015] Fifthly, embodiments of this application provide a video encoder, including: Memory, used to store computer programs; A processor, configured to execute the video encoding method as described in any of the first aspects when running the computer program.

[0016] Sixthly, embodiments of this application provide a video decoder, including: Memory, used to store computer programs; A processor, configured to execute the video decoding method as described in any of the second aspects when running the computer program.

[0017] In a seventh aspect, embodiments of this application provide a computer-readable storage medium storing a computer program / instructions and a bit stream thereon, wherein the computer program / instructions, when executed by a processor, are capable of generating the bit stream according to the method described in any one of the first or second aspects.

[0018] Eighthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method as described in any one of the first or second aspects.

[0019] In this embodiment, based on each historical reconstructed frame, multiple candidate reference frames are determined for each image block in the current frame. These candidate reference frames include: real reference frames selected from each historical reconstructed frame; virtual reference frames generated by feature extraction from historical reconstructed frames using an image generation model; and fused reference frames determined based on real and / or virtual reference frames. This setup not only fully utilizes the novel content provided by historical reconstructed frames but also creates virtual reference frames that can supplement or enhance historical reconstructed frames through an image generation model. Furthermore, by fusing real and / or virtual reference frames, fused reference frames are generated. This effectively expands the content richness of candidate reference frames while ensuring data reliability, thereby improving reconstruction quality.

[0020] Next, for each image block in the current frame, a target reference frame for reconstruction is selected from multiple candidate reference frames. Then, a frame type identifier for the target reference frame corresponding to each image block is added to the bitstream. Through this setting, each image block in the current frame can adaptively select the reference frame type most suitable for its content features, thereby reducing the amount of residual data and improving reconstruction quality. Simultaneously, by adding a small number of bits of frame type identifier to the bitstream to indicate the type of the target reference frame corresponding to the image block, the corresponding target reference frame can be quickly obtained during the decoding stage based on the frame type identifier for subsequent reconstruction operations, effectively reducing bitrate overhead. Attached Figure Description

[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments of this application will be briefly introduced below. The drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a schematic diagram of the video encoding and decoding process in related technologies.

[0023] Figure 2 This is an overall flowchart of a video encoding method provided in an embodiment of this application.

[0024] Figure 3 This is a schematic diagram illustrating the process of obtaining a real reference frame for an embodiment of this application.

[0025] Figure 4 This is an overall flowchart of a video decoding method provided in an embodiment of this application.

[0026] Figure 5 This is a structural block diagram of a video encoding device provided in an embodiment of this application.

[0027] Figure 6This is a structural block diagram of a video decoding device provided in an embodiment of this application.

[0028] Figure 7 This is a structural block diagram of a video decoder or video encoder provided in an embodiment of this application. Detailed Implementation

[0029] The technical solutions in the embodiments of this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " will mean "or", for example, A / B can mean A or B; "and / or" in the text is merely a description of the relationship between related objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "multiple" means two or more.

[0030] In the description of the embodiments of this application, unless otherwise stated, the term "multiple" refers to two or more, and other quantifiers are similarly understood. The preferred embodiments described herein are for illustration and explanation only and are not intended to limit this application. Furthermore, the embodiments and features in the embodiments of this application can be combined with each other without conflict.

[0031] To further illustrate the technical solutions provided in the embodiments of this application, a detailed description is provided below in conjunction with the accompanying drawings and specific implementation methods. Although the embodiments of this application provide method operation steps as shown in the following embodiments or drawings, more or fewer operation steps may be included in the method based on conventional or non-inventive effort. For steps that do not logically have a necessary causal relationship, the execution order of these steps is not limited to the execution order provided in the embodiments of this application. In actual processing or when the control device executes the method, it may be executed sequentially or in parallel according to the method shown in the embodiments or drawings.

[0032] As mentioned earlier, in video encoding and decoding tasks, each video frame is not usually encoded independently and completely. Instead, inter-frame prediction is used to select reference frames from the decoded historical reconstructed frames and reconstruct the image content of the current frame based on the reference frames, thereby reducing bitrate overhead.

[0033] Specifically, in current video codec standards (such as H.264 / AVC, H.266 / VVC, etc.), the workflow of a video encoder is typically as follows: Figure 1 As shown: First, the current frame is divided into multiple image blocks (called coding units (CUs) or prediction units (PUs). Then, for each image block, motion estimation is performed. The region with the highest matching degree between the image block and the historical reconstructed frames (i.e., decoded video frames) cached in the Decoded Picture Buffer (DPB) is traversed. The historical reconstructed frame containing this region is the reference frame used to reconstruct the current frame.

[0034] Next, based on the displacement difference between the image patch and the region, motion vector data (MV) of the image patch is generated. Motion compensation is then performed on this motion vector data to generate the corresponding prediction block. Then, the residual data between the image patch and the prediction block is calculated. This residual data, motion vector data, and reference frame identifier (or index) are transformed, quantized, and then entropy encoded to form a compressed bitstream, which is transmitted to the video decoder.

[0035] Continue as Figure 1 As shown, the video decoder parses the motion vector data, reference frame identifier, and residual data of the current image block from the received compressed bitstream. Then, based on the reference frame identifier, it retrieves the corresponding historical reconstructed frame from the local DPB, thereby obtaining the reference frame corresponding to the current image block.

[0036] Then, motion compensation calculations are performed on the motion vector data to generate prediction blocks consistent with the video encoder. The parsed residual data is then inversely quantized and inversely transformed to obtain reconstructed residual data. Subsequently, the prediction blocks and reconstructed residual data are added together to obtain the reconstructed block corresponding to the current image block. Finally, all reconstructed image blocks are combined into a complete reconstructed frame and stored in the DPB for use in encoding / decoding subsequent video frames or image display.

[0037] Because of the above method, only real reference frames can be selected from historical reconstructed frames as the basis for reconstruction. Other types of reference frames cannot be introduced. The selected reference frames are often difficult to accurately represent the image content of the current frame, thus introducing a large amount of residual data. This not only increases the bit rate overhead, but also seriously affects the reconstruction quality of the current frame.

[0038] Currently, although some manufacturers have attempted to introduce fusion reference frames to improve prediction performance, such methods still have significant limitations in practical applications. Specifically, the current method of generating fusion reference frames mainly involves dynamically selecting several frames from multiple cached historical reconstruction frames, assigning them specific fusion weights, and then fusing them into a single image by weighted averaging or linear combination. This image serves as the fusion reference frame used to reconstruct the current frame.

[0039] On the one hand, the reference sources of the aforementioned fusion reference frames are entirely limited to historical reconstructed frames. Therefore, as mentioned earlier, in complex image scenarios such as fast motion and object occlusion, even after fusion, its predictive ability is still limited and it is difficult to effectively represent the true features of the current image patch.

[0040] On the other hand, since DPB typically caches multiple historical reconstructed frames, the number of selectable frame combinations and their corresponding weight configuration schemes grows exponentially. To accurately indicate the fusion strategy employed, the video encoder needs to explicitly transmit multi-bit identification information (such as a reference frame index list, weight index, etc.) in the bitstream. This not only significantly increases the complexity of syntax parsing but also introduces a considerable additional bitrate overhead.

[0041] In view of this, the inventive concept of the present application is as follows: based on each historical reconstruction frame, determine multiple candidate reference frames corresponding to each image block in the current frame; wherein, the multiple candidate reference frames include: real reference frames selected from each historical reconstruction frame, virtual reference frames generated based on feature extraction of historical reconstruction frames by an image generation model, and fused reference frames determined based on real reference frames and / or virtual reference frames.

[0042] The above settings not only make full use of the novel content provided by historical reconstructed frames, but also create virtual reference frames that can supplement or enhance historical reconstructed frames through image generation models. By fusing real reference frames and / or virtual reference frames, fused reference frames are generated. This can effectively expand the content richness of candidate reference frames while ensuring data reliability, thereby improving reconstruction quality.

[0043] Next, for each image block in the current frame, a target reference frame for reconstructing the image block is selected from multiple candidate reference frames, and then the frame type identifier of the target reference frame corresponding to each image block is added to the bitstream.

[0044] With the above settings, each image block in the current frame can adaptively select the reference frame type that best matches its content features, thereby reducing the amount of residual data and improving reconstruction quality. Simultaneously, by adding a small number of frames as a frame type identifier to the bitstream, the type of the target reference frame corresponding to the image block is indicated. This allows for the rapid acquisition of the corresponding target reference frame during the decoding stage for subsequent reconstruction operations, effectively reducing bitrate overhead.

[0045] After introducing the inventive concept of the embodiments of this application, the following is a detailed description of a video encoding method provided by the embodiments of this application. This method can be applied to a video encoder, specifically as follows: Figure 2 As shown, the method includes the following steps: Step S21: Based on each historical reconstruction frame, determine multiple candidate reference frames corresponding to each image block in the current frame; wherein, the multiple candidate reference frames include: real reference frames selected from each historical reconstruction frame, virtual reference frames generated by extracting features from historical reconstruction frames based on an image generation model, and fusion reference frames determined based on the real reference frames and / or the virtual reference frames. As mentioned earlier, in video encoding and decoding tasks, decoded video frames are stored in the DPB (Database Block), and these video frames are called historical reconstructed frames. It should be noted that the construction of the above-mentioned candidate reference frames depends on the historical reconstructed frames that already exist in the DPB. Specifically, it can be applied to video frames encoded using inter-frame prediction (such as P-frames or B-frames). For I-frames encoded using intra-frame prediction, since the DPB is empty or unavailable, the reference frame generation and selection mechanism in this embodiment is not enabled, and the intra-frame prediction process specified in the standard is directly followed.

[0046] To provide a reliable and rich source of reconstruction data for the current frame, thereby improving the reconstruction quality of the current frame, this application provides three types of candidate reference frames: real reference frames, virtual reference frames, and fused reference frames.

[0047] In this embodiment, the true reference frame refers to the true reference frame selected from the cached historical reconstructed frames. The method for obtaining the true reference frame can refer to the aforementioned traditional process of obtaining reference frames through inter-frame prediction. Specifically, the current frame is divided into multiple image blocks. For each image block, the region most similar to the current image block is searched among the multiple historical reconstructed frames cached in the DPB, and the historical reconstructed frame containing that region is used as the true reference frame for that image block.

[0048] For example, for a 16×16 image patch, the video encoder slides a matching window within a preset search range (e.g., horizontal / vertical ±64 pixels) with subpixel precision within each frame of the reference image in the DPB. It determines the similarity between the two based on evaluation metrics such as the sum of absolute pixel differences (SAD) and transform domain cost (SATD) between each window region and the current image patch.

[0049] This allows us to find the window region with the highest similarity to the current image patch in each historical reconstructed frame. Finally, from the window regions corresponding to each historical reconstructed frame, we select the region with the highest similarity. The historical reconstructed frame containing this region is the true reference frame corresponding to that image patch.

[0050] However, considering that the aforementioned sliding window search requires traversing a large number of candidate positions, resulting in high computational complexity, it is prone to matching semantically inconsistent erroneous regions due to local pixel similarity, thus causing accuracy issues. Therefore, this application provides a novel process for obtaining a real reference frame. By introducing the positional information of image blocks in the current frame, the search range of the reference frame is constrained, and the spatiotemporal consistency of the matching is improved.

[0051] In some embodiments, a first image feature of an image patch can be pre-extracted, and then, based on the position information of the image patch in the current frame, a second image feature at the same position in each historical reconstructed frame can be determined. Next, the feature similarity between the first image feature and the corresponding second local image feature of each historical reconstructed frame is calculated. Finally, based on the feature similarity, the true reference frame is selected from each historical reconstructed frame.

[0052] Specifically, it can be as follows Figure 3 As shown, the first image features of the current image patch A are extracted in advance using feature extraction algorithms such as Scale Invariant Feature Transform (SIFT) and Fast Window Rotation (ORB), and then the position information (x, y) of the image patch in the current frame is obtained. For each historical reconstruction frame, image patch B at the same position is obtained. n And extract its second image features.

[0053] Next, the feature similarity between the first image feature and the corresponding second local image feature of each historical reconstructed frame is calculated. Then, the image block B1 with the highest feature similarity is selected. The historical reconstructed frame 1 where image block B1 is located is the real reference frame corresponding to the image block.

[0054] The above process constrains the matching region in historical reference frames by introducing the positional information of the current image patch, eliminating the need for extensive sliding window traversal in each frame and effectively reducing the computational complexity of reference frame search. Furthermore, by performing similarity comparison in the feature domain rather than the original pixel domain, it can distinguish regions with the same semantics but different pixel distributions, avoiding semantic inconsistencies and mismatches caused by local texture duplication or occlusion. This ensures that the selected real reference frames more accurately reflect the content characteristics of the current image patch.

[0055] In this embodiment, the virtual reference frame refers to a virtual reference frame generated based on a trained image generation model. Specifically, from the cached historical reconstructed frames, N frames closest to the current frame can be selected as model input (e.g., the two frames preceding the current frame in the video frame sequence). The trained image generation model then extracts features from the input N frames to generate the virtual reference frame.

[0056] To facilitate understanding, the training process of the image generation model in this application is explained below: A deep neural network-based image generation model is pre-built. The model structure can be set according to actual business needs. For example, it can adopt any of the conventional architectures such as Convolutional Recurrent Neural Network (ConvRNN) or Spatiotemporal Transformer architecture. This application does not limit it in this regard.

[0057] Next, the cleaned historical video frame set (e.g., a complete video frame sequence of a certain video) is used to train the image generation model to be trained in multiple rounds of iteration until the preset iteration stopping condition is met, and the trained image generation model is obtained.

[0058] The process for each iteration is as follows: First, select N consecutive historical video frames from the historical video frame set, where N≥3. Then, select one frame (excluding the first frame) from the N historical video frames as the training label, and use the remaining historical video frames as training samples. In practice, the last frame or any intermediate frame can be selected as the training label; for example, with N=3, this means using the second or third frame out of these three consecutive frames as the training label.

[0059] It should be noted that when the middle frame is used as the training label, it corresponds to the bidirectionally predicted B-frame in inter-frame prediction; when the last frame is used as the training label, it corresponds to the forward predicted P-frame in inter-frame prediction. The specific rules for setting the training labels can be set based on actual business needs, and this application does not limit them.

[0060] Next, features are extracted from the input training samples based on the current model parameters to generate virtual video frames for this round. Then, based on the feature differences between the virtual video frames and the sample labels, the iteration loss for this round is determined. Finally, the current model parameters are corrected based on the iteration loss for this round.

[0061] After each iteration, it is determined whether the current iteration meets the preset iteration convergence condition. This convergence condition may include the iteration loss being less than a threshold, the number of iterations reaching a preset number, etc., which are not limited in this application. If the iteration convergence condition is not met after the current iteration, the current model parameters are adjusted based on the iteration loss of this iteration, and the next iteration begins. If the iteration does not meet the convergence condition, the iteration can be stopped, and the trained image generation model is obtained.

[0062] By employing a trained image generation model, virtual reference frames that are sequentially continuous with the current frame are predicted based on historical reconstructed frames. This introduces virtual reference frames that conform to the motion patterns and scene evolution of the video, effectively expanding the content sources of candidate reference frames.

[0063] Meanwhile, since the virtual reference frame is a reasonable prediction result generated by multi-frame temporal modeling rather than random or static interpolation content, it can effectively reflect local image features that may appear in the current frame but have not been actually captured. Therefore, by using the virtual reference frame as a candidate reference frame, the content diversity of the reconstructed data source can be improved while ensuring data reliability.

[0064] In the embodiments of this application, the fusion reference frame refers to a fusion reference frame determined based on a real reference frame and / or a virtual reference frame. In some embodiments, the fusion reference frame may be generated based on a trained image fusion model.

[0065] To facilitate understanding, the training process of the image fusion model in this application is explained below: A deep neural network-based image fusion model is pre-constructed, and its network structure can be flexibly configured according to actual fusion needs. For example, it can adopt any of the conventional structures such as channel attention fusion network, multi-scale feature weighting network, or Transformer-based cross-modal fusion architecture. This application does not limit this. The model is designed to dynamically adjust the weights of the real reference frame and the virtual reference frame in the fusion process, in addition to utilizing the pixel content of the real reference frame and the virtual reference frame, and combining multi-dimensional clues such as the spatial texture distribution in the current decoding context, the motion continuity of neighboring blocks, and the quantization intensity and residual activity of the current coding unit, thereby generating a reference frame that is more suitable for the reconstruction needs of the current image patch.

[0066] Next, training samples and sample labels are constructed by combining the trained image generation model.

[0067] In practice, M consecutive frames (M≥3) can be selected from the aforementioned set of historical video frames. Then, one frame (excluding the first frame) is selected from the M historical video frames as the training label, and the remaining historical video frames are used as training samples. Specifically, the last frame or any intermediate frame can be chosen as the training label; for example, M=3 means the second or third frame of these three consecutive frames is used as the training label.

[0068] It should be noted that when the middle frame is used as the training label, it corresponds to the bidirectional prediction B-frame in inter-frame prediction; when the last frame is used as the training label, it corresponds to the forward prediction P-frame in inter-frame prediction. The specific rules for setting the training labels can be set based on actual business needs, and this application does not limit them.

[0069] Next, the selected training samples are input into the trained image generation model to generate corresponding virtual reference frames. Simultaneously, using the aforementioned method for obtaining real reference frames, real reference frames are selected from the historical reconstructed frames corresponding to the remaining frames.

[0070] When constructing training samples, in addition to the input real reference frame and virtual reference frame, the features of the current image patch in the temporal and spatial domains can also be introduced. For example, these may include: texture data of the current image patch and image patches at the same position in the two input frames, block motion vector data of the current image patch in the current frame of the adjacent images, and quantization parameters used by the current image patch.

[0071] Subsequently, feature fusion processing is performed on the input based on the current model parameters to generate the fusion reference frame for this round. Then, based on the feature differences between the fusion reference frame and the sample labels, the iteration loss for this round is determined. Finally, the current model parameters are corrected based on the iteration loss for this round.

[0072] After each iteration, it is determined whether the current iteration meets the preset iteration convergence condition. This convergence condition may include the iteration loss being less than a threshold, the number of iterations reaching a preset number, etc., which are not limited in this application. If the iteration convergence condition is not met after the current iteration, the current model parameters are adjusted based on the iteration loss of this iteration, and the next iteration begins. If the iteration does not meet the convergence condition, the iteration can be stopped, and the trained image generation model is obtained.

[0073] In practical encoding and decoding applications, this image fusion model can be pre-integrated into the video decoder without transmitting any model parameters in the bitstream. When the video encoder decides to use a fusion reference frame for a certain image block, it only needs to insert a lightweight syntax element into the bitstream to instruct the video decoder to enable fusion mode and select the corresponding reference frame combination. Based on this syntax signal, the video decoder calls the local model, combines the real reference frames in the DPB, the virtual reference frames synthesized in real time through the image generation model, and the coding context features of the current image block to generate the fusion reference frame.

[0074] After constructing corresponding candidate reference frames for each image block in the current frame through the above process, the target reference frame for image reconstruction can be selected from each candidate reference frame through step S22.

[0075] Step S22: For each image block, based on the multidimensional content features of the image block, select a target reference frame from the plurality of candidate reference frames to reconstruct the image block; To select the reference source that best balances compression efficiency and reconstruction quality from a variety of candidate reference frames, this application adopts Rate-Distortion Cost (RDO) as a unified evaluation criterion. Rate-Distortion Cost comprehensively reflects the degree of image distortion and the required coding bit overhead when using a reference frame for prediction, and is a core indicator for evaluating the quality of coding decisions in video coding.

[0076] In some embodiments, for each candidate reference frame, motion estimation can be performed on the current image block using the candidate reference frame to obtain motion vector data of the current image block, and the rate-distortion cost between the current image block and the candidate reference frame can be determined based on the motion vector data and the multidimensional content features of the current image block.

[0077] For any candidate reference frame, the video encoder attempts to use it to predict the content of the current image patch. This process, called motion estimation, aims to find the region in the candidate reference frame that is most similar to the current image patch and use the motion vector data of the current image patch to describe the positional offset of this region relative to the current image patch.

[0078] After obtaining the motion vector data of the current image block, a comprehensive score, namely the rate-distortion cost, can be calculated by combining the multi-dimensional content features of the current image block (such as texture complexity, neighborhood motion consistency, quantization parameters, etc.), thereby providing a basis for the selection of the target reference frame.

[0079] In this way, by using rate-distortion cost as a unified evaluation criterion, objective and quantifiable comparisons can be made among multiple candidate reference frames, thereby selecting the target reference frame that achieves the optimal balance between reconstruction quality and coding overhead. This approach avoids suboptimal selections that may result from relying on a single metric (such as only minimum residual or closest temporal distance), effectively improving the rationality of reference frame selection and encoding / decoding efficiency.

[0080] In practice, the prediction block corresponding to the current image block in the candidate reference frame can be determined based on the motion vector data of the current image block. Then, the residual data of the current image block is determined based on the pixel difference between the current image block and the prediction block. Finally, the rate-distortion cost between the current image block and the candidate reference frame is determined based on the multidimensional content features and residual data of the current image block.

[0081] To facilitate understanding, the specific calculation process for the rate-distortion cost described above will be explained below: First, motion vector data of the current image patch relative to the candidate reference frame is obtained through motion estimation operations. This motion vector data indicates the best matching position of the current image patch in the reference frame. Then, the prediction patch is determined from the candidate reference frame based on the coordinate offset indicated by the motion vector data.

[0082] Next, the difference between each pixel value of the current image block and the corresponding pixel value in the prediction block is calculated to obtain a residual block. The residual block is then transformed and quantized to obtain quantized transform coefficients. Based on these transform coefficients, the number of bits required to encode the residual block is estimated, yielding the code rate overhead R. Further, the quantized transform coefficients are dequantized and inverse transformed to obtain a reconstructed residual block. This reconstructed residual block is then pixel-by-pixel added to the prediction block to obtain the reconstructed block. Finally, the sum of squared errors (SSE) between the reconstructed block and the current image block is calculated as the reconstruction distortion D.

[0083] Subsequently, multidimensional content features T of the current image patch are extracted. These multidimensional content features can be set based on actual business needs, such as texture complexity in the spatial dimension, neighboring block motion consistency in the temporal dimension, and quantization parameters QP in the coding dimension. This application does not limit these features. Then, the multidimensional content features T are used to adaptively weight or correct the distortion term D and the bitrate term R. The corrected distortion term D' and bitrate R' are substituted into the rate-distortion cost formula J=D'+λ·R' to obtain the rate-distortion cost J corresponding to the current candidate reference frame. Here, λ is a Lagrange multiplier, the value of which is determined by the current quantization parameter QP and is used to balance the relative weights of distortion and bitrate.

[0084] After determining the rate-distortion cost of each candidate reference frame, the candidate reference frame with the lowest rate-distortion cost can be used as the target reference frame for the current image patch. Since the rate-distortion cost calculation in this embodiment is based on traditional distortion and bitrate calculations, and further incorporates the multi-dimensional content features of the current image patch for adaptive correction, the target reference frame selected in this way has stronger reliability in terms of content matching, structural integrity, and temporal consistency. This provides a more accurate prediction basis for the current image patch, thereby effectively improving the detail fidelity and overall visual quality of the reconstructed image.

[0085] Furthermore, considering that directly fusing real and virtual reference frames to generate a fused reference frame might introduce new interference data when there are significant differences in image quality between the two, for example, when the virtual reference frame is distorted due to complex motion or occlusion, while the real reference frame remains highly reliable, directly generating a fused reference frame based on the two might cause the fusion result to deviate from the true content, generating low-quality candidate reference frames and thus reducing the reconstruction quality.

[0086] Based on this, the embodiments of this application also provide another method for generating fused reference frames. In some embodiments, after determining the real reference frame and virtual reference frame corresponding to the current image block, the reconstruction quality difference between the two can be judged. The reconstruction quality difference can be evaluated by various loss metrics, such as pixel-level mean square error (MSE), perceptual loss, and rate distortion cost mentioned above.

[0087] Considering that the core optimization goal of video encoding and decoding tasks is to maximize reconstruction quality under a limited bit rate, and that rate-distortion cost can comprehensively reflect reconstruction distortion and coding bit overhead, and is directly related to coding efficiency and reconstruction performance, therefore, as a preferred embodiment of this application, rate-distortion cost is still used as an evaluation index to measure the adaptability of real reference frames and virtual reference frames to the reconstruction of the current image block.

[0088] In some embodiments, if the rate-distortion cost difference between the real reference frame and the virtual reference frame is less than a preset threshold, a trained image fusion model is used to fuse the features of the real reference frame and the virtual reference frame to obtain a fused reference frame. That is, the obtained real reference frame and virtual reference frame are input into the trained image fusion model, and the corresponding fused reference frame is output.

[0089] If the rate-distortion cost difference between the real reference frame and the virtual reference frame is not less than a preset threshold, it indicates that there is a significant difference in the reconstruction effect of the real reference frame and the virtual reference frame on the current image patch. For example, one of them may contain obvious distortion or structural deviations due to limitations such as motion estimation errors, occlusion, or generative models. If it is fused with another frame with high reliability, it will inject erroneous information into the fusion result and impair the overall reliability of the fused reference frame.

[0090] At this point, a fused reference frame can be obtained by performing image enhancement processing on a specified reference frame. This specified reference frame is the one with the lower rate-distortion cost between the real reference frame and the virtual reference frame.

[0091] The preset image enhancement processing may include at least one of adaptive filtering, local contrast enhancement, edge sharpening, or deblocking suppression. The specific selection rules can be set based on actual business needs, and this application does not limit them.

[0092] In the above process, the optimal fusion strategy is dynamically selected based on the actual reconstruction performance of the current image patch by the real and virtual reference frames (measured by rate-distortion cost). When the difference between the two is low, the image fusion model complements the data of the real and virtual reference frames to enhance the richness of the reference content. When the difference between the two is significant, to avoid the errors that may be introduced by blind fusion, targeted image enhancement is performed on a more reliable single frame to improve detail while preserving its structural correctness.

[0093] This mechanism effectively balances fusion gain and data reliability, preventing low-quality data from contaminating high-quality reference sources and avoiding missing performance improvement opportunities in complementary scenarios, thereby improving the overall quality stability and reconstruction adaptation capability of the fused reference frames.

[0094] Step S23: Add the frame type identifier of the target reference frame corresponding to each image block to the bitstream.

[0095] After determining the target reference frame for each image block, its corresponding frame type identifier can be added to the bitstream. This frame type identifier can be represented using binary codewords. For example, "0" represents the real reference frame, "1" represents the virtual reference frame, and "01" or "10" represents the fused reference frame. Alternatively, a multi-valued symbol system (such as enumerating values ​​"0, 1, 2" to correspond to the real reference frame, virtual reference frame, and fused reference frame, respectively) can be used. The representation method can be set according to actual business needs, and this application does not limit it in this regard.

[0096] Furthermore, as mentioned above, the embodiments of this application can select the appropriate generation method for the fused reference frame based on the rate-distortion cost between the current image patch and the real reference frame and the virtual reference frame to generate the fused reference frame. Considering practical applications, to reduce bitstream overhead, image patches are usually not added to the bitstream. That is, the video decoder cannot obtain the image patch, nor can it calculate the rate-distortion cost between the image patch and the real reference frame and the virtual reference frame.

[0097] Based on this, when determining the target reference frame as the fusion reference frame, a fusion identifier indicating the generation method of the fusion reference frame can be added to the bitstream so that the video decoder can select to use the trained image fusion model to perform feature fusion on the real reference frame and the virtual reference frame according to the identifier, so as to obtain the fusion reference frame; or, image enhancement processing can be performed on the one with the smaller rate-distortion cost between the real reference frame and the virtual reference frame to obtain the fusion reference frame.

[0098] It should be understood that, in order to support the video decoder in accurately reconstructing the current frame, other data necessary for decoding must also be encoded into the encoded data of the current frame, including but not limited to: motion vector data of each image block, residual data, prediction mode, quantization parameters, etc. After entropy encoding, the above data together constitute the complete encoded data of the current frame.

[0099] The above process explicitly indicates the target reference frame type used by each image block in the bitstream and synchronously transmits the corresponding motion vector data and residual data, enabling the video decoder to accurately reproduce the reference frame selection and prediction process of the video encoder, thereby ensuring the consistency and high quality of the reconstructed image. At the same time, this encoding method is compatible with any mainstream video coding framework and only requires a few additional syntax elements to support the flexible scheduling of multiple types of reference frames, effectively improving the reconstruction quality of the current frame while reducing bitrate overhead.

[0100] The following is a detailed description of a video decoding method provided in an embodiment of this application. This method can be applied to a video decoder, as detailed below. Figure 4As shown, the method includes the following steps: Step S41: Decode the frame type identifier of each image block in the current frame from the bitstream; wherein, the frame type identifier is the identifier of the target reference frame; the target reference frame is selected from multiple candidate reference frames, the multiple candidate reference frames include: real reference frames selected from each historical reconstructed frame, virtual reference frames generated based on feature extraction of historical reconstructed frames based on the image generation model, and fused reference frames determined based on the real reference frames and / or the virtual reference frames; As mentioned earlier, to support the video decoder in accurately reconstructing the current frame, the video encoder has written context information such as motion vector data, residual data, prediction mode, and quantization parameters corresponding to each image block into the bitstream. In addition to the frame type identifier, the video decoder also needs to synchronously decode the above information from the current frame's bitstream to reconstruct the current frame.

[0101] Step S42: For each image block, obtain the target reference frame indicated by the frame type identifier, and reconstruct the image block based on the target reference frame to obtain the reconstructed block corresponding to the image block; The video decoder in this embodiment also has the ability to construct candidate reference frames. The method of constructing candidate reference frames is the same as that of the video encoder described above, and will not be repeated here.

[0102] After decoding the encoded data of the current frame to obtain the frame type identifier of each image block, the video decoder can use the reference frame type indicated by the frame type identifier to generate the corresponding target reference frame.

[0103] For example, if the frame type identifier corresponding to the current image block is "0", it means that its corresponding target reference frame is a real reference frame. In this case, the video decoder can use the aforementioned process of the video encoder to generate a real reference frame for the image block as the target reference frame.

[0104] After obtaining the target reference frame, the video decoder locates and extracts a prediction block from the target reference frame that matches the spatial position and size of the current image patch, based on the decoded motion vector data. Then, based on the decoded residual data, quantization parameters, and prediction mode, standard inverse quantization and inverse transform operations are performed to restore the residual data from the compressed domain to the original pixel-domain residual representation. Finally, the prediction block and the pixel-domain residual representation are added pixel by pixel to generate the reconstructed block corresponding to the current image patch.

[0105] Furthermore, as mentioned above, when the video encoder in this embodiment determines that the target reference frame is a fusion reference frame, it adds a fusion identifier, which indicates the generation method of the fusion reference frame, to the bitstream.

[0106] Therefore, when the video decoder determines that the current target reference frame is a fusion reference frame, it can further obtain the fusion identifier from the bitstream and obtain the fusion reference frame according to the method indicated by the identifier (i.e., using a trained image fusion model to perform feature fusion on the real reference frame and the virtual reference frame, or by performing image enhancement processing on the one with the smaller rate-distortion cost between the real reference frame and the virtual reference frame).

[0107] It should be noted that if the frame type identifier indicates that the target reference frame is a virtual reference frame or a fused reference frame, the video decoder must first call the locally preset image generation model or image fusion model, and use the historical reconstructed frames already stored in the DPB to generate the corresponding virtual reference frame or fused reference frame according to a process completely consistent with the video encoder. Since the model structure, parameters, and input conditions are all pre-agreed upon at both the encoder and decoder ends, there is no need to transmit the model itself through the bitstream; the video decoder can be instructed to perform the corresponding type of reference frame synthesis solely based on the frame type identifier.

[0108] Step S43: Determine the reconstructed frame of the current frame based on the reconstructed block corresponding to each of the image blocks.

[0109] After obtaining the reconstructed blocks of all image blocks in the current frame, the reconstructed frame corresponding to the current frame is reconstructed by stitching and combining the image blocks according to their position information in the current frame. This reconstructed frame will serve as the historical reconstructed frame in subsequent encoding and decoding processes.

[0110] The above process can accurately reproduce the multi-type reference frame mechanism used by the video encoder based on the frame type identifier carried in the bitstream, without relying on the original image content, and completely reconstruct each image block of the current frame. Since the real reference frame, virtual reference frame, and fused reference frame are all dynamically generated by the video decoder based on the same model and historical reconstructed frames as the video encoder, there is no need to transmit the reference frame itself or model parameters. Moreover, by explicitly signaling the rate-distortion optimization decision results with lightweight syntax elements (such as frame type identifiers), the accurate identification of the target reference frame type and the correct execution of the reconstruction path by the video decoder are ensured, thereby improving reconstruction quality while reducing bitrate overhead.

[0111] Based on the same inventive concept, this application also provides a video encoding device 500, the structure of which is as follows: Figure 5 As shown, it includes: The reference frame acquisition unit 501 is configured to: determine multiple candidate reference frames corresponding to each image block in the current frame based on each historical reconstruction frame; wherein, the multiple candidate reference frames include: real reference frames selected from each historical reconstruction frame, virtual reference frames generated based on feature extraction of historical reconstruction frames by an image generation model, and fused reference frames determined based on the real reference frames and / or the virtual reference frames; The reference frame confirmation unit 502 is configured to: for each image block, select a target reference frame for reconstructing the image block from the plurality of candidate reference frames based on the multidimensional content features of the image block; The encoding unit 503 is configured to add a frame type identifier of the target reference frame corresponding to each image block to the bitstream.

[0112] In some embodiments, the multidimensional content features based on the image patch are executed to select a target reference frame for reconstructing the image patch from the plurality of candidate reference frames. The reference frame confirmation unit 502 is specifically configured to: For each candidate reference frame, motion estimation is performed on the image block using the candidate reference frame to obtain motion vector data of the image block, and the rate-distortion cost between the image block and the candidate reference frame is determined based on the motion vector data and the multidimensional content features of the image block. The candidate reference frame with the lowest rate-distortion cost is used as the target reference frame for the image block.

[0113] In some embodiments, the multidimensional content features based on the motion vector data and the image patch are used to determine the rate-distortion cost between the image patch and the candidate reference frame. Specifically, the reference frame confirmation unit 502 is configured to: Based on the motion vector data, the prediction block corresponding to the image block in the candidate reference frame is determined; Based on the pixel difference between the image block and the prediction block, residual data is determined; Based on the multidimensional content features and the residual data, the rate-distortion cost between the image patch and the candidate reference frame is determined.

[0114] In some embodiments, after performing the step of selecting a target reference frame for reconstructing the image patch from the plurality of candidate reference frames, the reference frame confirmation unit 502 is further configured to: If the target reconstructed frame is the fusion reference frame, then a fusion identifier indicating the generation method of the fusion reference frame is added to the bitstream; The fusion reference frame is determined in the following way: If the difference between the rate-distortion cost between the image patch and the real reference frame and the rate-distortion cost between the image patch and the virtual reference frame is less than a preset threshold, then the trained image fusion model is used to perform feature fusion on the real reference frame and the virtual reference frame to obtain the fused reference frame. If the difference is not less than the preset threshold, then image enhancement processing is performed on the specified reference frame to obtain the fused reference frame; the specified reference frame is the one with the smaller rate-distortion cost between the real reference frame and the virtual reference frame.

[0115] In some embodiments, the true reference frame is determined in the following manner: Extract the first image features of the image block; Based on the position data of the image block in the current frame, determine the second image features of each historical reconstructed frame at the same position; Determine the feature similarity between the first image feature and the corresponding second image feature of each of the historical reconstructed frames; Based on the feature similarity, the real reference frame is selected from each of the historical reconstructed frames.

[0116] In some embodiments, the image generation model is obtained through multiple rounds of iterative training based on a set of historical video frames; wherein each iteration is as follows: Select N consecutive historical video frames from the set of historical video frames; where N≥3; From the N historical video frames, select one frame as the training label and use the remaining historical video frames (excluding the first frame) as training samples. Based on the current model parameters, feature extraction is performed on the training samples to generate virtual video frames for this round; The iteration loss for this round is determined based on the feature differences between the virtual video frame and the sample label; The current model parameters are adjusted based on the loss from this iteration.

[0117] Based on the same inventive concept, this application also provides a video decoding device 600, the structure of which is as follows: Figure 6 As shown, it includes: Decoding unit 601 is configured to: decode from the bitstream to obtain the frame type identifier of each image block in the current frame; wherein the frame type identifier is the identifier of the target reference frame; the target reference frame is selected from multiple candidate reference frames, the multiple candidate reference frames including: real reference frames selected from each historical reconstructed frame, virtual reference frames generated based on feature extraction of historical reconstructed frames based on an image generation model, and fused reference frames determined based on the real reference frames and / or the virtual reference frames; The reconstruction block unit 602 is configured to: for each image block, obtain the target reference frame indicated by the frame type identifier, and reconstruct the image block based on the target reference frame to obtain the reconstruction block corresponding to the image block; The reconstruction frame unit 603 is configured to determine the reconstruction frame of the current frame based on the reconstruction block corresponding to each of the image blocks.

[0118] In some embodiments, the target reference frame is determined in the following manner: For each candidate reference frame, motion estimation is performed on the image block using the candidate reference frame to obtain motion vector data of the image block, and the rate-distortion cost between the image block and the candidate reference frame is determined based on the motion vector data and the multidimensional content features of the image block. The candidate reference frame with the lowest rate-distortion cost is used as the target reference frame for the image block.

[0119] In some embodiments, the decoding unit 601 is specifically configured to perform the multidimensional content features based on the motion vector data and the image patch to determine the rate-distortion cost between the image patch and the candidate reference frame, wherein the decoding unit 601 is specifically configured to: Based on the motion vector data, the prediction block corresponding to the image block in the candidate reference frame is determined; Based on the pixel difference between the image block and the prediction block, residual data is determined; Based on the multidimensional content features and the residual data, the rate-distortion cost between the image patch and the candidate reference frame is determined.

[0120] In some embodiments, the reconstruction block unit 602 is further configured to: If the target reference frame is the fusion reference frame, then the fusion identifier is decoded from the bitstream; The fusion reference frame is generated using the generation method indicated by the fusion identifier. The fusion reference frame is determined in the following way: If the difference between the rate-distortion cost between the image patch and the real reference frame and the rate-distortion cost between the image patch and the virtual reference frame is less than a preset threshold, then the trained image fusion model is used to perform feature fusion on the real reference frame and the virtual reference frame to obtain the fused reference frame. If the difference is not less than the preset threshold, then image enhancement processing is performed on the specified reference frame to obtain the fused reference frame; the specified reference frame is the one with the smaller rate-distortion cost between the real reference frame and the virtual reference frame.

[0121] In some embodiments, the true reference frame is determined in the following manner: Extract the first image features of the image block; Based on the position data of the image block in the current frame, determine the second image features of each historical reconstructed frame at the same position; Determine the feature similarity between the first image feature and the corresponding second image feature of each of the historical reconstructed frames; Based on the feature similarity, the real reference frame is selected from each of the historical reconstructed frames.

[0122] In some embodiments, the image generation model is obtained through multiple rounds of iterative training based on a set of historical video frames; wherein each iteration is as follows: Select N consecutive historical video frames from the set of historical video frames; where N≥3; From the N historical video frames, select one frame as the training label and use the remaining historical video frames (excluding the first frame) as training samples. Based on the current model parameters, feature extraction is performed on the training samples to generate virtual video frames for this round; The iteration loss for this round is determined based on the feature differences between the virtual video frame and the sample label; The current model parameters are adjusted based on the loss from this iteration.

[0123] Based on the same inventive concept, this application also provides a video encoder, which may include: a memory and a processor, wherein the memory stores a computer program, and the processor is configured to implement the video encoding method provided in any of the above embodiments when the computer program is configured to do so.

[0124] Based on the same inventive concept, this application also provides a video decoder, including: a memory and a processor, wherein the memory stores a computer program, and the processor is configured to implement the video decoding method provided in any of the above embodiments when the computer program is configured to do so.

[0125] Figure 7 Exemplary block diagrams of video encoders or video decoders from some of the embodiments described above are shown. Figure 7 As shown, the video encoder or video decoder 700 may include at least one of the following: a tuner / demodulator 701, a communicator 702, a detector 703, an external device interface 704, a processor 705, a display 706, an audio output interface 707, a memory 708, a power supply 709, and a user interface 710.

[0126] In some embodiments, processor 705 includes at least one of: a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), RAM (RandomAccess Memory), ROM (Read-Only Memory), a first to an nth interface for input / output, a communication bus, etc.

[0127] The display 706 includes a display screen assembly for presenting images, a driving assembly for driving image display, a component for receiving image signals from the processor output, and a user interface for displaying video content, image content, menu control interface, and user control UI.

[0128] The display 706 can be a liquid crystal display, an organic light-emitting diode (OLED) display, or a projection display, etc.

[0129] The communicator 702 is a component used to communicate with external devices or servers according to various communication protocol types. For example, the communicator may include at least one of the following: a Wi-Fi module, a Bluetooth module, a wired Ethernet module, other network communication protocol chips or near-field communication protocol chips, and an infrared receiver. The video encoder or video decoder 700 can use the communicator 702 to send and receive control signals and data signals with the control device or server.

[0130] The user interface can be used to receive control signals input by the user through a control device (such as an infrared remote control) or by touch or gesture.

[0131] Detector 703 can be used to acquire signals from the external environment or to interact with the external environment. For example, detector 703 includes a light receiver, which can be used to acquire ambient light intensity; or, detector 703 includes an image acquisition device, such as a camera, which can be used to acquire external environmental scenes, user attributes, or user interaction gestures; or, detector 703 includes a sound acquisition device, such as a microphone, for receiving external sounds.

[0132] External device interface 704 includes, but is not limited to, one or more interfaces such as: High-Definition Multimedia Interface (HDMI), analog or high-definition component input interface (component), Composite Video Broadcast Signal (CVBS), Universal Serial Bus (USB), RGB (Red, Green, Blue) port, etc. It can also be a composite input / output interface formed by multiple interfaces mentioned above.

[0133] The tuner / demodulator 701 receives broadcast television signals via wired or wireless means, and demodulates audio and video signals, such as EPG data signals, from multiple wireless or wired broadcast television signals. In some embodiments, the processor 705 and the tuner / demodulator 701 may be located in different separate devices, that is, the tuner / demodulator 701 may also be located in an external device of the main device where the processor 705 is located, such as an external set-top box.

[0134] Processor 705 controls the operation of electronic devices and responds to user operations through various software control programs stored in memory. Processor 705 controls the overall operation of video encoder or video decoder 700. For example, in response to receiving a user command to select a UI object to display on display 706, processor 705 can perform operations related to the object selected by the user command.

[0135] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a computing device, the computing device implements the video encoding method or the video decoding method described in any of the above embodiments.

[0136] This application also provides a computer program product that, when run on a computer, enables the computer to implement the video encoding method or video decoding method described in any of the above embodiments.

[0137] Those skilled in the art will understand that all or part of the steps of the foregoing method embodiments can be implemented by a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed, it performs the steps of the foregoing method embodiments. The storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0138] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application can essentially be embodied in the form of a software product, for example, a computer program product stored in a storage medium, including a computer program used to cause a computer device to execute all or part of the methods described in the embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

Claims

1. A video encoding method, characterized in that, Applied to a video encoder, the method includes: Based on each historical reconstructed frame, multiple candidate reference frames are determined for each image block in the current frame; wherein, the multiple candidate reference frames include: real reference frames selected from each historical reconstructed frame, virtual reference frames generated by extracting features from historical reconstructed frames based on an image generation model, and fused reference frames determined based on the real reference frames and / or the virtual reference frames. For each image block, a target reference frame for reconstructing the image block is selected from the plurality of candidate reference frames based on the multidimensional content features of the image block; Add a frame type identifier for the target reference frame corresponding to each image block to the bitstream.

2. The method according to claim 1, characterized in that, The step of selecting a target reference frame for reconstructing the image patch from the plurality of candidate reference frames based on the multidimensional content features of the image patch includes: For each candidate reference frame, motion estimation is performed on the image block using the candidate reference frame to obtain motion vector data of the image block, and the rate-distortion cost between the image block and the candidate reference frame is determined based on the motion vector data and the multidimensional content features of the image block. The candidate reference frame with the lowest rate-distortion cost is used as the target reference frame for the image block.

3. The method according to claim 2, characterized in that, The determination of the rate-distortion cost between the image patch and the candidate reference frame based on the motion vector data and the multidimensional content features of the image patch includes: Based on the motion vector data, the prediction block corresponding to the image block in the candidate reference frame is determined; Based on the pixel difference between the image block and the prediction block, residual data is determined; Based on the multidimensional content features and the residual data, the rate-distortion cost between the image patch and the candidate reference frame is determined.

4. The method according to claim 2, characterized in that, After selecting a target reference frame for reconstructing the image patch from the plurality of candidate reference frames, the method further includes: If the target reconstructed frame is the fusion reference frame, then a fusion identifier indicating the generation method of the fusion reference frame is added to the bitstream; The fusion reference frame is determined in the following way: If the difference between the rate-distortion cost between the image patch and the real reference frame and the rate-distortion cost between the image patch and the virtual reference frame is less than a preset threshold, then the trained image fusion model is used to perform feature fusion on the real reference frame and the virtual reference frame to obtain the fused reference frame. If the difference is not less than the preset threshold, then image enhancement processing is performed on the specified reference frame to obtain the fused reference frame; the specified reference frame is the one with the smaller rate-distortion cost between the real reference frame and the virtual reference frame.

5. The method according to any one of claims 1-4, characterized in that, The real reference frame is determined in the following way: Extract the first image features of the image block; Based on the position information of the image block in the current frame, determine the second image features of each historical reconstructed frame at the same position; Determine the feature similarity between the first image feature and the corresponding second image feature of each of the historical reconstructed frames; Based on the feature similarity, the real reference frame is selected from each of the historical reconstructed frames.

6. The method according to any one of claims 1-4, characterized in that, The image generation model is obtained through multiple rounds of iterative training based on a set of historical video frames; each iteration is as follows: Select N consecutive historical video frames from the set of historical video frames; where N≥3; From the N historical video frames, select one frame as the training label and use the remaining historical video frames (excluding the first frame) as training samples. Based on the current model parameters, feature extraction is performed on the training samples to generate virtual video frames for this round; The iteration loss for this round is determined based on the feature differences between the virtual video frame and the sample label; The current model parameters are adjusted based on the loss from this iteration.

7. A video decoding method, characterized in that, Applied to a video decoder, the method includes: The frame type identifier of each image block in the current frame is obtained by decoding from the bitstream; wherein, the frame type identifier is the identifier of the target reference frame; the target reference frame is selected from multiple candidate reference frames, the multiple candidate reference frames include: real reference frames selected from each historical reconstructed frame, virtual reference frames generated by extracting features from historical reconstructed frames based on the image generation model, and fused reference frames determined based on the real reference frames and / or the virtual reference frames; For each image block, obtain the target reference frame indicated by the frame type identifier, and reconstruct the image block based on the target reference frame to obtain the reconstructed block corresponding to the image block; The reconstructed frame of the current frame is determined based on the reconstructed block corresponding to each of the image blocks.

8. A video encoder, characterized in that, include: Memory, used to store computer programs; A processor, configured to execute the video encoding method as described in any one of claims 1-6 when running the computer program.

9. A video decoder, characterized in that, include: Memory, used to store computer programs; A processor, configured to execute the video decoding method as described in claim 7 when running the computer program.

10. A computer-readable storage medium storing a computer program / instructions and a bit stream thereon, characterized in that, When the computer program / instructions are executed by the processor, they are able to generate the bit stream according to any one of claims 1-6 or 7.