Video encoding method, video decoding method, and related apparatus

CN122824901APending Publication Date: 2026-09-25ZHONGXING INTELLIGENT SYST TECH CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610790711.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-02
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

在图像内容复杂度较高的情况下,不同图像块之间的局部特性差异较大,若对所有图像块均采用相同的残差变换方式进行处理,则可能导致残差数据冗余度较高,这不仅会增大码率开销,还会影响当前帧的重建质量

Benefits of technology

[0018]本申请实施例中,针对当前帧中各图像块,分别采用多种图像预测方式,生成各图像块对应的多个预测块。进一步地,针对每个预测块,采用多种残差变换方式对该预测块对应图像块与该预测块之间的像素差值进行量化处理,得到多个残差数据。上述设置通过对多种图像预测方式和残差变换方式进行组合适配,能够得到图像块在多种预测与量化组合下的残差表示,进而丰富编码决策的候选空间。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824901A_ABST
    Figure CN122824901A_ABST
Patent Text Reader

Abstract

The application provides a video encoding method, a video decoding method and related devices, which are used for improving the reconstruction quality of a current frame and reducing the code rate overhead. The method comprises the following steps: for each image block in the current frame, a plurality of prediction blocks corresponding to each image block are generated by using a plurality of image prediction modes respectively; for each prediction block, a plurality of residual data corresponding to the prediction block are obtained by using a plurality of residual transformation modes to quantize the pixel difference between the image block corresponding to the prediction block and the prediction block respectively; based on the residual data corresponding to each prediction block, a target prediction mode and a target transformation mode corresponding to each image block are selected from the plurality of image prediction modes and the plurality of residual transformation modes respectively; and a first identifier of the target prediction mode, a second identifier of the target transformation mode and the target residual data are added to a code stream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video encoding and decoding technology, and in particular to a video encoding method, a video decoding method, and related apparatus. Background Technology

[0002] To support fast start-up and smooth playback of video streams, video encoding and decoding tasks often generate prediction blocks by predicting each image block in the current frame. Subsequently, the pixel difference between the image block and the prediction block is quantized to obtain the corresponding residual data, which is added to the bitstream. This residual data is then used to reconstruct the current frame during the decoding stage, thereby reducing bitrate overhead.

[0003] Currently, most common encoding and decoding standards (such as H.264 / AVC, H.266 / VVC, etc.) use fixed residual transformation methods (such as Discrete Cosine Transform DCT, Discrete Sine Transform DST, etc.) to obtain residual data.

[0004] These standardized residual transformation methods all generate residual data by converting pixel-domain data to the frequency domain. When the image content is highly complex, the local characteristics of different image blocks vary greatly. If the same residual transformation method is used to process all image blocks, it may lead to high redundancy of residual data. This will not only increase the bit rate overhead, but also affect the reconstruction quality of the current frame.

[0005] Therefore, improving the reconstruction quality of the current frame and reducing bitrate overhead are urgent problems that need to be solved. Summary of the Invention

[0006] This application provides a video encoding method, a video decoding method, and related apparatus for improving the reconstruction quality of the current frame and reducing bit rate overhead.

[0007] In a first aspect, embodiments of this application provide a video encoding method, the method comprising: For each image block in the current frame, multiple image prediction methods are used to generate multiple prediction blocks corresponding to each image block; For each prediction block, multiple residual transformation methods are used to quantize the pixel difference between the image block corresponding to the prediction block and the prediction block to obtain multiple residual data corresponding to the prediction block. Based on the residual data corresponding to each prediction block, the target prediction method and target transformation method corresponding to the image block are selected from the multiple image prediction methods and the multiple residual transformation methods, respectively. The first identifier of the target prediction method, the second identifier of the target transformation method, and the target residual data are added to the bitstream; wherein, the target residual data is obtained by quantizing the pixel difference between the image block and the prediction block generated by the target prediction method using the target transformation method.

[0008] Secondly, embodiments of this application provide a video decoding method, the method comprising: Decode the first identifier, second identifier, and target residual data corresponding to each image block in the current frame from the bitstream; For each image patch, the target prediction method corresponding to the first identifier is selected from multiple image prediction methods, and the target transformation method corresponding to the second identifier is selected from multiple residual transformation methods; The target prediction method is used to generate a prediction block, and the target transformation method is used to perform inverse quantization on the target residual data to obtain reconstructed residual data; Based on the prediction block corresponding to each image block and the reconstruction residual data, the reconstruction block corresponding to each image block is obtained, and the reconstruction frame corresponding to the current frame is determined based on the reconstruction block.

[0009] In some embodiments, the multiple residual transformation methods include: a first transformation method that transforms pixel-domain data to the frequency domain, and a second transformation method that performs feature encoding on pixel-domain data based on a residual recognition model; the residual recognition model is obtained through multiple rounds of iterative training based on a set of historical video frames; Each iteration is as follows: Select a training image patch from any historical video frame in the set of historical video frames; Multiple training prediction blocks corresponding to the training image blocks are generated by using the various image prediction methods described above. For the pixel difference between the training image block and each training prediction block, the pixel difference is feature-encoded using the current model parameters to obtain the first residual feature of this round; and the first residual feature is feature-decoded using the current model parameters to obtain the training residual data of this round of reconstruction. Based on the training residual data and the pixel difference, determine the iteration loss for this round; The current model parameters are adjusted based on the loss from this iteration.

[0010] In some embodiments, the multiple image prediction methods include: inter-frame prediction and intra-frame prediction; the step of using the current model parameters to perform feature decoding on the first residual features to obtain the training residual data for this round of reconstruction includes: From the reference frames corresponding to the historical video frames, a first reference image block with the same position as the training image block is selected, and from the historical video frames where the training image block is located, at least one second reference image block adjacent to the training image block is selected; wherein, the reference frames are determined based on the encoded data of the historical video frame set; If the training prediction block is obtained using the inter-frame prediction method, then based on the pixel difference between the first reference image block and the training prediction block, and the pixel difference between the second reference image block and the training prediction block, feature fusion is performed on the first residual feature of the training prediction block to obtain a first feature fusion result; feature decoding is performed on the first feature fusion result to obtain the training residual data. If the training prediction block is obtained using the intra-frame prediction method, then based on the pixel difference between the second reference image block and the training prediction block, the first residual feature of the training prediction block is fused to obtain the second feature fusion result; the second feature fusion result is then decoded to obtain the training residual data.

[0011] In some embodiments, before decoding the first identifier, the second identifier, and the target residual data corresponding to each image block in the current frame from the bitstream, the method further includes: It was determined that the target identifier was not decoded in the bitstream; The method further includes: If the target identifier is decoded from the bitstream, a prediction block is generated using the target prediction method corresponding to the first identifier, and the prediction block is used as the reconstruction block of the image block; wherein, the target identifier is added to the bitstream along with the first identifier by the video encoder when it determines that the target residual data meets a preset residual threshold.

[0012] Thirdly, embodiments of this application provide a video encoding apparatus, the apparatus comprising: The prediction unit is configured to generate multiple prediction blocks corresponding to each image block by using multiple image prediction methods for each image block in the current frame. The quantization unit is configured to: for each prediction block, use multiple residual transformation methods to quantize the pixel difference between the image block corresponding to the prediction block and the prediction block to obtain multiple residual data corresponding to the prediction block; The confirmation unit is configured to: select the target prediction method and target transformation method corresponding to the image block from the multiple image prediction methods and the multiple residual transformation methods based on the residual data corresponding to each prediction block; The encoding unit is configured to add a first identifier of the target prediction method, a second identifier of the target transformation method, and target residual data to the bitstream; wherein the target residual data is obtained by quantizing the pixel difference between the image block and the prediction block generated by the target prediction method using the target transformation method.

[0013] Fourthly, embodiments of this application provide a video decoding apparatus, the apparatus comprising: The decoding unit is configured to decode the first identifier, the second identifier, and the target residual data corresponding to each image block in the current frame from the bitstream; The identification unit is configured to: for each image block, select the target prediction method corresponding to the first identifier from multiple image prediction methods, and select the target transformation method corresponding to the second identifier from multiple residual transformation methods; The residual unit is configured to: generate a prediction block using the target prediction method, and perform inverse quantization processing on the target residual data using the target transformation method to obtain reconstructed residual data; The reconstruction unit is configured to: obtain the reconstruction block corresponding to each image block based on the prediction block corresponding to each image block and the reconstruction residual data, and determine the reconstruction frame corresponding to the current frame based on the reconstruction blocks.

[0014] Fifthly, embodiments of this application provide a video encoder, including: Memory, used to store computer programs; A processor, configured to execute the video encoding method as described in any of the first aspects when running the computer program.

[0015] Sixthly, embodiments of this application provide a video decoder, including: Memory, used to store computer programs; A processor, configured to execute the video decoding method as described in any of the second aspects when running the computer program.

[0016] In a seventh aspect, embodiments of this application provide a computer-readable storage medium storing a computer program / instructions and a bit stream thereon, wherein the computer program / instructions, when executed by a processor, are capable of generating the bit stream according to the method described in any one of the first or second aspects.

[0017] Eighthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method as described in any one of the first or second aspects.

[0018] In this embodiment, for each image block in the current frame, multiple image prediction methods are used to generate multiple prediction blocks corresponding to each image block. Further, for each prediction block, multiple residual transformation methods are used to quantize the pixel differences between the corresponding image block and the prediction block, resulting in multiple residual data. By appropriately combining multiple image prediction methods and residual transformation methods, the above settings can obtain residual representations of image blocks under various prediction and quantization combinations, thereby enriching the candidate space for coding decisions.

[0019] Next, based on the residual data corresponding to each prediction block, target prediction and target transformation methods are selected from various image prediction and residual transformation methods. Since the residual data generated under different combinations of image prediction and residual transformation methods will differ in terms of energy distribution and compressibility, evaluation indicators can be established based on the residual data corresponding to each prediction block, and then the target prediction and target transformation methods that are more suitable for the image blocks can be selected.

[0020] Finally, the first identifier of the target prediction method, the second identifier of the target transformation method, and the corresponding target residual data are written into the bitstream. By adding lightweight identifiers to the bitstream, the video decoder is instructed to use the image prediction method and residual quantization method associated with the corresponding identifier to generate the reconstructed frame of the current frame. This not only reduces bitrate overhead but also effectively improves the reconstruction quality of the current frame. Attached Figure Description

[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments of this application will be briefly introduced below. Obviously, the drawings introduced below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a schematic diagram of the video encoding and decoding process in related technologies.

[0023] Figure 2 This is an overall flowchart of a video encoding method provided in an embodiment of this application.

[0024] Figure 3 This is a schematic diagram of the training process of the residual recognition model provided in the embodiments of this application.

[0025] Figure 4 This is a schematic diagram illustrating the process of obtaining the target prediction method and target transformation method provided in the embodiments of this application.

[0026] Figure 5This is an overall flowchart of a video decoding method provided in an embodiment of this application.

[0027] Figure 6 This is a structural block diagram of a video encoding device provided in an embodiment of this application.

[0028] Figure 7 This is a structural block diagram of a video decoding device provided in an embodiment of this application.

[0029] Figure 8 This is a structural block diagram of a video decoder or video encoder provided in an embodiment of this application. Detailed Implementation

[0030] The technical solutions in the embodiments of this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " will mean "or", for example, A / B can mean A or B; "and / or" in the text is merely a description of the relationship between related objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "multiple" means two or more.

[0031] In the description of the embodiments of this application, unless otherwise stated, the term "multiple" refers to two or more, and other quantifiers are similarly understood. The preferred embodiments described herein are for illustration and explanation only and are not intended to limit this application. Furthermore, the embodiments and features in the embodiments of this application can be combined with each other without conflict.

[0032] To further illustrate the technical solutions provided in the embodiments of this application, a detailed description is provided below in conjunction with the accompanying drawings and specific implementation methods. Although the embodiments of this application provide method operation steps as shown in the following embodiments or drawings, more or fewer operation steps may be included in the method based on conventional or non-inventive effort. For steps that do not logically have a necessary causal relationship, the execution order of these steps is not limited to the execution order provided in the embodiments of this application. In actual processing or when the control device executes the method, it may be executed sequentially or in parallel according to the method shown in the embodiments or drawings.

[0033] As mentioned earlier, in video encoding and decoding tasks, prediction blocks are often generated by predicting each image block in the current frame. Subsequently, the pixel difference between the image block and the prediction block is quantized to obtain the corresponding residual data, which is added to the bitstream. This residual data is then used to reconstruct the current frame during the decoding stage, thereby reducing bitrate overhead.

[0034] Specifically, in commonly used video codec standards (such as H.264 / AVC, H.266 / VVC, etc.), the workflow of a video encoder is typically as follows: Figure 1 As shown: First, the current frame is divided into multiple image blocks (called coding units (CUs) or prediction units (PUs). For each image block, inter-frame prediction or intra-frame prediction can be used to generate the prediction block corresponding to the current image block. The relevant process for generating prediction blocks using inter-frame prediction can be as follows: Figure 1 As shown, it performs motion estimation on the current image patch, traversing the historical reconstructed frames cached in the Decoded Picture Buffer (DPB) to find the region with the highest image similarity to the current image patch. Then, the historical reconstructed frame containing that region is used as a reference frame for reconstructing the current frame. Next, based on the displacement difference between the image patch and that region, motion vector information (MV) of the image patch is generated. Motion compensation calculations are then performed on the motion vector data of the image patch to generate the corresponding prediction block.

[0035] The process of generating prediction blocks using intra-frame prediction ( Figure 1 (Not shown in the image) This mainly includes obtaining reference pixels from the reconstructed neighboring regions in the current frame by performing an intra-frame prediction mode selection operation on the current image patch. Then, based on the set prediction direction mode (such as DC mode, angle prediction mode, etc.), the prediction block corresponding to the image patch is generated by performing interpolation or copying operations on the reference pixels according to the selected prediction mode.

[0036] After generating the prediction block, the pixel difference between the current image block and the prediction block (also known as the residual block) is calculated. This pixel difference is used to reflect the missing details of the prediction block relative to the image block.

[0037] To improve compression efficiency, residual blocks need to be quantized. Specifically, current video coding standards mostly use Integer Discrete Cosine Transform (DCT) or its approximation to transform the pixel-domain residual data to the frequency domain, obtaining a set of transform coefficients. Then, according to preset quantization parameters (QP), these transform coefficients are quantized, i.e., division and rounding operations are performed with a fixed step size, significantly reducing the number of non-zero coefficients and lowering representation precision. The quantized frequency domain coefficients serve as the final residual data. Finally, the residual data, image prediction method identification information, etc., are entropy-encoded together to form a compressed bitstream, which is then transmitted to the video decoder.

[0038] After receiving the bitstream, the video decoder decodes parameters such as the prediction mode identifier and residual data of the current image block, generates prediction blocks based on these parameters, and reconstructs the current frame based on the generated prediction blocks.

[0039] Continue as Figure 1 As shown, if the image prediction method is inter-frame prediction, the video decoder parses the motion vector data, reference frame identifier, and residual data of the current image block from the received compressed bitstream. Then, based on the reference frame identifier, it retrieves the corresponding historical reconstructed frame from the local DPB, thereby obtaining the reference frame corresponding to the current block. Finally, based on the motion vector data, motion compensation is performed on the pixels pointed to by the motion vector data in the reference frame to generate a prediction block consistent with the video encoder.

[0040] Correspondingly, if the image prediction method is intra-frame prediction, the video decoder will parse the identifier of the prediction direction mode used by the current image block and the residual data from the received compressed bitstream. Subsequently, it will obtain reference pixels from the reconstructed adjacent pixel regions in the current frame, and perform interpolation or copying operations on the reference pixels based on the prediction direction mode indicated by the prediction mode identifier to generate a prediction block consistent with the video encoder.

[0041] After determining the prediction block through the above process, the video decoder dequantizes the residual data parsed from the bitstream to obtain the reconstructed residual data. Then, the prediction block and the reconstructed residual data are added together to obtain the reconstructed block corresponding to the current image block. Finally, all reconstructed image blocks are combined into a complete reconstructed frame and stored in the DPB for use in encoding / decoding subsequent video frames or image display.

[0042] In related technologies, fixed residual transformation methods (such as Discrete Cosine Transform (DCT) and Discrete Sine Transform (DST)) are often used to obtain residual data. These standardized residual transformation methods all generate residual data by converting pixel-domain data to the frequency domain. When the image content is highly complex, the local characteristics of different image blocks vary greatly. If the same residual transformation method is used to process all image blocks, it may lead to high redundancy in the quantized residual data. This will not only increase the bit rate overhead but also affect the reconstruction quality of the current frame.

[0043] In view of this, the inventive concept of this application is as follows: for each image block in the current frame, multiple image prediction methods are used to generate multiple prediction blocks corresponding to each image block. Further, for each prediction block, multiple residual transformation methods are used to quantize the pixel difference between the corresponding image block and the prediction block, resulting in multiple residual data. By appropriately combining multiple image prediction methods and residual transformation methods, the above setup can obtain the residual representation of the image block under multiple prediction and transformation combinations, thereby enriching the candidate space for coding decisions.

[0044] Next, based on the residual data corresponding to each prediction block, target prediction and target transformation methods are selected from various image prediction and residual transformation methods. Since the residual data generated under different combinations of image prediction and residual transformation methods will differ in terms of energy distribution and compressibility, evaluation indicators can be established based on the residual data corresponding to each prediction block, and then the target prediction and target transformation methods that are more suitable for the image blocks can be selected.

[0045] Finally, the first identifier of the target prediction method, the second identifier of the target transformation method, and the corresponding target residual data are written into the bitstream. By adding lightweight identifiers to the bitstream, the video decoder is instructed to use the image prediction method and residual quantization method associated with the corresponding identifier to generate the reconstructed frame of the current frame. This not only reduces bitrate overhead but also effectively improves the reconstruction quality of the current frame.

[0046] After introducing the inventive concept of the embodiments of this application, the following is a detailed description of a video encoding method provided by the embodiments of this application. This method can be applied to a video encoder, specifically as follows: Figure 2 As shown, the method includes the following steps: Step S21: For each image block in the current frame, multiple image prediction methods are used to generate multiple prediction blocks corresponding to each image block; As mentioned earlier, many related technologies use fixed image prediction methods to predict each image patch and generate corresponding prediction blocks. However, due to the high diversity of video content in terms of texture structure, motion characteristics, and local correlations, fixed prediction methods cannot be dynamically adjusted according to the specific characteristics of image patches. This may result in insufficient prediction accuracy, high residual energy, and consequently, increased bitrate overhead or decreased reconstruction quality.

[0047] Based on this, the embodiments of this application employ inter-frame prediction and intra-frame prediction methods for each image block, generating prediction blocks corresponding to that image block under each prediction method. That is, the various image prediction methods in the embodiments of this application include inter-frame prediction and intra-frame prediction.

[0048] The relevant procedures for generating image patches using inter-frame prediction and intra-frame prediction methods have been described above. Figure 1 Some parts have been explained, so they will not be repeated here.

[0049] Step S22: For each prediction block, multiple residual transformation methods are used to quantize the pixel difference between the image block corresponding to the prediction block and the prediction block to obtain multiple residual data corresponding to the prediction block. As mentioned earlier, most related technologies employ fixed residual transformation methods to obtain residual data. These standardized residual transformation methods all generate residual data by converting pixel-domain data to the frequency domain. When image content is highly complex, the local characteristics of different image blocks vary significantly. If the same residual transformation method is applied to all image blocks, it may lead to high redundancy in the quantized residual data. This not only increases the bitrate overhead but also affects the reconstruction quality of the current frame.

[0050] Based on this, this application provides a new residual transformation method in addition to the traditional residual transformation method. Specifically, the residual transformation method of this application includes: a first transformation method and a second transformation method.

[0051] The first transformation method refers to the traditional residual transformation method that converts pixel domain data to the frequency domain, such as the Discrete Cosine Transform (DCT) and Discrete Sine Transform (DST).

[0052] The second transformation method is to use a trained residual recognition model to encode the features of pixel domain data in order to construct residual data with high texture complexity and rich feature details.

[0053] It should be noted that the residual transformation method in this application embodiment refers not only to the transformation process from the pixel domain to the transform domain (or feature domain), but also includes the complete processing chain for subsequent quantization of the transformation result. Therefore, whether it is the first transformation method (such as DCT, DST, etc.) or the second transformation method (feature encoding based on the trained residual recognition model), the residual data generated by each method is already in the final form after quantization processing and can be directly used for entropy coding and bitstream output.

[0054] To facilitate understanding, the training process of the residual recognition model in this application embodiment will be explained below: A pre-constructed autoencoder-type training model containing a model encoder and a model decoder is constructed. The model encoder is used to extract and compress features from the input pixel domain data to generate compact residual features. The model decoder is used to decode the residual features back to the pixel domain to obtain the reconstructed residual data.

[0055] Next, the cleaned historical video frame set (e.g., a complete video frame sequence of a certain video) is used to perform multiple rounds of iterative training on the residual recognition model to be trained until the preset iteration stopping condition is met, and the trained residual recognition model is obtained.

[0056] Each iteration is as follows: From any historical video frame in the historical video frame set, a training image patch is selected. Multiple training prediction blocks are generated corresponding to this training image patch using various image prediction methods. Then, the pixel differences between this training image patch and each training prediction block are feature-encoded using the current model parameters to obtain the first residual feature for this round. Next, the first residual feature is decoded using the current model parameters to obtain the training residual data for this round of reconstruction. The iterative loss for this round is obtained by calculating the loss between the training residual data and the pixel differences, and the current model parameters are adjusted based on this iterative loss.

[0057] For ease of understanding, the specifics are as follows: Figure 3 As shown, for training image block Q, training prediction block A is generated using inter-frame prediction, and training prediction block B is generated using intra-frame prediction.

[0058] The model encoder uses the current model parameters to encode the pixel difference (QA) between the training image block Q and the training prediction block A, obtaining the first residual feature Q. A Then, feature encoding is performed on the pixel difference (QB) between the training image block Q and the training prediction block B to obtain the first residual feature Q. B The model decoder uses the current model parameters to process the first residual feature Q. A and the first residual feature Q B Feature decoding is performed to obtain the training residual data C for this round of reconstruction. A and C B .

[0059] Next, the training residual data C are calculated using a preset loss function (such as mean squared error MSE, mean absolute error MAE, etc.). A The loss between pixel difference (QA) and training residual data C B The loss between the pixel difference (QB) and the current iteration loss is obtained, and the current model parameters are adjusted based on the current iteration loss.

[0060] Furthermore, to further improve the prediction accuracy and robustness of the model, this embodiment of the application can introduce different contextual features based on the acquisition method of the training prediction blocks before performing feature decoding on the first residual features in the pixel domain, thereby enhancing the first residual features. This feature enhancement method aims to enrich the information content of the first residual features by utilizing information from adjacent image blocks or inter-frames, thereby more accurately capturing detailed changes and motion patterns in the video sequence. Then, feature decoding is performed on the enhanced features to improve reconstruction quality and reduce coding redundancy.

[0061] In some embodiments, a first reference image block with the same position as the training image block can be selected from the reference frame corresponding to the historical video frame, and at least one second reference image block adjacent to the training image block can be selected from the historical video frame where the current training image block is located.

[0062] The reference frame is determined based on the encoded data of a set of historical video frames. Specifically, in each iteration, inter-frame prediction and intra-frame prediction are used to generate corresponding training prediction blocks. During inter-frame prediction, one frame is selected from the encoding results of each historical video frame to generate the training prediction block.

[0063] If the training prediction block is obtained using an inter-frame prediction method, then based on the pixel difference between the first reference image block and the training prediction block, and the pixel difference between the second reference image block and the training prediction block, feature fusion is performed on the first residual features of the training prediction block to obtain the first feature fusion result; feature decoding is performed on the first feature fusion result to obtain the training residual data. Correspondingly, if the training prediction block is obtained by intra-frame prediction, then based on the pixel difference between the second reference image block and the training prediction block, the first residual feature of the training prediction block is fused to obtain the second feature fusion result; the second feature fusion result is then decoded to obtain the training residual data.

[0064] For ease of understanding, the aforementioned Figure 3 To illustrate the feature decoding process described above, we find a first reference image block Q1 located at the same position as the training image block Q in the reference frame corresponding to the historical video frame where the training image block Q is located; and select at least one spatially adjacent second reference image block Q2 from the historical video frame where the training image block is located, for example, selecting a decoded image block located to the left or above the training image block Q.

[0065] For training prediction block A, which is generated through inter-frame prediction and has a reference frame, spatial and temporal contextual information can be introduced for feature enhancement. Specifically, the pixel difference (Q1-A) between the first reference image block Q1 and the training prediction block A, and the pixel difference (Q2-A) between the second reference image block Q2 and the training prediction block A can be used to enhance the first residual feature Q of the training prediction block. A Feature fusion is performed. Subsequently, the feature fusion result is decoded using a model decoder to obtain the training residual data corresponding to training prediction block A.

[0066] Correspondingly, for the training prediction block B, which is generated through intra-frame prediction and lacks a reference frame, spatial contextual information can be introduced for feature enhancement. Specifically, the pixel difference (Q2-B) between the second reference image block Q2 and the training prediction block B can be used to enhance the first residual feature Q of the training prediction block. B Feature fusion is then performed. Subsequently, the fusion result is decoded using a model decoder to obtain the training residual data corresponding to training prediction block B.

[0067] Therefore, during model training, contextual residual information from spatial neighborhoods or spatiotemporal reference locations can be adaptively fused according to different image prediction methods, making the reconstructed training residual data more accurately approximate the original pixel differences. This mechanism enhances the context-awareness and detail-preserving ability of the first residual feature output by the model encoder. Thus, during the inference stage, when this trained residual recognition model is integrated into the video encoder as a second transformation method, the generated residual data can more fully represent the true differences in image patches, thereby reducing bitrate overhead or improving reconstruction quality with the same reconstruction quality and bit budget.

[0068] After each iteration, it is determined whether the current iteration meets the preset iteration convergence condition. This convergence condition may include the iteration loss being less than a threshold, the number of iterations reaching a preset number, etc., which are not limited in this application. If the iteration convergence condition is not met after the current iteration, the current model parameters are adjusted based on the iteration loss of this iteration, and the next iteration begins. If the iteration convergence condition is met, the iteration can be stopped, and the trained residual recognition model is obtained.

[0069] In the above model training process, multiple training prediction blocks are generated for the same training image block under various image prediction methods. Based on the pixel difference between each training prediction block and the training image block, the feature encoding capability of the model encoder for the pixel difference and the reconstruction capability of the model decoder for the first residual feature are jointly optimized, so that the model encoder can more fully retain the texture details, edge information and local structural features in the pixel difference during the feature encoding stage.

[0070] It should be noted that during the inference phase, the trained residual recognition model can be integrated into the video encoder. When the pixel difference is quantized using the second transformation method, the model encoder of the residual recognition model encodes the pixel difference as a feature and outputs it. The output result is the residual data obtained using the second transformation method (corresponding to the first residual feature in the training phase), which participates in subsequent entropy encoding.

[0071] Step S23: Based on the residual data corresponding to each prediction block, select the target prediction method and target transformation method corresponding to the image block from the multiple image prediction methods and the multiple residual transformation methods respectively; As mentioned earlier, in this embodiment, inter-frame prediction and intra-frame prediction methods are used in advance to obtain the inter-frame prediction block and intra-frame prediction block corresponding to the current image block. Subsequently, the first transformation method and the second transformation method are used to quantize each prediction block to obtain the residual data corresponding to each prediction block.

[0072] This process, as Figure 4 As shown, multiple image prediction methods and multiple residual transformation methods are combined to obtain corresponding residual data. For example, residual data 1 is obtained by using inter-frame prediction and the first transformation method; residual data 2 is obtained by using inter-frame prediction and the second transformation method; residual data 3 is obtained by using intra-frame prediction and the first transformation method; and residual data 4 is obtained by using intra-frame prediction and the second transformation method.

[0073] In this step, for each residual data point, the prediction block corresponding to that residual data is reconstructed to obtain the reconstructed block. Then, based on the reconstructed block and the current image block, the rate-distortion cost corresponding to that residual data is determined. After obtaining the rate-distortion cost for each residual data point, the image prediction method corresponding to the residual data with the lowest rate-distortion cost is selected as the target prediction method, and the residual transformation method corresponding to the residual data with the lowest rate-distortion cost is selected as the target transformation method.

[0074] Specifically, Rate-Distortion Optimization (RDO) is a decision criterion used in video coding to balance coding bit overhead and reconstruction quality. Its core idea is to estimate the number of coding bits required for each of the multiple coding candidate schemes for the same image patch, as well as the degree of distortion between the reconstructed result and the original image patch. The weighted sum of these two factors is then used as the rate-distortion cost.

[0075] In this embodiment, for the residual data obtained by combining each image prediction method and residual transformation method, the prediction block corresponding to the residual data is first reconstructed using the residual data to obtain the reconstructed block. Taking the current image block as Q as an example, its corresponding prediction block is P. The prediction block is quantized by a residual transformation method, and the resulting residual data is R. Then the reconstructed block can be represented as: Q'=P+R.

[0076] Subsequently, the distortion D corresponding to this combination is calculated, which is the pixel difference between the reconstructed block Q' and the image block Q, for example, using mean square error or mean absolute error. Simultaneously, the number of bits required to write the residual data R, the first identifier of the corresponding image prediction method, and the second identifier of the corresponding residual transformation method into the bitstream is recorded as the rate overhead R'. Finally, the rate-distortion cost (J) corresponding to this combination is calculated using the following formula: J = D + λ·R', yielding the rate-distortion cost J corresponding to the residual data. Here, λ is a Lagrange multiplier, whose value is determined by the set quantization parameter QP, used to balance the relative weights of distortion and bit rate.

[0077] After obtaining the rate-distortion cost corresponding to each residual data point, the image prediction method and residual transformation method corresponding to the residual data with the minimum rate-distortion cost are used as the corresponding target prediction method and target transformation method. For example... Figure 4 The rate-distortion cost corresponding to residual data 3 shown is the lowest. At this time, the inter-frame prediction method corresponding to residual data 3 is taken as the target prediction method, and the first transformation method corresponding to residual data 3 is taken as the target transformation method.

[0078] Since each combination of image prediction and residual transformation methods corresponds to a calculable rate-distortion cost, and this cost comprehensively reflects the number of bits required to write the residual data, the first identifier of the image prediction method, and the second identifier of the residual transformation method into the bitstream, as well as the pixel difference between the reconstructed block obtained by this combination and the current image block, all candidate combinations can be quantitatively compared under the same evaluation criterion. Thus, by selecting the image prediction method corresponding to the residual data with the minimum rate-distortion cost as the target prediction method, and its corresponding residual transformation method as the target transformation method, it can be ensured that the selected scheme achieves optimal rate-distortion performance while balancing reconstruction quality and bit rate consumption. This avoids suboptimal choices that might result from decisions based solely on the size of the prediction residual or solely on the complexity of the transformation method, improving the accuracy and adaptability of coding decisions.

[0079] Step S24: Add the first identifier of the target prediction method, the second identifier of the target transformation method, and the target residual data to the bitstream; wherein, the target residual data is obtained by quantizing the pixel difference between the image block and the prediction block generated by the target prediction method using the target transformation method.

[0080] The target residual data is the residual data with the lowest rate-distortion cost mentioned above. In this embodiment of the application, after selecting the target prediction method and the target transformation method, the first identifier corresponding to the target prediction method, the second identifier corresponding to the target transformation method, and the target residual data obtained by quantizing the pixel difference between the current image block and the target prediction block using the target transformation method are written into the bitstream together.

[0081] It should be noted that in the current mainstream video codec standards, the bitstream syntax does not reserve a field for indicating residual data generated by the model. Therefore, the decoder cannot recognize or correctly parse the residual data generated by model-driven methods (such as the second transformation method in this application).

[0082] Based on this, this application explicitly introduces a first identifier and a second identifier as custom control information in the bitstream to instruct the video decoder how to select the correct residual transformation method to reconstruct the current image block.

[0083] The first identifier indicates the image prediction method used (such as inter-frame prediction or intra-frame prediction), and the second identifier indicates the residual transformation method used in the encoding stage (such as the first transformation method or the second transformation method). With this design, even if the target residual data is output by a residual recognition model and is a non-traditional data representation, the video decoder can accurately determine which inverse quantization method should be invoked for processing based on the second identifier.

[0084] For example, the second identifier corresponding to the first transformation method can be set to "0", and the second identifier corresponding to the second transformation method can be set to "1". When an image block uses inter-frame prediction (first identifier is "0") and the second transformation method (second identifier is "1"), "0", "1" and the corresponding target residual data will be written into the bitstream. After the video decoder reads the second identifier "1", it can activate the inverse quantization module corresponding to the trained residual recognition model to restore the residual data, instead of using the traditional inverse quantization process.

[0085] In this way, after the video decoder receives the bitstream, it can determine the image prediction method to be used based on the first identifier to generate a prediction block, and determine the inverse transform method to be used based on the second identifier to perform inverse quantization processing on the target residual data. Then, the inverse quantized residual data is added to the prediction block to reconstruct the current image block. Since the bitstream explicitly carries the prediction method and transform method information that matches the target residual data, reconstruction deviations caused by missing or misjudged mode information during decoding are avoided, ensuring the consistency and accuracy of the encoding and decoding process.

[0086] Furthermore, to further improve the processing speed, in this embodiment of the application, before quantizing each prediction block using each residual transformation method, it is necessary to screen each prediction block to select candidate prediction blocks that meet the quantization conditions. Subsequently, only the candidate prediction blocks are quantized.

[0087] In some embodiments, after obtaining each prediction block (i.e., inter-frame prediction block and intra-frame prediction block) corresponding to the current image block, for each pixel in the current image block, the pixel difference between that pixel and a pixel at the same position in the inter-frame prediction block, and the pixel difference between that pixel and a pixel at the same position in the intra-frame prediction block are determined respectively. Subsequently, based on the sum of the absolute values ​​of the corresponding pixel differences of each pixel in the current image block in the inter-frame prediction block, and the sum of the absolute values ​​of the corresponding pixel differences in the intra-frame prediction block, candidate prediction blocks for subsequent quantization processing are determined from the inter-frame prediction block and the intra-frame prediction block.

[0088] Specifically, each pixel in the current image block can be represented as ( , , … Each pixel in the inter-frame prediction block can be represented as ( , , … Each pixel in the intra-prediction block can be represented as ( , , … ).

[0089] Calculate the pixel difference between each pixel in the current image patch and the pixel at the same position in the inter-frame prediction patch: {( - ), ( - ), ( - )...( - )}, and the pixel difference between each pixel in the current image block and the pixel at the same position in the intra-predicted block: { ( - ), ( - ), ( - )...( - )}.

[0090] Next, calculate the sum of the absolute values ​​of the pixel differences in the corresponding pixel values ​​in the inter-frame prediction block for each pixel in the current image block: + + +… And the sum of the absolute values ​​of the pixel differences between the corresponding pixels in the current image block and the corresponding pixels in the intra-frame prediction block: + + +… Finally, based on the absolute value of this difference, candidate prediction blocks are selected from inter-frame prediction blocks and intra-frame prediction blocks.

[0091] Since the pixel differences between each pixel in the current image block and different prediction blocks reflect the content similarity between each prediction block and the current image block, the total amount of residual data that may be generated by the image prediction method corresponding to each prediction block can be quantified by summing the absolute values ​​of the pixel differences between each pixel in the current image block and different prediction blocks. Therefore, by pre-selecting candidate prediction blocks that have a significant impact on reconstruction quality from inter-frame and intra-frame prediction blocks based on the absolute values ​​of the pixel differences between the current image block and each prediction block, the number of prediction blocks participating in the quantization process can be reduced before entering the residual quantization stage, thereby reducing the amount of computation and improving the overall coding speed.

[0092] In some embodiments, candidate prediction blocks can be selected from intra-frame prediction blocks and inter-frame prediction blocks in the following manner: Pre-determine whether the difference between the sum of the absolute values ​​of the pixel differences of each pixel in the current image block in the inter-frame prediction block and the sum of the absolute values ​​of the pixel differences of each pixel in the intra-frame prediction block is greater than a preset sum value threshold. Right now, + + +… and + + +… If the difference between the two prediction blocks is greater than a threshold, and is less than or equal to the threshold, it means that the sum of the absolute values ​​of the pixel differences between the inter-frame prediction block and the intra-frame prediction block is similar. In this case, both the inter-frame prediction block and the intra-frame prediction block are considered as candidate prediction blocks.

[0093] Correspondingly, if the difference is greater than the threshold, that is, there is a significant difference in the sum of the absolute values ​​of the pixel differences between the inter-frame prediction block and the intra-frame prediction block, it indicates that there is a significant difference in the overall residual energy of the two prediction blocks, and one prediction method is significantly better than the other. In addition, to further improve the prediction accuracy, the comparison of the above difference with the threshold can be replaced by comparing the absolute value of the difference with the threshold, which is not limited in this application.

[0094] Considering that relying solely on the sum of absolute values ​​may overlook the characteristics of local error distribution, leading to misjudgment, this application further calculates the first standard deviation of the pixel difference between each pixel in the current image block and the pixel at the same position in the inter-frame prediction block, and the second standard deviation of the pixel difference between each pixel at the same position in the intra-frame prediction block. Based on the first and second standard deviations, candidate prediction blocks are determined from the inter-frame and intra-frame prediction blocks.

[0095] Specifically, the standard deviation reflects the dispersion of pixel differences among pixels in the current image block. Therefore, by comparing the first standard deviation and the second standard deviation, the distribution characteristics of pixel differences corresponding to inter-frame prediction blocks and intra-frame prediction blocks can be further distinguished. If the first standard deviation is less than the second standard deviation, the inter-frame prediction block is considered a candidate prediction block; if the first standard deviation is greater than the second standard deviation, the intra-frame prediction block is considered a candidate prediction block; if the first standard deviation is equal to the second standard deviation, both the inter-frame prediction block and the intra-frame prediction block are considered candidate prediction blocks.

[0096] The above process quantizes all prediction blocks when the sum of the absolute values ​​of the pixel differences between the current image block and each prediction block is similar, in order to avoid missing potentially better candidates. Meanwhile, when the sum of the absolute values ​​differs significantly, the standard deviation is used to assist in the judgment, reducing the risk of incorrectly discarding suitable prediction blocks due to relying solely on the sum of absolute values, and effectively controlling the number of prediction blocks entering subsequent quantization processes.

[0097] To facilitate understanding, an example is used to illustrate this. After generating the inter-frame prediction block and intra-frame prediction block corresponding to the current image block using inter-frame prediction and intra-frame prediction methods, if the inter-frame prediction block is determined as a candidate prediction block after the above candidate prediction block judgment, then only the quantization processing of the inter-frame prediction block needs to be performed. This eliminates the need for the quantization processing of the intra-frame prediction block and the related process of calculating the rate-distortion cost of the residual data obtained from its quantization processing, thereby reducing bit rate overhead and improving processing efficiency.

[0098] Furthermore, considering that if the current image block contains a large, flat area, the target prediction block obtained through the above process may be completely identical to the image content of the current image block. In this case, the value of the target residual data obtained after quantization is zero. Writing the target residual data with a value of zero into the bitstream would occupy unnecessary codewords, causing coding redundancy.

[0099] Before performing the aforementioned step S24 in this embodiment, it is necessary to determine in advance whether the target residual data meets the preset residual threshold. In some embodiments, if the target residual data meets the preset residual threshold (i.e., the amount of data in the target residual data is 0), then the target identifier and the first identifier are added to the bitstream; wherein, the target identifier is used to instruct the video decoder to generate a prediction block using the target prediction method corresponding to the first identifier, and to use the prediction block as a reconstruction block of the image block.

[0100] Specifically, when the amount of target residual data is 0, it means that the current image patch can be completely predicted by the target prediction method without relying on residual data for correction.

[0101] At this point, by writing the target identifier and the first identifier into the bitstream, the video decoder can be instructed to determine the target prediction method based on the first identifier to generate a prediction block. Following the indication of the target identifier, this prediction block is directly used as the reconstruction block of the current image block, without needing to read or dequantize any residual data. This eliminates the need to write the target residual data when it is zero, effectively reducing redundant bits in the bitstream.

[0102] The following is a detailed description of a video decoding method provided in an embodiment of this application. This method can be applied to a video decoder, as detailed below. Figure 5 As shown, the method includes the following steps: Step S51: Decode the first identifier, the second identifier, and the target residual data corresponding to each image block in the current frame from the bitstream; As mentioned earlier, if the content of the current image block is a large flat area, the target prediction block obtained by the video encoder may be completely consistent with the image content of the current image block, that is, the value of the target residual data is zero. If the target residual data with a value of zero is written into the bitstream, it will occupy unnecessary codewords and cause coding redundancy. Based on this, the video encoder of this application embodiment will write the first identifier and the target identifier representing that the value of the target residual data is zero into the bitstream when the value of the target residual data is zero.

[0103] Therefore, before executing step S51, it must be determined in advance that the target identifier has not been decoded from the bitstream. If the target identifier is decoded from the bitstream, a prediction block is generated using the target prediction method corresponding to the first identifier, and this prediction block is directly used as the reconstruction block of the current image block.

[0104] Step S52: For each image block, select the target prediction method corresponding to the first identifier from multiple image prediction methods, and select the target transformation method corresponding to the second identifier from multiple residual transformation methods; As mentioned in step S24 above, after determining the target residual data, the video encoder will add the target residual data, as well as the identification information of the target prediction method and the target transformation method that generated the target residual data, to the bitstream.

[0105] After the video decoder extracts the above information from the bitstream, it can determine the target prediction method and target transformation method for generating target residual data based on the identifier indication.

[0106] Step S53: Generate a prediction block using the target prediction method, and perform inverse quantization processing on the target residual data using the target transformation method to obtain reconstructed residual data; The video decoder of this application is configured with multiple image prediction methods and multiple residual transformation methods, the same as those in the video encoder. After determining the target prediction method based on the first identifier, a prediction block can be generated using the same method as the video encoder. The generation method of the prediction block will not be described in detail here.

[0107] After determining the target transformation method based on the second identifier, the target residual data can be inversely quantized based on the target transformation method to obtain the reconstructed residual data.

[0108] Specifically, if the target transformation method is the first transformation method (i.e., the traditional residual transformation method that converts pixel domain data to frequency domain), the reconstructed residual data can be obtained by performing an inverse transformation operation from frequency domain to pixel domain on the target residual data.

[0109] Correspondingly, if the target transformation method is the second transformation method (i.e., using the trained residual recognition model to encode the features of the pixel domain data), then the model decoder in the residual recognition model is called, the target residual data is used as input, the model decoder performs feature decoding on the input content, and outputs the corresponding reconstructed residual data.

[0110] In this way, the video decoder can accurately reproduce the prediction and dequantization paths used by the video encoder, ensuring that the reconstructed residual data is semantically and numerically similar to the original pixel differences, thereby improving reconstruction quality. Simultaneously, since the target residual data has already undergone the optimal transformation method adaptively selected by the video encoder based on the image content characteristics, both reconstruction quality and bitrate efficiency can be balanced without increasing additional signaling overhead.

[0111] Step S54: Based on the prediction block corresponding to each image block and the reconstruction residual data, obtain the reconstruction block corresponding to each image block, and determine the reconstruction frame corresponding to the current frame based on the reconstruction blocks.

[0112] For each image block to be reconstructed indicated in the bitstream, the video decoder generates a corresponding prediction block based on the target prediction method. Subsequently, the target residual data is dequantized using the target transformation method to obtain the reconstruction residual data. The reconstruction residual data is then added pixel by pixel to the pixels in the prediction block to obtain the reconstruction block corresponding to the image block.

[0113] After obtaining the reconstructed blocks corresponding to each image block in the current frame through the above process, the image blocks are combined according to their spatial position relationship in the current frame to form a complete reconstructed frame. The reconstructed frame is then stored in the decoding image buffer for use in encoding / decoding or displaying subsequent video frames.

[0114] In the above process, it is determined in advance whether the bitstream contains a target identifier: if a target identifier exists, the prediction block is directly generated using the target prediction method corresponding to the first identifier, and the prediction block is used as the reconstruction block of the current image block. There is no need to parse and dequantize the target residual data, thereby saving decoding computing resources and avoiding redundant operations caused by processing all-zero residuals.

[0115] Correspondingly, if no target identifier exists, the first identifier, the second identifier, and the target residual data are parsed from the bitstream. Based on the first identifier, the target prediction method is determined to generate a prediction block. Based on the second identifier, the target transformation method is determined, and the target residual data is dequantized accordingly. Then, the resulting reconstructed residual data is added pixel-by-pixel to the prediction block to obtain the reconstructed block. Through this mechanism, the decoder can collaborate with the video encoder to skip residual transmission and processing when the residual is zero, and accurately restore the adaptive quantization result when the residual is non-zero. This reduces the complexity of bitstream parsing while ensuring reconstruction quality.

[0116] Based on the same inventive concept, this application also provides a video encoding device 600, specifically as follows: Figure 6 As shown, the device includes: Prediction unit 601 is configured to: generate multiple prediction blocks corresponding to each image block by using multiple image prediction methods for each image block in the current frame; The quantization unit 602 is configured to: for each prediction block, use multiple residual transformation methods to quantize the pixel difference between the image block corresponding to the prediction block and the prediction block to obtain multiple residual data corresponding to the prediction block; The confirmation unit 603 is configured to: select the target prediction method and target transformation method corresponding to the image block from the multiple image prediction methods and the multiple residual transformation methods based on the residual data corresponding to each prediction block; The encoding unit 604 is configured to add a first identifier of the target prediction method, a second identifier of the target transformation method, and target residual data to the bitstream; wherein the target residual data is obtained by quantizing the pixel difference between the image block and the prediction block generated by the target prediction method using the target transformation method.

[0117] In some embodiments, the multiple residual transformation methods include: a first transformation method that transforms pixel-domain data to the frequency domain, and a second transformation method that performs feature encoding on pixel-domain data based on a residual recognition model; the residual recognition model is obtained through multiple rounds of iterative training based on a set of historical video frames; Each iteration is as follows: Select a training image patch from any historical video frame in the set of historical video frames; Multiple training prediction blocks corresponding to the training image blocks are generated by using the various image prediction methods described above. For the pixel difference between the training image block and each training prediction block, the pixel difference is feature-encoded using the current model parameters to obtain the first residual feature of this round; and the first residual feature is feature-decoded using the current model parameters to obtain the training residual data of this round of reconstruction. Based on the training residual data and the pixel difference, determine the iteration loss for this round; The current model parameters are adjusted based on the loss from this iteration.

[0118] In some embodiments, the multiple image prediction methods include: inter-frame prediction and intra-frame prediction; performing feature decoding on the first residual features using the current model parameters to obtain the training residual data for this round of reconstruction, the quantization unit 602 is specifically configured as follows: From the reference frames corresponding to the historical video frames, a first reference image block with the same position as the training image block is selected, and from the historical video frames where the training image block is located, at least one second reference image block adjacent to the training image block is selected; wherein, the reference frames are determined based on the encoded data of the historical video frame set; If the training prediction block is obtained using the inter-frame prediction method, then based on the pixel difference between the first reference image block and the training prediction block, and the pixel difference between the second reference image block and the training prediction block, feature fusion is performed on the first residual feature of the training prediction block to obtain a first feature fusion result; feature decoding is performed on the first feature fusion result to obtain the training residual data. If the training prediction block is obtained using the intra-frame prediction method, then based on the pixel difference between the second reference image block and the training prediction block, the first residual feature of the training prediction block is fused to obtain the second feature fusion result; the second feature fusion result is then decoded to obtain the training residual data.

[0119] In some embodiments, the target residual data is determined in the following manner: For each residual data, the prediction block corresponding to the residual data is reconstructed based on the residual data to obtain the reconstructed block corresponding to the residual data; and, based on the reconstructed block corresponding to the residual data and the image block, the rate-distortion cost corresponding to the residual data is determined. The residual data with the lowest rate-distortion cost is taken as the target residual data.

[0120] In some embodiments, the multiple image prediction methods include: inter-frame prediction and intra-frame prediction: before performing the quantization processing on the pixel difference between the image block corresponding to the prediction block and the prediction block by employing multiple residual transformation methods for each prediction block, the quantization unit 602 is further configured as follows: The predicted block is determined to be a candidate predicted block; The candidate prediction blocks are determined in the following way: For each pixel in the image block, the pixel difference between the pixel and a pixel at the same position in the inter-frame prediction block, and the pixel difference between the pixel and a pixel at the same position in the intra-frame prediction block are determined respectively; wherein, the inter-frame prediction block is a prediction block generated using the inter-frame prediction method; and the intra-frame prediction block is a prediction block generated using the intra-frame prediction method. The candidate prediction block is determined from the inter-frame prediction block and the intra-frame prediction block based on the sum of the absolute values ​​of the corresponding pixel differences of each pixel in the image block in the inter-frame prediction block and the sum of the absolute values ​​of the corresponding pixel differences in the intra-frame prediction block.

[0121] In some embodiments, the process of determining the candidate prediction block from the inter-frame prediction block and the intra-frame prediction block based on the sum of the absolute values ​​of the corresponding pixel differences of each pixel point in the image block in the inter-frame prediction block and the sum of the absolute values ​​of the corresponding pixel differences in the intra-frame prediction block, wherein the quantization unit 602 is specifically configured to: If the sum of the absolute values ​​of the pixel differences corresponding to each pixel in the inter-frame prediction block and the sum of the absolute values ​​of the pixel differences corresponding to each pixel in the intra-frame prediction block are less than or equal to a preset sum threshold, then both the inter-frame prediction block and the intra-frame prediction block are considered as candidate prediction blocks. If the difference is greater than the preset sum threshold, then the first standard deviation of the pixel difference between each pixel in the image block and the pixel at the same position in the inter-frame prediction block, and the second standard deviation of the pixel difference between each pixel and the pixel at the same position in the intra-frame prediction block are determined; based on the first standard deviation and the second standard deviation, the candidate prediction block is determined from the inter-frame prediction block and the intra-frame prediction block.

[0122] In some embodiments, before adding the first identifier of the target prediction method, the second identifier of the target transformation method, and the target residual data to the bitstream, the encoding unit 604 is further configured to: The target residual data is determined to not meet the preset residual threshold. The encoding unit is further configured as follows: If the target residual data satisfies the preset residual threshold, then the target identifier and the first identifier are added to the bitstream; wherein, the target identifier is used to instruct the video decoder to generate a prediction block using the target prediction method corresponding to the first identifier, and to use the prediction block as the reconstruction block of the image block.

[0123] Based on the same inventive concept, this application also provides a video decoding device 700, specifically as follows: Figure 7 As shown, the device includes: Decoding unit 701 is configured to decode the first identifier, the second identifier, and the target residual data corresponding to each image block in the current frame from the bitstream; The identification unit 702 is configured to: for each image block, select the target prediction method corresponding to the first identifier from multiple image prediction methods, and select the target transformation method corresponding to the second identifier from multiple residual transformation methods; The residual unit 703 is configured to: generate a prediction block using the target prediction method, and perform inverse quantization processing on the target residual data using the target transformation method to obtain reconstructed residual data; The reconstruction unit 704 is configured to: obtain the reconstruction block corresponding to each image block based on the prediction block corresponding to each image block and the reconstruction residual data, and determine the reconstruction frame corresponding to the current frame based on the reconstruction blocks.

[0124] In some embodiments, the multiple residual transformation methods include: a first transformation method that transforms pixel-domain data to the frequency domain, and a second transformation method that performs feature encoding on pixel-domain data based on a residual recognition model; the residual recognition model is obtained through multiple rounds of iterative training based on a set of historical video frames; Each iteration is as follows: Select a training image patch from any historical video frame in the set of historical video frames; Multiple training prediction blocks corresponding to the training image blocks are generated by using the various image prediction methods described above. For the pixel difference between the training image block and each training prediction block, the pixel difference is feature-encoded using the current model parameters to obtain the first residual feature of this round; and the first residual feature is feature-decoded using the current model parameters to obtain the training residual data of this round of reconstruction. Based on the training residual data and the pixel difference, determine the iteration loss for this round; The current model parameters are adjusted based on the loss from this iteration.

[0125] In some embodiments, the multiple image prediction methods include: inter-frame prediction and intra-frame prediction; performing feature decoding on the first residual feature using the current model parameters to obtain the training residual data for this round of reconstruction, wherein the residual unit 703 is specifically configured as follows: From the reference frames corresponding to the historical video frames, a first reference image block with the same position as the training image block is selected, and from the historical video frames where the training image block is located, at least one second reference image block adjacent to the training image block is selected; wherein, the reference frames are determined based on the encoded data of the historical video frame set; If the training prediction block is obtained using the inter-frame prediction method, then based on the pixel difference between the first reference image block and the training prediction block, and the pixel difference between the second reference image block and the training prediction block, feature fusion is performed on the first residual feature of the training prediction block to obtain a first feature fusion result; feature decoding is performed on the first feature fusion result to obtain the training residual data. If the training prediction block is obtained using the intra-frame prediction method, then based on the pixel difference between the second reference image block and the training prediction block, the first residual feature of the training prediction block is fused to obtain the second feature fusion result; the second feature fusion result is then decoded to obtain the training residual data.

[0126] In some embodiments, before performing the decoding of the first identifier, the second identifier, and the target residual data corresponding to each image block in the current frame from the bitstream, the decoding unit 701 is further configured to: It was determined that the target identifier was not decoded in the bitstream; The decoding unit 701 is further configured to: if the target identifier is decoded from the bitstream, generate a prediction block using the target prediction method corresponding to the first identifier, and use the prediction block as the reconstruction block of the image block.

[0127] Based on the same inventive concept, this application also provides a video encoder, which may include: a memory and a processor, wherein the memory stores a computer program, and the processor is configured to implement the video encoding method provided in any of the above embodiments when executing the computer program.

[0128] Based on the same inventive concept, this application also provides a video decoder, including: a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the computer program to implement the video decoding method provided in any of the above embodiments.

[0129] Figure 8 Exemplary block diagrams of video encoders or video decoders from some of the embodiments described above are shown. Figure 8 As shown, the video encoder or video decoder 800 may include at least one of the following: a tuner / demodulator 801, a communicator 802, a detector 803, an external device interface 804, a processor 805, a display 806, an audio output interface 807, a memory 808, a power supply 809, and a user interface 810.

[0130] In some embodiments, processor 805 includes at least one of: a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), RAM (randomaccess memory), ROM (read-only memory), a first to an nth interface for input / output, a communication bus, etc.

[0131] The display 806 includes a display screen assembly for presenting images, a driving assembly for driving image display, a component for receiving image signals from the processor output, and a user interface for displaying video content, image content, menu control interface, and user control UI.

[0132] The display 806 can be a liquid crystal display, an organic light-emitting diode (OLED) display, or a projection display, etc.

[0133] The communicator 802 is a component used to communicate with external devices or servers according to various communication protocol types. For example, the communicator may include at least one of the following: a Wi-Fi module, a Bluetooth module, a wired Ethernet module, other network communication protocol chips or near-field communication protocol chips, and an infrared receiver. The video encoder or video decoder 800 can use the communicator 802 to send and receive control signals and data signals with the control device or server.

[0134] The user interface can be used to receive control signals input by the user through a control device (such as an infrared remote control) or by touch or gesture.

[0135] Detector 803 can be used to acquire signals from the external environment or to interact with the external environment. For example, detector 803 may include a light receiver, which can be used to acquire ambient light intensity; or, detector 803 may include an image acquisition device, such as a camera, which can be used to acquire external environmental scenes, user attributes, or user interaction gestures; or, detector 803 may include a sound acquisition device, such as a microphone, for receiving external sounds.

[0136] The external device interface 804 may include, but is not limited to, one or more of the following: High-Definition Multimedia Interface (HDMI), analog or high-definition component input interface (component), Composite Video Broadcast Signal (CVBS), Universal Serial Bus (USB), RGB (Red, Green, Blue) port, etc. It may also be a composite input / output interface formed by multiple of the above interfaces.

[0137] The tuner / demodulator 801 receives broadcast television signals via wired or wireless means, and demodulates audio and video signals, such as EPG data signals, from multiple wireless or wired broadcast television signals. In some embodiments, the processor 805 and the tuner / demodulator 801 may be located in different separate devices, that is, the tuner / demodulator 801 may also be located in an external device of the main device where the processor 805 is located, such as an external set-top box.

[0138] Processor 805 controls the operation of electronic devices and responds to user operations through various software control programs stored in memory. Processor 805 controls the overall operation of video encoder or video decoder 800. For example, in response to receiving a user command to select a UI object to be displayed on display 806, processor 805 can perform operations related to the object selected by the user command.

[0139] In some embodiments, processor 805 includes at least one of: a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), RAM (RandomAccess Memory), ROM (Read-Only Memory), a first to an nth interface for input / output, a communication bus, etc.

[0140] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a computing device, the computing device implements the video encoding method or the video decoding method described in any of the above embodiments.

[0141] This application also provides a computer program product that, when run on a computer, enables the computer to implement the video encoding method or the video decoding method described in any of the above embodiments.

[0142] Those skilled in the art will understand that all or part of the steps of the foregoing method embodiments can be implemented by a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed, it performs the steps of the foregoing method embodiments. The storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0143] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to the prior art, can be embodied in the form of software products, for example, through computer program products. These computer program products are stored in a storage medium and include computer programs used to cause a computer device to execute all or part of the methods described in the embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

Claims

1. A video encoding method, characterized in that, Applied to a video encoder, the method includes: For each image block in the current frame, multiple image prediction methods are used to generate multiple prediction blocks corresponding to each image block; For each prediction block, multiple residual transformation methods are used to quantize the pixel difference between the image block corresponding to the prediction block and the prediction block to obtain multiple residual data corresponding to the prediction block. Based on the residual data corresponding to each prediction block, the target prediction method and target transformation method corresponding to the image block are selected from the multiple image prediction methods and the multiple residual transformation methods, respectively. The first identifier of the target prediction method, the second identifier of the target transformation method, and the target residual data are added to the bitstream; wherein, the target residual data is obtained by quantizing the pixel difference between the image block and the prediction block generated by the target prediction method using the target transformation method.

2. The method according to claim 1, characterized in that, The various residual transformation methods include: a first transformation method that converts pixel domain data to the frequency domain, and a second transformation method that encodes the features of pixel domain data based on a residual recognition model; the residual recognition model is obtained through multiple rounds of iterative training based on a set of historical video frames. Each iteration is as follows: Select a training image patch from any historical video frame in the set of historical video frames; Multiple training prediction blocks corresponding to the training image blocks are generated by using the various image prediction methods described above. For the pixel difference between the training image block and each training prediction block, the pixel difference is feature-encoded using the current model parameters to obtain the first residual feature of this round; and the first residual feature is feature-decoded using the current model parameters to obtain the training residual data of this round of reconstruction. Based on the training residual data and the pixel difference, determine the iteration loss for this round; The current model parameters are adjusted based on the loss from this iteration.

3. The method according to claim 2, characterized in that, The various image prediction methods include: inter-frame prediction and intra-frame prediction; the step of using the current model parameters to perform feature decoding on the first residual features to obtain the training residual data for this round of reconstruction includes: From the reference frames corresponding to the historical video frames, a first reference image block with the same position as the training image block is selected, and from the historical video frames where the training image block is located, at least one second reference image block adjacent to the training image block is selected; wherein, the reference frames are determined based on the encoded data of the historical video frame set; If the training prediction block is obtained using the inter-frame prediction method, then based on the pixel difference between the first reference image block and the training prediction block, and the pixel difference between the second reference image block and the training prediction block, feature fusion is performed on the first residual feature of the training prediction block to obtain a first feature fusion result; feature decoding is performed on the first feature fusion result to obtain the training residual data. If the training prediction block is obtained using the intra-frame prediction method, then based on the pixel difference between the second reference image block and the training prediction block, the first residual feature of the training prediction block is fused to obtain the second feature fusion result; the second feature fusion result is then decoded to obtain the training residual data.

4. The method according to claim 1, characterized in that, The target residual data is determined in the following way: For each residual data, the prediction block corresponding to the residual data is reconstructed based on the residual data to obtain the reconstructed block corresponding to the residual data; and, based on the reconstructed block corresponding to the residual data and the image block, the rate-distortion cost corresponding to the residual data is determined. The residual data with the lowest rate-distortion cost is taken as the target residual data.

5. The method according to any one of claims 1-4, characterized in that, The multiple image prediction methods include: inter-frame prediction and intra-frame prediction. Before quantizing the pixel difference between the image block corresponding to the prediction block and the prediction block by employing multiple residual transformation methods for each prediction block, the method further includes: The predicted block is determined to be a candidate predicted block; The candidate prediction blocks are determined in the following way: For each pixel in the image block, the pixel difference between the pixel and a pixel at the same position in the inter-frame prediction block, and the pixel difference between the pixel and a pixel at the same position in the intra-frame prediction block are determined respectively; wherein, the inter-frame prediction block is a prediction block generated using the inter-frame prediction method; and the intra-frame prediction block is a prediction block generated using the intra-frame prediction method. The candidate prediction block is determined from the inter-frame prediction block and the intra-frame prediction block based on the sum of the absolute values ​​of the corresponding pixel differences of each pixel in the image block in the inter-frame prediction block and the sum of the absolute values ​​of the corresponding pixel differences in the intra-frame prediction block.

6. The method according to claim 5, characterized in that, The step of determining the candidate prediction block from the inter-frame prediction block and the intra-frame prediction block based on the sum of the absolute values ​​of the corresponding pixel differences of each pixel point in the image block in the inter-frame prediction block and the sum of the absolute values ​​of the corresponding pixel differences in the intra-frame prediction block includes: If the sum of the absolute values ​​of the pixel differences corresponding to each pixel in the inter-frame prediction block and the sum of the absolute values ​​of the pixel differences corresponding to each pixel in the intra-frame prediction block are less than or equal to a preset sum threshold, then both the inter-frame prediction block and the intra-frame prediction block are considered as candidate prediction blocks. If the difference is greater than the preset sum threshold, then the first standard deviation of the pixel difference between each pixel in the image block and the pixel at the same position in the inter-frame prediction block, and the second standard deviation of the pixel difference between each pixel and the pixel at the same position in the intra-frame prediction block are determined; based on the first standard deviation and the second standard deviation, the candidate prediction block is determined from the inter-frame prediction block and the intra-frame prediction block.

7. The method according to any one of claims 1-4, characterized in that, Before adding the first identifier of the target prediction method, the second identifier of the target transformation method, and the target residual data to the bitstream, the method further includes: The target residual data is determined to not meet the preset residual threshold. The method further includes: If the target residual data satisfies the preset residual threshold, then the target identifier and the first identifier are added to the bitstream; wherein, the target identifier is used to instruct the video decoder to generate a prediction block using the target prediction method corresponding to the first identifier, and to use the prediction block as the reconstruction block of the image block.

8. A video decoding method, characterized in that, Applied to a video decoder, the method includes: Decode the first identifier, second identifier, and target residual data corresponding to each image block in the current frame from the bitstream; For each image patch, the target prediction method corresponding to the first identifier is selected from multiple image prediction methods, and the target transformation method corresponding to the second identifier is selected from multiple residual transformation methods; The target prediction method is used to generate a prediction block, and the target transformation method is used to perform inverse quantization on the target residual data to obtain reconstructed residual data; Based on the prediction block corresponding to each image block and the reconstruction residual data, the reconstruction block corresponding to each image block is obtained, and the reconstruction frame corresponding to the current frame is determined based on the reconstruction block.

9. A video encoder, characterized in that, include: Memory, used to store computer programs; A processor, configured to execute the video encoding method according to any one of claims 1-7 when running the computer program.

10. A video decoder, characterized in that, include: Memory, used to store computer programs; A processor, configured to execute the video decoding method according to claim 8 when running the computer program.