Image decoding method, image coding method and related equipment

By appropriately disabling some buffered information during image encoding and decoding, and fusing advanced prior decoding motion features and residual features, the decoding residual features are optimized, thus solving the problem of low decoding accuracy in existing technologies and achieving higher decoding precision and error control.

CN121644833APending Publication Date: 2026-03-10ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-20
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing image encoding and decoding methods have low decoding accuracy.

Method used

When the current frame is a refresh frame, some types of cached information are appropriately disabled, and only some types of cached information are used during the decoding process; when the current frame is a refresh frame, only some types of cached information, including reference frame features, are enabled; during the decoding process, super-prior decoding motion features and residual features are fused to improve decoding accuracy; and temporal residual features and prediction features are used to optimize decoding residual features to improve decoding performance.

Benefits of technology

It improves the accuracy of image decoding, reduces error accumulation, and enhances the precision of the encoding and decoding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121644833A_ABST
    Figure CN121644833A_ABST
Patent Text Reader

Abstract

The invention discloses an image decoding method, an image coding method and related equipment. The image decoding method comprises the following steps: in response to the fact that a current frame is not a refresh frame, opening all types of cache information; in response to the fact that the current frame is a refresh frame, only part of types of cache information are started, and the part of types of cache information at least comprise reference frame features; and decoding the current frame by using the opened cache information. According to the scheme, the image decoding accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image encoding and decoding, and in particular to an image decoding method, an image encoding method, and related equipment. Background Technology

[0002] Because video images are large in size, they typically need to be encoded and compressed. The compressed video image data is called a video stream. The video stream can be transmitted to a decoding end via wired or wireless network for decoding and viewing. The entire image encoding and compression process can include prediction, transformation, quantization, and encoding. This reduces the amount of video data, thereby reducing network bandwidth usage during transmission and minimizing storage space.

[0003] However, current image encoding and decoding methods suffer from problems such as low decoding accuracy. Summary of the Invention

[0004] The main technical problem addressed by this application is to provide an image decoding method, an image encoding method, and related equipment that can improve the accuracy of image decoding.

[0005] To address the aforementioned problems, the first aspect of this application provides an image decoding method, which includes:

[0006] In response to the fact that the current frame is not a refresh frame, enable all types of cache information;

[0007] In response to the current frame being a refresh frame, only certain types of cache information are enabled, and these certain types of cache information include at least the features of the reference frame.

[0008] The current frame is decoded using the enabled buffer information.

[0009] To address the aforementioned problems, a second aspect of this application provides an image encoding method, which includes:

[0010] In response to the fact that the current frame is not a refresh frame, enable all types of cache information;

[0011] In response to the current frame being a refresh frame, certain types of cached information are disabled.

[0012] The current frame is encoded using the enabled buffer information.

[0013] To address the aforementioned problems, a third aspect of this application provides an image decoding method, which includes:

[0014] Determine the priori decoded motion features of the current frame;

[0015] The first fused feature is obtained by fusing the prior decoding motion features and the residual features;

[0016] Based on the first fusion feature, the motion feature bitstream of the current frame is decoded to obtain the motion information of the current frame;

[0017] The current frame is decoded based on its motion information.

[0018] To address the aforementioned problems, a fourth aspect of this application provides an image encoding method, the method comprising:

[0019] Determine the priori decoded motion features of the current frame;

[0020] The first fused feature is obtained by fusing the prior decoding motion features and the residual features;

[0021] Based on the first fusion feature, the motion feature bitstream of the current frame is decoded to obtain the motion information of the current frame.

[0022] To address the aforementioned problems, a fifth aspect of this application provides an image decoding method, the method comprising:

[0023] Determine the priori decoding residual features of the current frame;

[0024] The second fused feature is obtained by fusing the super-prior decoding residual features of the current frame and the first alignment feature; the first alignment feature is obtained by motion compensation of the first temporal residual feature.

[0025] Based on the second fusion feature, the residual feature bitstream of the current frame is decoded to obtain the residual information of the current frame;

[0026] The current frame is decoded based on the residual information of the current frame.

[0027] To address the aforementioned problems, a sixth aspect of this application provides an image encoding method, the method comprising:

[0028] Determine the priori decoding residual features of the current frame;

[0029] The second fused feature is obtained by fusing the super-prior decoding residual features of the current frame and the first alignment feature; the first alignment feature is obtained by motion compensation of the first temporal residual feature.

[0030] The residual feature bitstream of the current frame is decoded based on the second fusion feature to obtain the residual information of the current frame.

[0031] To address the aforementioned problems, a seventh aspect of this application provides an image decoding method, the method comprising:

[0032] The decoding residual features of the current frame are optimized using the second temporal residual features and / or prediction features to obtain the first optimized features;

[0033] Residual information is reconstructed from the first optimized feature to obtain the residual information of the current frame;

[0034] The current frame is decoded based on the residual information of the current frame.

[0035] To address the aforementioned problems, an eighth aspect of this application provides an image encoding method, the method comprising:

[0036] The decoding residual features of the current frame are optimized using the second temporal residual features and / or prediction features to obtain the first optimized features;

[0037] The residual information of the current frame is obtained by reconstructing the residual information of the first optimized feature.

[0038] To address the aforementioned problems, a ninth aspect of this application provides an image decoding method, which is applied to video decoding. The method provides at least three consecutive frame number intervals for the video to be decoded, the three consecutive frame number intervals being a first interval, a second interval, and a third interval, respectively. The method includes:

[0039] In response to the current frame's frame number being located in either the first interval or the third interval, feature extraction is performed on the image of a reference frame of the current frame to obtain the reference features of the current frame.

[0040] In response to the fact that the frame number of the current frame is located in the second interval, feature fusion extraction is performed on the image and features of the reference frame to obtain the reference features of the current frame.

[0041] In response to the current frame's frame number being located in the other of the first interval and the third interval, feature extraction is performed on the features of the reference frame to obtain the reference features of the current frame;

[0042] Motion compensation is performed based on the reference features to obtain the prediction information for the current frame;

[0043] The reconstructed image of the current frame is obtained by decoding based on the predicted information.

[0044] To address the aforementioned problems, a tenth aspect of this application provides an image encoding method, wherein the image decoding method is applied to video decoding, and at least three consecutive frame number intervals are set for the video to be decoded, the three consecutive frame number intervals being a first interval, a second interval, and a third interval, respectively, the method comprising:

[0045] In response to the current frame's frame number being located in either the first interval or the third interval, feature extraction is performed on the image of a reference frame of the current frame to obtain the reference features of the current frame.

[0046] In response to the fact that the frame number of the current frame is located in the second interval, feature fusion extraction is performed on the image and features of the reference frame to obtain the reference features of the current frame.

[0047] In response to the current frame's frame number being located in the other of the first interval and the third interval, feature extraction is performed on the features of the reference frame to obtain the reference features of the current frame;

[0048] Motion compensation is performed based on the reference features to obtain the prediction information for the current frame.

[0049] To address the aforementioned problems, the eleventh aspect of this application provides a computer device comprising a memory and a processor coupled to each other, wherein the memory stores program data and the processor executes the program data to implement any step of any of the methods described above.

[0050] To address the aforementioned problems, the twelfth aspect of this application provides a computer-readable storage medium storing program data executable by a processor, the program data being used to implement any step of any of the methods described above.

[0051] In the above-described scheme, when the current frame is a refresh frame, this application can disable some types of cache information, that is, only enable some types of cache information. The enabled types of cache information may include the reconstruction features of the reference frame. In this way, when the refresh frame is decoded, only the enabled cache information can be used to decode the current frame. By setting the refresh frame and making the refresh frame refer to only some types of cache information, that is, appropriately disabling the cache information passed from the previous frame, the error transmission from the previous frame is reduced, and the accumulation of error can be truncated.

[0052] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description

[0053] To more clearly illustrate the technical solutions in this application, the accompanying drawings required in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Among them:

[0054] Figure 1 This is a schematic diagram of the structure of an embodiment of the image encoding and decoding system of this application;

[0055] Figure 2 This is a schematic diagram of the structure of the first embodiment of the end-to-end video compression framework of this application;

[0056] Figure 3 This is a schematic diagram of the structure of the second embodiment of the end-to-end video compression framework of this application;

[0057] Figure 4 This is a flowchart illustrating the first embodiment of the image decoding method of this application;

[0058] Figure 5 This is a schematic diagram of the structure of a motion vector entropy module in the end-to-end video compression framework of this application;

[0059] Figure 6 This is a flowchart illustrating the first embodiment of the image encoding method of this application;

[0060] Figure 7 This is a flowchart illustrating the second embodiment of the image decoding method of this application;

[0061] Figure 8 This is a schematic diagram of the structure of a residual information entropy module in the end-to-end video compression framework of this application;

[0062] Figure 9 This is a flowchart illustrating the second embodiment of the image encoding method of this application;

[0063] Figure 10 This is a schematic diagram of the structure of a compensation module of the residual information entropy module in the end-to-end video compression framework of this application.

[0064] Figure 11 This is a flowchart illustrating the third embodiment of the image decoding method of this application;

[0065] Figure 12 This is a schematic diagram of the structure of a compensation module in an embodiment of the optimization module in the end-to-end video compression framework of this application;

[0066] Figure 13 This is a flowchart illustrating the third embodiment of the image encoding method of this application;

[0067] Figure 14This is a flowchart illustrating the fourth embodiment of the image decoding method of this application;

[0068] Figure 15 This is a schematic diagram of reference information switching in the image encoding and decoding method of this application;

[0069] Figure 16 This is a schematic diagram of the reference information feature extraction process in the image encoding and decoding method of this application;

[0070] Figure 17 This is a flowchart illustrating the fourth embodiment of the image encoding method of this application;

[0071] Figure 18 This is a flowchart illustrating the fifth embodiment of the image decoding method of this application;

[0072] Figure 19 This is a flowchart illustrating the fifth embodiment of the image encoding method of this application;

[0073] Figure 20 This is a schematic diagram of the structure of an embodiment of the computer device of this application;

[0074] Figure 21 This is a schematic diagram of the structure of an embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0075] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0076] The terms "first" and "second" in this application are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.

[0077] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0078] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this document means two or more. Moreover, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of objects. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0079] This application provides the following embodiments, and each embodiment is described in detail below.

[0080] Please see Figure 1 , Figure 1 This is a schematic diagram of the structure of an embodiment of the image encoding and decoding system of this application.

[0081] The image encoding / decoding system 100 includes an encoding end 101 and a decoding end 102. The encoding end 101 and the decoding end 102 can be computer equipment, electronic equipment, etc., and can be any device with processing capabilities, such as a computer, server, mobile phone, tablet, etc. This application does not impose any limitations on this. The encoding end 101 and the decoding end 102 can communicate with each other and can be used to perform encoding and / or decoding operations on images / videos.

[0082] The encoding end 101 can be used to perform preprocessing and encoding / compression steps for images / videos to obtain bitstream data. The encoding end 101 can transmit the bitstream data to the decoding end 102. The decoding end 102 can receive the bitstream data from the encoding end 101 and perform decoding and other related steps involving the bitstream data, as well as steps related to backend vision tasks, such as image / video processing and classification.

[0083] Alternatively, the image encoding and decoding system can perform image encoding and decoding operations through an end-to-end video compression framework.

[0084] In one embodiment, such as Figure 2 As shown, the structure of an end-to-end video compression framework includes, but is not limited to, the following parts:

[0085] 1. First motion information compression and reconstruction:

[0086] (1) Motion estimation: The current frame and the reference frame obtain the first motion information between the frames through the motion estimation network.

[0087] (2) First motion information Encoder: Dimensionally reduce and extract compact information from the first motion information (or the first motion information residual) to obtain the first motion information to be encoded.

[0088] (3) Motion Vector Entropy Module: Obtain the probability of each character appearing in the first motion information to be encoded after quantization, perform arithmetic encoding, and output the first motion information bitstream.

[0089] (4) First motion information Decoder: The first motion information is obtained by upscaling and reconstructing the low-dimensional motion features of the decoded data.

[0090] 2. Motion compensation and temporal prediction: Motion compensation operations such as warping are performed on the features of the reference frame using the first motion information. Relevant information is extracted from the reference frame to obtain the predicted frame of the current frame.

[0091] 3. Residual compression and reconstruction:

[0092] (1) Residual Encoder: Performs dimensionality reduction and compression on the residual information between the current frame and the predicted frame. The residual information includes, but is not limited to, the difference information or concatenation information between the two frames. The purpose is to remove the temporal domain correlation between frames and the spatial domain correlation within frames.

[0093] (2) Residual Entropy Module: Obtain the probability of each character appearing in the quantized residual information to be encoded, perform arithmetic encoding, and output the residual information bitstream.

[0094] (3) Residual Decoder: The low-dimensional residual features obtained from decoding the above bitstream are upsized and reconstructed to obtain residual information. In this process, the residual information is combined with the prediction frame to obtain preliminary reconstruction information. For example, if the residual at the Encoder end is a difference value, the Decoder end will add the prediction frame information to obtain the preliminary reconstruction information of the current frame.

[0095] 4. Frame reconstruction: Further processing of the preliminary reconstruction information yields the reconstructed image and reconstructed features of the current frame. Both are stored in the cache information and can be used as a reference frame for the next frame.

[0096] In another embodiment, the structure of the end-to-end video compression framework can be as follows: Figure 3 As shown.

[0097] Optionally, the end-to-end video compression framework may also include an optimization module to optimize the output of the residual information encoder and / or the residual information entropy module. Optionally, the decoding residual features and / or the residual features to be encoded in the current frame can be optimized using prediction information and / or residual information obtained from decoding previous frames to enhance the residual features, which helps to reduce the bit rate while improving reconstruction quality. Here, residual features can also be called context features, and residual information can also be called context information.

[0098] Optionally, the motion vector entropy module can incorporate current frame features, reference features, and / or residual features from previous frame decoding as input to calculate more accurate probability parameters (mean, variance) and / or quantization parameters, etc.

[0099] Optionally, the current frame information can be introduced as prior information into the residual information entropy module to improve the accuracy of probability parameter prediction. Alternatively, the temporal decoding residual features introduced into the residual information entropy module can be compensated before being used to calculate the probability parameters. This optimizes the temporal information, removes temporal redundancy, enhances feature information, and improves coding performance. The motion information required for compensation can be derived from the decoded motion information. The decoded motion information includes, but is not limited to, the decoded motion features output by the motion vector entropy module and / or the motion information reconstructed by the motion vector decoder.

[0100] The aforementioned temporal decoding residual features refer to the feature information output by the residual information entropy module during the decoding process of previous frames.

[0101] Optionally, for the feature extraction module 2, feature extraction can be performed on the reference frame features, or the reference frame image, or a combination of the reference frame image and features, to obtain the reference features of the current frame.

[0102] Optionally, during the video encoding and decoding process, the reference information in the cache can be refreshed according to a preset refresh cycle, and only at least one cached information can be referenced to truncate the accumulation of errors.

[0103] Optionally, the various solutions mentioned above can be implemented independently or in combination. Specific details of these solutions will be provided below.

[0104] This application provides an image decoding method, an image encoding method, and related equipment to improve the accuracy of image decoding during the encoding and decoding process. In some embodiments, the encoding end 101 and the decoding end 102 of this embodiment can be used to implement any step of the following embodiments.

[0105] Please see Figure 4 , Figure 4This is a flowchart illustrating the first embodiment of the image decoding method of this application. The specific steps of this embodiment can be executed using the decoding end described above. The method may include the following steps:

[0106] S11: Analyze the hyper-prior decoding motion features of the current frame.

[0107] The encoding end can encode the current image frame to obtain the current frame bitstream. The decoding end obtains the current frame bitstream and decodes it to obtain the current reconstructed frame of the current image frame. The current frame bitstream can include a motion bitstream and / or a context bitstream. The motion bitstream can include a motion feature bitstream and / or a priori motion bitstream. The residual bitstream can include a residual feature bitstream and / or a priori decoding residual bitstream. The current frame bitstream can be used to decode and obtain the reconstructed frame of the current image frame, i.e., the decoded image.

[0108] The hyperprior motion bitstream of the current frame can be parsed to obtain the hyperprior decoded motion features of the current frame. Optionally, entropy decoding and hyperprior decoding can be performed on the hyperprior motion bitstream of the current frame to obtain the hyperprior decoded motion features of the current frame.

[0109] S12: Fuse the prior decoding motion features and residual features to obtain the first fused feature.

[0110] Optionally, the hyper-priority decoded motion features and residual features can be fused to obtain a first fused feature, so that the motion feature bitstream of the current frame can be decoded based on the first fused feature to obtain the motion information of the current frame.

[0111] In one implementation, the residual features can be scaled using a first scaling factor to obtain scaled residual features; then, the hyperprior decoded motion features and the scaled residual features are fused to obtain a first fused feature. Optionally, this application executes step S12 through a motion vector entropy module. In this case, after the motion vector entropy module obtains the residual features from the buffer, it can scale the obtained residual features using a pre-stored first scaling factor, and then input the scaled residual features into the fusion module. The first scaling factor pre-stored by the motion vector entropy module can be a learnable parameter obtained during training. The motion vector entropy module can pre-store at least one first scaling factor. When the motion vector entropy module pre-stores multiple first scaling factors, it can use these multiple first scaling factors sequentially. If multiple first scaling factors have been used sequentially, they can be used again sequentially, and so on.

[0112] Of course, in some embodiments, the hyperprior decoded motion features and the unscaled residual features can also be directly fused to obtain the first fused feature.

[0113] Optionally, the image encoding and decoding system can simultaneously input one of the unscaled residual features and the scaled residual features, along with the super-prior decoded motion features, into the fusion module to generate the first fused feature.

[0114] In other embodiments, one of the unscaled residual features and the scaled residual features, along with the prior decoded motion features and the reference features, can be fused to obtain a first fused feature. In this embodiment, the image encoding / decoding system can simultaneously input the prior decoded motion features, the reference features, and the scaled residual features into the fusion module to generate the first fused feature.

[0115] The reference feature can be a reconstructed feature from a frame preceding the current frame, such as a reconstructed feature from a reference frame of the current frame. The image frame corresponding to the reference feature can be any frame; for example, it can be any image frame within a preset number of one-way or two-way frames of the current image frame. It is understandable that, since adjacent frames have the highest correlation, adjacent frames can be prioritized, meaning the reconstructed features of adjacent frames of the current image frame can be used as reference features.

[0116] In addition, the residual features of this embodiment may include the residual features of the current frame and / or the residual features of image frames preceding the current frame.

[0117] Furthermore, the first fusion feature may include probability parameters (e.g., mean and / or variance) and / or quantization parameters of the motion feature to be encoded. The probability parameters can be used for entropy decoding of the motion feature bitstream, etc. The quantization parameters can be used for dequantization of the motion feature bitstream to recover the decoded motion features. Thus, by introducing spatial domain feature information (current frame features, residual features, and reference features, etc.), the probability parameters of the motion feature to be encoded are predicted more effectively.

[0118] Optionally, the fusion module of this application may include at least one of a residual network, a recurrent network, or an attention network. For example, the fusion module may include a convolutional network and a residual network; the convolutional network can be used for channel dimensionality reduction, and the residual network can be used to obtain residuals, i.e., difference information. It is understood that the fusion module of this application may also be other network structures, and this application does not limit the fusion module.

[0119] S13: Based on the first fusion feature, the motion feature bitstream of the current frame is decoded to obtain the motion information of the current frame.

[0120] After fusing the hyper-priority decoding motion features and residual features to obtain the first fused feature, the motion feature bitstream of the current frame can be decoded based on the first fused feature to obtain the motion information of the current frame.

[0121] The motion information of the current frame can be obtained by performing entropy decoding, inverse quantization, and / or motion information reconstruction on the motion feature bitstream of the current frame based on the first fusion feature.

[0122] As described above, the first fusion feature may include probability parameters (e.g., mean and / or variance, etc.) and / or quantization parameters of the motion features to be encoded. In one embodiment, such as Figure 5 As shown, entropy decoding can be performed on the motion feature bitstream based on the quantization parameters and / or variance in the first fusion feature to obtain entropy-decoded motion features; the entropy-decoded motion features can be dequantized based on the mean in the first fusion feature to obtain dequantized motion features; the dequantized motion features and the mean in the first fusion feature are added together to obtain decoded motion features; motion information reconstruction processing is performed on the decoded motion features to obtain the motion information of the current frame. Optionally, in other embodiments, some operations in entropy decoding, dequantization, and mean addition may not be performed.

[0123] Corresponding to the first embodiment of the image decoding method, this application also provides an image encoding method according to the first embodiment, such as... Figure 6 As shown, the image encoding method includes the following steps.

[0124] S21: Determine the hyperprior decoding motion features of the current frame.

[0125] The hyperprior motion bitstream of the current frame can be parsed to obtain the hyperprior decoded motion features of the current frame. Optionally, entropy decoding and hyperprior decoding can be performed on the hyperprior motion bitstream of the current frame to obtain the hyperprior decoded motion features of the current frame.

[0126] For the encoding end, before step S21, the motion features to be encoded output by the motion vector encoder can be processed by a priori coding to obtain the a priori motion bitstream of the current frame.

[0127] Optionally, a priori coding process can be performed on the motion features to be encoded based on the current frame features and / or residual features to obtain the priori motion bitstream of the current frame. In one embodiment, the current frame features and the motion features to be encoded can be input into a priori encoder to process the current frame features and the motion features to be encoded through the priori encoder to obtain the priori coding result of the current frame; the priori coding result of the current frame can then be quantized and / or entropy encoded to obtain the priori motion bitstream of the current frame.

[0128] Here, "current frame features" refers to the features obtained by extracting features from the current frame image using the feature extraction module. For example, using... Figure 2 and Figure 3 The feature extraction module 1 in the end-to-end video compression framework shown performs feature extraction on the current frame image to obtain the current frame features.

[0129] In one implementation, the residual features can be scaled using a first scaling factor to obtain scaled residual features. Then, based on the current frame features and / or the scaled residual features, a priori coding process is performed on the motion features to be encoded to obtain the priori motion bitstream of the current frame. Optionally, this application executes step S12 through a motion vector entropy module. In this case, after the motion vector entropy module obtains the residual features from the buffer, it can scale the obtained residual features using a pre-stored first scaling factor, and then input the scaled residual features into the internal priori coding module. This allows the priori coding module to perform a priori coding on the motion features to be encoded based on the current frame features and / or the scaled residual features to obtain the priori motion bitstream of the current frame.

[0130] Of course, in some embodiments, the motion features to be encoded can also be processed by super-prior coding based on the current frame features and / or the unscaled residual features to obtain the super-prior motion bitstream of the current frame.

[0131] In other implementations, the motion features to be encoded can be directly processed using a priori coding to obtain the a priori motion bitstream of the current frame. That is, during the a priori coding process of the motion features to be encoded, it is possible to choose to refer to the features of the current frame or not.

[0132] S22: The super-prior decoding motion features and residual features are fused to obtain the first fused feature.

[0133] See step S12 for details, which will not be repeated here.

[0134] S23: Based on the first fusion feature, the motion feature bitstream of the current frame is decoded to obtain the motion information of the current frame.

[0135] See step S13 for details, which will not be repeated here.

[0136] Optionally, for an image encoding and decoding system, after obtaining the motion features (i.e., the decoding motion features mentioned above) and / or motion information of the current frame, the motion features and / or motion information can be input into a buffer so that the motion features and / or motion information in the buffer can be used as a reference for encoding and decoding subsequent frames.

[0137] At least some of the processing steps S11, S12, and S13, and at least some of the processing steps S21, S22, and S23, can be performed in the motion vector entropy module of the end-to-end video compression framework. Please refer to [link / reference]. Figure 5 , Figure 5This is a schematic diagram of the framework of an embodiment of the motion vector entropy module of this application. The motion vector entropy module includes a quantization module 1, an entropy encoding module 1, an entropy decoding module 1, an inverse quantization module 1, a fusion network 1, a super-prior encoding module 1, a super-prior decoding module 1, a quantization module 2, an entropy encoding module 2, and an entropy decoding module 2.

[0138] The system comprises the following modules: **Super-Prior Encoding 1:** This module performs super-prior encoding on the motion features to be encoded, obtaining the super-prior encoding result for the current frame. During super-prior encoding, features of the current frame can be selectively referenced. **Quantization 2:** This module quantizes the super-prior encoding result of the current frame. **Entropy Encoding 2:** This module performs entropy encoding on the output of the quantization module, obtaining the super-prior motion bitstream. **Entropy Decoding 2:** This module performs entropy decoding on the super-prior motion bitstream. **Super-Prior Decoding 1:** This module performs super-prior decoding on the output of the entropy decoding module, obtaining the super-prior decoded motion features. **Fusion 1:** This module processes the super-prior decoded motion features to obtain the probability parameters and / or quantization parameters of the motion features to be encoded. During the fusion 1 process, residual features and / or reference features can be selectively introduced. The introduced residual features can be residual features of the current frame or residual features of image frames preceding the current frame. The Quantization 1 module is used to quantize the first difference feature, which can be the difference between the motion feature to be encoded and the mean data in the first fused feature. Optionally, the Quantization 1 module can quantize the first difference feature based on the quantization parameters in the first fused feature. The Entropy Encoding 1 module is used to entropy encode the output of the Quantization 1 module to obtain the motion feature bitstream. Optionally, the Entropy Encoding 1 module can be used to entropy encode the output of the Quantization 1 module based on the variance in the first fused feature. The Entropy Decoding 1 module is used to entropy decode the motion feature bitstream. Optionally, the Entropy Decoding 1 module can be used to entropy decode the motion feature bitstream based on the variance in the first fused feature. The Inverse Quantization 1 module is used to inverse quantize the output of the Entropy Decoding 1 module to obtain the inverse quantization result of the motion feature to be encoded. Optionally, the Inverse Quantization 1 module can inverse quantize the output of the Entropy Decoding 1 module based on the quantization parameters in the first fused feature. The inverse quantization result can be added to the mean in the first fused feature to obtain the decoded motion feature.

[0139] Please see Figure 7 , Figure 7 This is a flowchart illustrating a second embodiment of the image decoding method of this application. The specific steps of this embodiment can be executed using the decoding end described above. The method may include the following steps:

[0140] S31: Analyze the super-prior decoding residual features of the current frame.

[0141] The super-prior residual bitstream of the current frame can be parsed to obtain the super-prior decoding residual features of the current frame. Optionally, entropy decoding and super-prior decoding can be performed on the super-prior residual bitstream of the current frame to obtain the super-prior decoding residual features of the current frame.

[0142] S32: Fuse the super-prior decoding residual features and the first alignment features to obtain the second fused features.

[0143] In this regard, considering that the first temporal residual feature and the super-prior decoding information have motion offset, motion compensation can be performed on the first temporal residual feature to align the motion-compensated first temporal residual feature with the super-prior decoding residual feature of the current frame, thereby improving the accuracy of the second fused feature.

[0144] As shown above, the first alignment feature is a feature obtained by performing motion compensation on the first temporal residual feature.

[0145] The first temporal residual feature can be the residual feature of image frames preceding the current frame. Furthermore, the second temporal residual feature can be the residual feature generated during the decoding process of previously decoded frames preceding the current frame, such as the residual feature of a reference frame for the current frame. Optionally, the image encoding / decoding system can obtain the first temporal residual feature from a buffer.

[0146] In one implementation, the first temporal residual feature can be scaled using a second scaling factor to obtain a scaled first temporal residual feature; then, motion compensation is performed on the scaled first temporal residual feature to obtain a first alignment feature. Optionally, this application executes step S32 through the residual information entropy module. In this case, after the residual information entropy module obtains the first temporal residual feature from the buffer, it can scale the obtained first temporal residual feature using a pre-stored second scaling factor, and then input the scaled first temporal residual feature into the fusion module. The second scaling factor pre-stored by the residual information entropy module can be a learnable parameter obtained during training. The motion vector entropy module can pre-store at least one second scaling factor. If the motion vector entropy module pre-stores multiple second scaling factors, it can use these multiple second scaling factors sequentially. If all multiple second scaling factors have been used sequentially, they can be used again sequentially, and so on. Optionally, the second scaling factor pre-stored by the residual information entropy module can be the same as or different from the first scaling factor; no restriction is placed here.

[0147] Of course, in some embodiments, motion compensation can also be performed directly on the unscaled first temporal residual features to obtain the first alignment features.

[0148] In one implementation, motion compensation can be performed on either the unscaled first temporal residual feature or the scaled first temporal residual feature using first motion information to obtain a first alignment feature. Optionally, the first motion information can be transformed to obtain second motion information, which can be used to perform dimensional transformation and / or information adjustment on the first motion information. Then, the second motion information can be used to perform motion compensation on the first temporal residual feature to obtain the first alignment feature. In other embodiments, the first motion information may not be transformed; that is, the first motion information can be directly used to perform motion compensation on the first temporal residual feature to obtain the first alignment feature. Optionally, the first motion information may include motion information output by the motion decoder, and / or motion features output by the motion vector entropy module, etc. Further, the first motion information may include, but is not limited to, motion information of the current frame output by the motion decoder, motion information of image frames preceding the current frame output by the motion decoder, motion features of the current frame output by the motion vector entropy module (i.e., decoded motion features of the current frame), and / or motion features of image frames preceding the current frame output by the motion vector entropy module.

[0149] In another implementation, the third motion information can be determined by matching the first temporal residual features and the prior decoding residual features. Optionally, the first temporal residual features and the prior decoding residual features can be input into the fine-tuning module to process the first temporal residual features and the prior decoding residual features to obtain the third motion information; the third motion information is then used to perform motion compensation on the first temporal residual features to obtain the first alignment feature.

[0150] In another implementation, the third motion information can be determined by matching the scaled first temporal residual features and the super-prior decoding residual features. Optionally, the scaled first temporal residual features and the super-prior decoding residual features can be input into a fine-tuning module to process the scaled first temporal residual features and the super-prior decoding residual features to obtain the third motion information; the third motion information is then used to perform motion compensation on the scaled first temporal residual features to obtain the first alignment feature.

[0151] In another implementation, a first alignment feature can be obtained by performing motion compensation on either the unscaled or scaled first temporal residual feature based on the first motion information and the prior decoding residual feature. Optionally, one of the first and second motion information, the first temporal residual feature, and the prior decoding residual feature can be input into a fine-tuning module. The fine-tuning module processes the first and second motion information, the first temporal residual feature, and the prior decoding residual feature to obtain more accurate third motion information. The third motion information is then used to perform motion compensation on the first temporal residual feature to obtain the first alignment feature. Alternatively, one of the first motion information and the second motion information, the scaled first temporal residual feature, and the prior decoding residual feature can be input into the fine-tuning module. The fine-tuning module processes one of the first motion information and the second motion information, the scaled first temporal residual feature, and the prior decoding residual feature to obtain a more accurate third motion information. The third motion information is then used to perform motion compensation on the scaled first temporal residual feature to obtain the first alignment feature.

[0152] The aforementioned fine-tuning module may include neural network structures such as residual networks and / or attention networks. For example, the fine-tuning module may include a residual network. It is understood that the fine-tuning module of this application may also be other network structures, and this application does not limit the fine-tuning module.

[0153] In the above-described steps for motion compensation of the first temporal residual features, one or a combination of at least two alignment methods, such as interpolation-based warp alignment and deformable convolution, can be used to perform motion compensation of the first temporal residual features.

[0154] After obtaining the first alignment feature, the super-prior decoding residual feature and the first alignment feature can be fused to obtain the second fused feature, so that the residual feature bitstream of the current frame can be decoded based on the second fused feature to obtain the residual information of the current frame.

[0155] Optionally, the image encoding and decoding system can simultaneously input the super-prior decoding residual features and the first alignment features into the fusion module to generate the second fused features.

[0156] In other embodiments, the super-prior decoding residual features, the first alignment features, and the prediction features can be fused to obtain a second fused feature. In this embodiment, the image encoding and decoding system can simultaneously input the super-prior decoding residual features, the prediction features of the current frame, and the first alignment features into the fusion module to generate the second fused feature.

[0157] The second fusion feature can include probability parameters (e.g., mean and / or variance) and / or quantization parameters of the residual features of the current frame. The probability parameters can be used for entropy decoding of the residual feature bitstream. The quantization parameters can be used for dequantization of the residual feature bitstream to recover the decoded residual features. Thus, spatial domain feature information (current frame features, first temporal residual features, etc.) is introduced as a priori and a compensation module (introducing motion information) to assist in predicting the probability parameters of the residual information (contextual information).

[0158] S33: Based on the second fusion feature, the residual feature bitstream of the current frame is decoded to obtain the residual information of the current frame.

[0159] After fusing the super-prior decoding residual features and the first alignment features to obtain the second fused features, the residual feature bitstream of the current frame can be decoded based on the second fused features to obtain the residual information of the current frame.

[0160] The residual information of the current frame can be obtained by performing entropy decoding, inverse quantization, and / or residual information reconstruction on the residual feature bitstream of the current frame based on the second fusion feature.

[0161] As described above, the second fusion feature may include probability parameters (e.g., mean and / or variance, etc.) and / or quantization parameters of the residual features of the current frame. In one embodiment, such as Figure 8 As shown, entropy decoding can be performed on the residual feature bitstream based on the quantization parameters and / or variance in the second fusion feature to obtain entropy-decoded residual features; the entropy-decoded residual features can be dequantized based on the mean in the second fusion feature to obtain dequantized residual features; the dequantized residual features and the mean in the second fusion feature are added together to obtain decoded residual features; residual information reconstruction processing is performed on the decoded residual features to obtain the residual information of the current frame. Optionally, in other embodiments, some operations in entropy decoding, dequantization, and mean addition may not be performed.

[0162] In image decoding methods, the residual information and / or motion information of the current frame can be used to construct the reconstructed image of the current frame.

[0163] Corresponding to the second embodiment of the image decoding method, this application also provides an image encoding method according to the second embodiment, such as... Figure 9 As shown, the image encoding method includes the following steps.

[0164] S41: Analyze the super-prior decoding residual features of the current frame.

[0165] The super-prior residual bitstream of the current frame can be parsed to obtain the super-prior decoding residual features of the current frame. Optionally, entropy decoding and super-prior decoding can be performed on the super-prior residual bitstream of the current frame to obtain the super-prior decoding residual features of the current frame.

[0166] For the encoding end, before step S41, the residual features to be encoded output by the residual information encoder can be processed by a priori coding to obtain the a priori residual bitstream of the current frame.

[0167] Optionally, a priori coding process can be performed on the residual features to be encoded based on the features of the current frame to obtain the priori residual bitstream of the current frame. In one embodiment, the features of the current frame and the residual features to be encoded can be input into a priori encoder to process the features of the current frame and the residual features to be encoded to obtain the priori coding result of the current frame; the priori coding result of the current frame is then quantized and / or entropy encoded to obtain the priori residual bitstream of the current frame.

[0168] In another implementation, the super-prior coding process can be performed on the residual features to be encoded based on the second alignment feature to obtain the super-prior residual bitstream of the current frame. In one embodiment, the second alignment feature and the residual features to be encoded can be input into the super-prior encoder to process the second alignment feature and the residual features to be encoded to obtain the super-prior coding result of the current frame; the super-prior coding result of the current frame is then quantized and / or entropy encoded to obtain the super-prior residual bitstream of the current frame.

[0169] Similar to the first alignment feature shown in step S32, the second alignment feature is obtained by motion compensation of the first temporal residual feature. Preferably, during the encoding and decoding process of the current frame, the first temporal residual feature used at the encoding end and the first temporal residual feature used at the decoding end can be residual features of the same image, or they can be residual features of different images.

[0170] Optionally, in one implementation, the first temporal residual feature can be scaled using a second scaling factor to obtain scaled residual features; then, motion compensation is performed on the scaled first temporal residual feature to obtain the second alignment feature.

[0171] Of course, in some embodiments, motion compensation can also be performed directly on the unscaled first temporal residual features to obtain the second alignment features.

[0172] The second alignment feature can be obtained in a variety of ways, without limitation, such as the three methods described below.

[0173] The first method involves using the first motion information to perform motion compensation on the first temporal residual feature to obtain the second alignment feature. The specific steps are shown in step S32 and will not be elaborated upon here.

[0174] The second approach is to determine the fourth motion information by matching the first time-domain residual features and the residual features to be encoded. Optionally, the first time-domain residual features and the residual features to be encoded can be input into the fine-tuning module to process the first time-domain residual features and the residual features to be encoded to obtain the fourth motion information; the fourth motion information is then used to perform motion compensation on the first time-domain residual features to obtain the second alignment feature.

[0175] The third approach involves performing motion compensation on the first temporal residual features based on the first motion information and the residual features to be encoded, thereby obtaining the second alignment feature. Optionally, such as... Figure 10 As shown, one of the first motion information and the second motion information, the first time-domain residual feature, and the residual feature to be encoded can be input into the fine-tuning module. The fine-tuning module processes one of the first motion information and the second motion information, the first time-domain residual feature, and the residual feature to be encoded to obtain a more accurate fourth motion information. The fourth motion information is then used to perform motion compensation on the first time-domain residual feature to obtain the second alignment feature.

[0176] The aforementioned fine-tuning module may include neural network structures such as residual networks and / or attention networks. For example, the fine-tuning module may include a residual network. It is understood that the fine-tuning module of this application may also be other network structures, and this application does not limit the fine-tuning module. In the aforementioned step of motion compensation for the first temporal residual features, one or a combination of at least two alignment methods, such as interpolation-based warp alignment and deformable convolution, may be used to perform motion compensation for the first temporal residual features.

[0177] In another implementation, a priori coding process can be performed on the residual features to be encoded based on the current frame image and the second alignment feature to obtain the priori residual bitstream of the current frame. In one embodiment, the current frame features, the second alignment feature, and the residual features to be encoded can be input into a priori encoder to process the current frame features, the second alignment feature, and the residual features to be encoded, thereby obtaining the priori coding result of the current frame; the priori coding result of the current frame is then quantized and / or entropy encoded to obtain the priori residual bitstream of the current frame.

[0178] In other implementations, the residual features to be encoded can be directly processed by prior coding to obtain the prior residual bitstream of the current frame. That is, during the prior coding process of the residual features to be encoded, it is possible to refer to the features of the current frame and / or the second alignment features, or to choose not to refer to the features of the current frame and the second alignment features.

[0179] S42: Fuse the super-prior decoding residual features and the first alignment features to obtain the second fused features.

[0180] See step S32 for details, which will not be repeated here.

[0181] S43: Based on the second fusion feature, the residual feature bitstream of the current frame is decoded to obtain the residual information of the current frame.

[0182] See step S33 for details, which will not be repeated here.

[0183] Optionally, for an image encoding and decoding system, after obtaining the residual features (i.e., the decoding residual features mentioned above) and / or residual information of the current frame, the residual features and / or residual information can be input into a buffer so that the residual features and / or residual information of the buffer can be used as a reference for encoding and decoding subsequent frames.

[0184] At least some of the processing steps S31, S32, and S33, and at least some of the processing steps S41, S42, and S43, can be performed in the residual information entropy module of the end-to-end video compression framework. Please refer to [link / reference]. Figure 8 , Figure 8 This is a schematic diagram of the framework of an embodiment of the residual information entropy module of this application. The residual information entropy module includes a quantization module 3, an entropy encoding module 3, an entropy decoding module 3, an inverse quantization module 3, a fusion network 2, a super-prior encoding module 2, a super-prior decoding module 2, a quantization module 4, an entropy encoding module 4, and an entropy decoding module 4.

[0185] The super-prior coding module 2 is used to perform super-prior coding on the residual features to be encoded, obtaining the super-prior residual coding result of the current frame. During the super-prior coding process, the current frame features and / or a second alignment feature can be selectively referenced. The quantization module 4 performs quantization processing on the super-prior residual coding result of the current frame. The entropy coding module 4 performs entropy coding on the output of the quantization module 4, obtaining the super-prior residual bitstream. The entropy decoding module 4 performs entropy decoding on the super-prior residual bitstream. The super-prior coding module 2 performs super-prior decoding on the output of the entropy decoding module 4, obtaining the super-prior decoded residual features. The fusion module 2 (i.e., the fusion module in the above embodiment) processes the super-prior decoded residual features to obtain the probability parameters of the residual features to be encoded. During the processing of the fusion module 2, the second alignment feature and / or motion information can be selectively introduced. The first temporal residual feature corresponding to the introduced second alignment feature can be the residual feature of image frames prior to the current frame. The quantization module 3 is used to quantize the second difference feature, which can be the difference between the residual feature to be encoded and the mean data in the second fused feature. The entropy encoding module 3 is used to entropy encode the output of the quantization module 3 to obtain the residual feature bitstream. Optionally, the entropy encoding module 3 can be used to entropy encode the output of the quantization module 3 based on the variance in the second fused feature. The entropy decoding module 3 is used to entropy decode the residual feature bitstream. Optionally, the entropy decoding module 3 can be used to entropy decode the residual feature bitstream based on the variance in the second fused feature. The dequantization module 3 is used to dequantize the output of the entropy decoding module 3 to obtain the dequantized result of the residual feature to be encoded. The dequantized result can be added to the mean in the second fused feature to obtain the decoded residual feature.

[0186] Please see Figure 11 , Figure 11 This is a flowchart illustrating a third embodiment of the image decoding method of this application. The specific steps of this embodiment can be executed using the decoding end described above. The method may include the following steps:

[0187] S51: Obtain the decoding residual features of the current frame.

[0188] Optionally, the image decoding method of the second embodiment can be used to determine the decoding residual features of the current frame. That is, the first alignment feature and the super-prior decoding residual features can be fused to obtain the second fused feature. Then, based on the second fused feature, the residual feature bitstream can be entropy decoded and / or dequantized to obtain the decoding residual features of the current frame.

[0189] In other implementations, the first alignment feature may not be referenced when determining the probability parameters of the residual features. For example, in one embodiment, the predicted features of the current frame and the super-prior decoding residual features can be fused to obtain a second fused feature; then, the decoding residual features of the current frame can be obtained based on the second fused feature. In another embodiment, the super-prior decoding residual features can be directly processed by the fusion module to obtain the second fused feature; then, the decoding residual features of the current frame can be obtained based on the second fused feature.

[0190] S52: Optimize the decoding residual features using the second time-domain residual features and / or prediction features to obtain the first optimized features.

[0191] After obtaining the decoding residual features of the current frame, the decoding residual features can be optimized using the second temporal residual features and / or prediction features to obtain the first optimized features. These first optimized features are then used to reconstruct residual information to obtain the residual information of the current frame. The second temporal residual features can be residual features generated during the decoding process of previously decoded frames.

[0192] In one implementation, the second time-domain residual feature can be scaled using a third scaling factor to obtain scaled residual features; then, the decoded residual feature can be optimized using the scaled second time-domain residual feature and / or prediction feature to obtain the first optimized feature.

[0193] Of course, in some embodiments, the decoding residual features can also be directly optimized using the unscaled second temporal residual features and / or prediction features to obtain the first optimized features.

[0194] In one implementation, the predicted features can be used to optimize the decoded residual features to obtain a first optimized feature. In a specific example, the predicted features and the decoded residual features can be fused to obtain the first optimized feature. In another specific example, the predicted features can be transformed, such as by adjusting the dimensions and / or information of the predicted features to obtain transformed predicted features; the transformed predicted features and the decoded residual features can then be fused to obtain the first optimized feature.

[0195] In another implementation, the decoding residual features can be optimized using the second time-domain residual features to obtain the first optimized features.

[0196] In a specific example, the second time-domain residual feature and the decoding residual feature can be directly fused to obtain the first optimized feature.

[0197] In another specific example, motion compensation can be performed on the second temporal residual feature to obtain the third alignment feature; then the third alignment feature and the decoding residual feature can be fused to obtain the first optimized feature.

[0198] Similar to the first alignment feature shown in step S32, the third alignment feature is obtained by motion compensation of the second temporal residual feature. Preferably, during the encoding and decoding process of the current frame, the second temporal residual feature and the first temporal residual feature can be residual features of the same image or residual features of different images, without limitation.

[0199] The third alignment feature can be obtained in a variety of ways, without limitation, such as the three methods described below.

[0200] The first method involves using the first motion information to perform motion compensation on the second temporal residual features to obtain the third alignment feature. The specific steps are shown in step S32 and will not be elaborated upon here.

[0201] The second method involves matching the second time-domain residual features and the decoding residual features to determine the fifth motion information. Optionally, the second time-domain residual features and the decoding residual features can be input into the fine-tuning module to process the second time-domain residual features and the decoding residual features to obtain the fifth motion information. The fifth motion information is then used to perform motion compensation on the second time-domain residual features to obtain the third alignment feature.

[0202] The third method involves performing motion compensation on the second temporal residual features based on the first motion information and the decoded residual features to obtain the third alignment feature. Optionally, such as... Figure 12 As shown, one of the first motion information and the second motion information, the second time-domain residual feature, and the decoding residual feature can be input into the fine-tuning module. The fine-tuning module processes one of the first motion information and the second motion information, the second time-domain residual feature, and the decoding residual feature to obtain a more accurate fifth motion information. The fifth motion information is then used to perform motion compensation on the second time-domain residual feature to obtain the third alignment feature.

[0203] The aforementioned fine-tuning module may include neural network structures such as residual networks and / or attention networks. For example, the fine-tuning module may include a residual network. It is understood that the fine-tuning module of this application may also be other network structures, and this application does not limit the fine-tuning module. In the aforementioned step of motion compensation for the second temporal residual features, one or a combination of at least two alignment methods, such as interpolation-based warp alignment and deformable convolution, may be used. For example, warp + convolution may be used to perform motion compensation for the second temporal residual features.

[0204] In another implementation, the decoding residual features can be optimized based on one of the temporal residual features and the third alignment features, and one of the prediction features and the transform prediction features, to obtain optimized features. In a specific example, such as... Figure 12 As shown, the third alignment feature, transform prediction feature, and decoding residual feature can be fused to obtain the optimized feature.

[0205] In the above implementation, a fusion module can be used to fuse the decoding residual features, the second temporal residual features, the third alignment features, the prediction features, and / or the transform prediction features.

[0206] S53: Reconstruct the residual information of the first optimized feature to obtain the residual information of the current frame.

[0207] After obtaining the first optimized feature based on step S52, residual information can be reconstructed from the first optimized feature to obtain the residual information of the current frame.

[0208] Optionally, the residual information decoder can be used to reconstruct the residual information of the first optimized feature to obtain the residual information of the current frame. That is, the first optimized feature can be input into the residual information decoder to reconstruct the residual information of the first optimized feature.

[0209] After obtaining the residual information of the current frame, the reconstructed image of the current frame can be obtained based on the residual information.

[0210] The steps for reconstructing the current frame based on residual information may include: calculating the reconstructed image of the current frame based on the residual information and the prediction information of the current frame.

[0211] Optionally, such as Figure 3 As shown, the first optimization module can optimize the decoded residual features output by the residual information entropy module using the second temporal residual features and / or prediction features to obtain the first optimized features; then, the residual information decoder can reconstruct the residual information of the current frame by reconstructing the first optimized features output by the first optimization module.

[0212] Corresponding to the third embodiment of the image decoding method, this application also provides an image encoding method according to the third embodiment, such as... Figure 13 As shown, the image encoding method includes the following steps.

[0213] S61: Obtain the decoding residual features of the current frame.

[0214] The residual feature bitstream of the current frame can be parsed to obtain the decoding residual features of the current frame. Optionally, the residual feature bitstream of the current frame can be entropy decoded and / or dequantized to obtain the decoding residual features of the current frame.

[0215] For the encoding end, before step S61, the residual features to be encoded output by the residual information encoder can be quantized and / or entropy encoded to obtain the residual feature bitstream of the current frame. In a specific example, the residual features to be encoded output by the residual information encoder can be directly quantized and / or entropy encoded to obtain the residual feature bitstream of the current frame. In another specific example, the residual features to be encoded output by the residual information encoder can be optimized first using the second temporal residual features and / or prediction features to obtain the second optimized features; the second optimized features are then input into the residual information entropy module to quantize and / or entropy encode the second optimized features to obtain the residual feature bitstream of the current frame. Optionally, as... Figure 3 As shown, the residual features to be encoded output by the residual information encoder can be optimized by the second optimization module using the second temporal residual features and / or prediction features to obtain the second optimized features; then, the second optimized features output by the second optimization module are quantized and / or entropy encoded by the residual information entropy module to obtain the residual feature bitstream of the current frame.

[0216] In addition, in one implementation, the second temporal residual feature can be scaled using a third scaling factor to obtain scaled residual features; then, the scaled second temporal residual feature and / or prediction feature are used to optimize the residual feature to be encoded to obtain the second optimized feature.

[0217] Of course, in some embodiments, the unscaled second temporal residual features and / or prediction features can also be directly used to optimize the residual features to be encoded to obtain the second optimized features.

[0218] In one implementation, the predicted features can be used to optimize the residual features to be encoded, resulting in a second optimized feature. In a specific example, the predicted features and the residual features to be encoded can be fused to obtain the second optimized feature. In another specific example, the predicted features can be transformed, such as by adjusting their dimensions and / or information, to obtain transformed predicted features; these transformed predicted features and the residual features to be encoded are then fused to obtain the second optimized feature.

[0219] In another implementation, the second temporal residual feature can be used to optimize the residual feature to be encoded, thus obtaining the second optimized feature.

[0220] In a specific example, the second temporal residual feature and the residual feature to be encoded can be directly fused to obtain the second optimized feature.

[0221] In another specific example, motion compensation can be performed on the second temporal residual feature to obtain the fourth alignment feature; then the fourth alignment feature and the residual feature to be encoded can be fused to obtain the second optimized feature.

[0222] Similar to the first alignment feature shown in step S32, the fourth alignment feature is obtained by motion compensation of the second temporal residual feature. Preferably, during the encoding and decoding process of the current frame, the second temporal residual feature and the first temporal residual feature can be residual features of the same image or residual images of different images, without limitation.

[0223] The fourth alignment feature can be obtained in a variety of ways, without limitation, such as the three methods described below.

[0224] The first method involves using the first motion information to perform motion compensation on the second temporal residual feature to obtain the fourth alignment feature. The specific steps are shown in step S52 and will not be elaborated upon here.

[0225] The second method involves matching the second temporal residual features with the residual features to be encoded to determine the fifth motion information. Optionally, the second temporal residual features and the residual features to be encoded can be input into a fine-tuning module to process the second temporal residual features and the residual features to be encoded to obtain the fifth motion information. The fifth motion information is then used to perform motion compensation on the second temporal residual features to obtain the fourth alignment feature.

[0226] The third approach involves using the first motion information and the residual feature to be encoded to perform motion compensation on the second temporal residual feature to obtain the fourth alignment feature. Optionally, as shown in the figure, one of the first and second motion information, the second temporal residual feature, and the residual feature to be encoded can be input into the fine-tuning module. The fine-tuning module processes one of the first and second motion information, the second temporal residual feature, and the residual feature to be encoded to obtain a more accurate fifth motion information. The fifth motion information is then used to perform motion compensation on the second temporal residual feature to obtain the fourth alignment feature.

[0227] The aforementioned fine-tuning module may include neural network structures such as residual networks and / or attention networks. For example, the fine-tuning module may include a residual network. It is understood that the fine-tuning module of this application may also be other network structures, and this application does not limit the fusion module. In the aforementioned step of motion compensation for the second temporal residual features, one or a combination of at least two alignment methods, such as interpolation-based warp alignment and deformable convolution, may be used. For example, warp + convolution may be used to perform motion compensation for the second temporal residual features.

[0228] In another implementation, the residual features to be encoded can be optimized based on one of the temporal residual features and the fourth alignment feature, and one of the prediction features and the transform prediction features, to obtain optimized features. In a specific example, as shown in the figure, the fourth alignment feature, the transform prediction feature, and the residual features to be encoded can be fused to obtain optimized features.

[0229] In the above implementation, a fusion module can be used to fuse the residual features to be encoded, the temporal residual features, the fourth alignment features, the prediction features, and / or the transform prediction features.

[0230] S62: Optimize the decoding residual features using the second time-domain residual features and / or prediction features to obtain the first optimized features.

[0231] See step S52 for details, which will not be repeated here.

[0232] S63: Reconstruct the residual information of the first optimized feature to obtain the residual information of the current frame.

[0233] See step S53 for details, which will not be repeated here.

[0234] Optionally, for an image encoding and decoding system, after obtaining the residual information of the current frame, the residual information can be input into a buffer so that the residual information in the buffer can be used as a reference for encoding and decoding subsequent frames.

[0235] Please see Figure 14 , Figure 14 This is a flowchart illustrating the fourth embodiment of the image decoding method of this application. This image decoding method can be applied to video decoding, wherein at least three consecutive frame number intervals can be set for the video to be decoded, and these three consecutive frame number intervals can be, in sequence, a first interval, a second interval, and a third interval. The specific steps of this embodiment can be executed using the aforementioned decoding end. The method may include the following steps:

[0236] S71: In response to the current frame's frame number being located in either the first interval or the third interval, feature extraction is performed on the image of the reference frame of the current frame to obtain the reference features of the current frame.

[0237] S72: In response to the frame number of the current frame being located in the second interval, feature fusion extraction is performed on the image and features of the reference frame of the current frame to obtain the reference features of the current frame.

[0238] S73: In response to the current frame's frame number being located in either the first interval or the third interval, feature extraction is performed on the features of the reference frame of the current frame to obtain the reference features of the current frame.

[0239] As described in steps S71, S72, and S73, at least three consecutive frame number intervals are set for the video to be decoded. These three consecutive frame number intervals can be the first interval, the second interval, and the third interval, respectively. If the frame number is in the first interval, feature extraction can be performed on the image / features of the reference frame of the current frame to obtain the reference features of the current frame. If the frame number is in the second interval, feature fusion extraction can be performed on the image and features of the reference frame of the current frame to obtain the reference features of the current frame. If the frame number is in the third interval, feature extraction can be performed on the features / images of the reference frame of the current frame to obtain the reference features of the current frame. In this way, when the reference information changes from an image to a feature, or from a feature to an image, the image and features of the reference frame are fused to serve as intermediate reference information to avoid information loss caused by a large reference span when the reference information changes directly from an image to a feature, or from a feature to an image. This can reduce information loss caused by sudden changes in the reference relationship.

[0240] It is understandable that "the remaining one between the first interval and the third interval" in step S73 refers to the one remaining after selecting one of the first interval and the third interval in step S71. For example, if "the remaining one between the first interval and the third interval" in step S71 refers to the first interval, then "the remaining one between the first interval and the third interval" in step S73 refers to the third interval. Or, for another example, if "the remaining one between the first interval and the third interval" in step S71 refers to the third interval, then "the remaining one between the first interval and the third interval" in step S73 refers to the first interval.

[0241] For example, such as Figure 15 As shown, assuming there are L+M+N frames to decode, L video frames have reference information consisting of the image (or features) of the reference frame, and N frames use the features (or image) of the reference frame as the reference information for the current frame. Between L and N (when there is a significant change in reference information), the M frames are first treated with a combination of the image and features of the reference frame as the reference information for the current frame. In one example, L=2, meaning the reference information for the i-th p-th frame is the reconstructed image of either the I-th or P-th frame; M=2, meaning the reference information for the (i+1)-th and (i+2)-th P-th frames can be the reconstructed features and image of the previous p-th frame; N=30, meaning the reference information for the (i+2)-(i+32)-th P-th frames can be the reconstructed features of the previous p-th frame.

[0242] Preferably, if the frame number is in the first interval, feature extraction can be performed on the image of the reference frame of the current frame to obtain the reference features of the current frame; if the frame number is in the second interval, feature fusion extraction can be performed on the image and features of the reference frame of the current frame to obtain the reference features of the current frame; if the frame number is in the third interval, feature extraction can be performed on the features of the reference frame of the current frame to obtain the reference features of the current frame. In a specific example, the first p-frame only references the reconstructed image of the I-frame / P-frame, the second p-frame can select the image and features of the reference frame, and the third p-frame and some subsequent frames only reference the features of the reference frame.

[0243] Optionally, the reference information for the current frame can be determined based on the frame number of the current frame. Furthermore, the ranges of the first, second, and third intervals mentioned above can be pre-agreed upon by the encoder and decoder, or they can be parsed from the bitstream by the decoder. In one example, the encoder and decoder agree that: the first p-frame only references the reconstructed image of the I-frame / P-frame; the second p-frame can choose the image and features of the reference frame; and the third p-frame and some subsequent frames only reference the features of the reference frame. In this case, the encoder does not need to include the range information of the first, second, and third intervals in the bitstream; the decoder can know from the agreement that: the first p-frame only references the reconstructed image of the I-frame / P-frame; the second p-frame can choose the image and features of the reference frame; and the third p-frame and some subsequent frames only reference the features of the reference frame. In another example, when the encoding end confirms that the distortion of an image frame encoded with the features of a reference frame exceeds a preset value, it can use the image of the reference frame as reference information to encode the image frame, and can encode the frame number of the image frame. The subsequent frames of the image frame are image encoded using the image and features of the reference frame as reference information, and the encoding end can encode the frame number of the first frame among these frames. The next few frames are image encoded using the features of the reference frame as reference information, and the encoding end can encode the frame number of the first frame among the frames image encoded using the features of the reference frame as reference information. In this way, the encoding end can inform the decoding end of the range of each interval.

[0244] In one embodiment, feature fusion extraction is performed on the image and features of a reference frame of the current frame to obtain reference features of the current frame. This may include: using a first feature extraction network to perform feature fusion extraction on the image and features of the reference frame of the current frame to obtain a first intermediate feature of the current frame; and using a unified feature extraction network to extract features from the first intermediate feature to obtain reference features of the current frame. Feature extraction on the image of the reference frame of the current frame to obtain reference features of the current frame includes: using a second feature extraction network to extract features from the image of the reference frame of the current frame to obtain a second intermediate feature of the current frame; and using a unified feature extraction network to extract features from the second intermediate feature to obtain reference features of the current frame. Feature extraction on the features of the reference frame of the current frame to obtain reference features of the current frame includes: using a third feature extraction network to extract features from the features of the reference frame of the current frame to obtain a third intermediate feature of the current frame; and using a unified feature extraction network to extract features from the third intermediate feature to obtain reference features of the current frame. That is, as follows... Figure 16 As shown, reference information for the current frame can be selected from three candidate reference information options: reference frame image, reference frame image + features, and reference frame features. After selecting the reference information for the current frame, the feature extraction module corresponding to the reference information can be used to adjust the information and / or transform the dimensions of the reference information to obtain the intermediate features of the current frame (such as the first intermediate feature, second intermediate feature, or third intermediate feature mentioned above). Then, the intermediate features of the current frame are input into a unified feature extraction module to generate the reference features of the current frame.

[0245] In another embodiment, the first feature extraction network can be directly used to perform feature fusion extraction on the image and features of the reference frame of the current frame to obtain the reference features of the current frame. Alternatively, the second feature extraction network can be directly used to extract features from the image of the current frame to obtain the reference features of the current frame. Furthermore, the third feature extraction network can be directly used to extract features from the image of the current frame to obtain the reference features of the current frame.

[0246] The aforementioned feature extraction module may include, but is not limited to, networks such as simple convolutional networks, residual networks, and / or attention modules. For example, the feature extraction module may include convolutional networks and residual networks. Alternatively, the feature extraction module may include convolutional networks, residual networks, and channel attention networks. It is understood that the feature extraction module of this application may also be other network structures, and this application does not limit the feature extraction module. In a specific example, the first feature extraction network includes convolutional units, residual units, and channel attention units; the second feature extraction network includes convolutional units and residual units; the third feature extraction network includes convolutional units and residual units; and the unified feature extraction network includes convolutional units and residual units.

[0247] S74: Perform motion compensation based on the reference features to obtain the prediction information for the current frame.

[0248] After obtaining the reference features of the current frame based on the above steps, motion compensation can be performed based on the reference features to obtain the prediction information of the current frame.

[0249] Optionally, motion compensation can be performed on the reference features based on the motion information of the current frame to extract relevant information from the reference frame and obtain the prediction information of the current frame.

[0250] Optionally, the motion information of the current frame can be determined by the image decoding method of the first embodiment. That is, the hyper-priority decoded motion features and residual information can be fused to obtain a first fused feature, and then the motion feature bitstream of the current frame can be decoded based on the first fused feature to obtain the motion information of the current frame.

[0251] In other implementations, residual features may not be referenced when determining the probability parameters of motion features. For example, in one embodiment, the super-prior decoded motion features can be directly processed by the fusion module to obtain the first fused feature; then, the decoded motion features of the current frame are obtained based on the first fused feature, thereby obtaining the motion information of the current frame.

[0252] Alternatively, the method of motion compensation for the reference features is not limited, and can be, for example, an interpolation-based warp alignment method, deformable convolution, or a combination of both.

[0253] S75: Decode the reconstructed image of the current frame based on the prediction information.

[0254] The steps of decoding the current frame based on the prediction information to obtain the reconstructed image of the current frame may include: calculating the reconstructed image of the current frame based on the prediction information and the residual information of the current frame.

[0255] Corresponding to the fourth embodiment of the image decoding method, such as Figure 17As shown, this application also provides an image encoding method according to a fourth embodiment, which can be applied to video encoding. The method includes at least three consecutive frame number intervals for the video to be encoded, which can be sequentially designated as a first interval, a second interval, and a third interval. The specific steps of this embodiment can be executed using the aforementioned encoding terminal. The method may include the following steps:

[0256] S81: In response to the current frame's frame number being located in either the first interval or the third interval, feature extraction is performed on the image of the reference frame of the current frame to obtain the reference features of the current frame.

[0257] S82: In response to the current frame's frame number being located in the second interval, feature fusion extraction is performed on the image and features of the reference frame of the current frame to obtain the reference features of the current frame.

[0258] S83: In response to the current frame's frame number being located in either the first interval or the third interval, feature extraction is performed on the features of the reference frame of the current frame to obtain the reference features of the current frame.

[0259] For a detailed description of steps S71, S72 and S73, please refer to them; they will not be repeated here.

[0260] S84: Perform motion compensation based on the reference features to obtain the prediction information for the current frame.

[0261] See the detailed description of step S74, which will not be repeated here.

[0262] Please see Figure 18 , Figure 18 This is a flowchart illustrating the fifth embodiment of the image decoding method of this application. The specific steps of this embodiment can be executed using the decoding end described above. The method may include the following steps:

[0263] S91: In response to the current frame not being a refresh frame, enable all types of cache information;

[0264] S92: In response to the current frame being a refresh frame, disable certain types of cached information;

[0265] S93: Decode the current frame using the enabled buffer information.

[0266] Please see Figure 19 , Figure 19 This is a flowchart illustrating the fifth embodiment of the image encoding method of this application. The specific steps of this embodiment can be executed using the decoding end described above. The method may include the following steps:

[0267] S101: In response to the current frame not being a refresh frame, enable all types of cache information;

[0268] S102: In response to the current frame being a refresh frame, disable certain types of cached information;

[0269] S103: Encode the current frame using the enabled buffer information.

[0270] For the first frame in the video stream, this frame is designated as an I-frame (keyframe). The characteristic of an I-frame is that it contains complete image information. Therefore, after the encoder sends the I-frame to the decoder, the decoder only needs the data from the I-frame to reconstruct the complete image during the decoding process. For the second frame in the video stream, when encoding and decoding the second frame, only the reconstructed image of the I-frame can be referenced. At this time, there is no other cached information to refer to (the cached information is empty). Compression and reconstruction are performed based on the P-frame type processing model. Subsequently, the cached information output by the P-frame model (reconstructed image, reconstruction features, motion context, and / or decoding features, etc.) can be obtained and used as reference information for subsequent frames in the video stream. For the third frame and subsequent frames in the video stream (i.e., frames following the second frame), compression and reconstruction are performed based on the P-frame model. At this time, cached information from the P-frame model can be referenced (i.e., cached information generated during the decoding process of frames preceding the frame to be decoded, such as reconstruction features, motion context, and / or decoding features from the P-model, etc.). In some scenarios, to improve data processing efficiency, reconstructed images are not the first choice when using cached information.

[0271] It should be noted that in video encoding and decoding methods, the cached information commonly used may include, but is not limited to, some cached features, such as reconstruction feature information, motion context information and / or decoding feature information of decoded frames. Although encoding and decoding based on this information can improve the efficiency of the encoding and decoding process, most of the time when using this information, such as prediction processing and estimation processing, error accumulation may occur under normal circumstances.

[0272] Therefore, in the image decoding method of this embodiment, some types of cached information can be turned off and only some types of cached information can be turned on when decoding refresh frames. In this way, by setting to turn off at least some types of cached information when decoding refresh frames during video encoding and decoding, the error transmission of previous frames can be reduced, and the accumulation of errors can be cut off. This reduces the accumulation of errors during video encoding and decoding, and improves the data compression effect and the subsequent data reconstruction effect.

[0273] As described above, the cached information may include at least one of the following: reconstructed video frame images generated during the decoding process, reconstructed feature information, motion context information, and other temporal priors (such as motion features and residual features output by the entropy module). That is, the cached information may include, but is not limited to, intermediate information generated during the decoding of previous frames, such as the reconstructed frame (reference frame image) of the reference frame, reference frame features (reconstructed features), motion context, and other temporal priors (motion features, residual features, etc. output by the entropy module).

[0274] If the current frame is not a refresh frame, all types of buffer information can be enabled. In this way, during decoding, the current frame can be decoded using the reconstructed frame (reference frame image), reference frame features (reconstructed features), motion context and / or other temporal priors (motion features, residual features, etc. output by the entropy module).

[0275] If the current frame is a refresh frame, some types of cached information can be disabled, meaning only some types of cached information can be enabled. These enabled types of cached information may include the reconstructed features of the reference frame. Thus, when decoding the refresh frame, only the enabled cached information can be used to decode the current frame. By setting the refresh frame and making it refer only to some types of cached information, i.e., appropriately disabling cached information passed from previous frames, the error propagation from previous frames is reduced, and the accumulation of errors can be truncated. In one embodiment, when decoding the refresh frame, only the reconstructed features of the reference frame can be enabled as cached information, so only the reconstructed features of the reference frame can be used to decode the current frame. In another embodiment, when decoding the refresh frame, only the reconstructed image of the reference frame can be enabled as cached information, so only the reconstructed image of the reference frame can be used to decode the current frame. In yet another embodiment, when decoding the refresh frame, only the reconstructed features and the reconstructed image of the reference frame can be enabled as cached information, so only these two types of cached information can be used to decode the current frame. It is understandable that disabling certain types of cached information means temporarily preventing the access of certain types of cached information by disabling the cached information.

[0276] Optionally, it can be determined whether the current frame is a refresh frame based on the encoding / decoding status parameters of the current frame. The encoding / decoding status parameters refer to relevant parameters of the current encoding / decoding status that can be obtained when encoding processing is required for the current video frame to be encoded (e.g., the frame number of the current video frame to be encoded, the frame number of the previous video frames to be encoded, loss information for each encoding / decoding process, loss before and after quantization, or bitstream size, etc.; no limitation is made here, and one or more of these data can be used as encoding / decoding status parameters).

[0277] The frame refresh rate can be set according to the encoding / decoding parameters.

[0278] In one embodiment, the refresh frame can be set according to the frame number.

[0279] For example, a preset refresh cycle can be used to determine whether the current frame is a refresh frame based on the frame number of the current frame.

[0280] In a specific example, the preset refresh interval is N, such as N can be set to 2, 4, 6, 8, 10 or 32, etc. When the number of frames of the current frame has not reached the refresh cycle (e.g. N or a multiple of N), the current frame is not a refresh frame. When the number of frames of the current frame reaches the refresh cycle, the current frame can be a refresh frame.

[0281] Optionally, the codec can agree on a refresh interval, so that the codec can automatically determine the refresh frame according to the agreed refresh interval.

[0282] In another embodiment, the encoding end determines the refresh period based on the video and can include refresh period-related parameters (such as refresh interval) into the bitstream so that the decoding end knows the refresh period of the image frames in the video.

[0283] Optionally, the encoding end can determine the refresh cycle based on the distortion of image frames in the video. For example, the encoding end can determine how many frames (which can be set to K frames) of continuous encoding with all types of buffer information in the video frame will cause the image frame distortion to exceed the threshold. After the encoding end confirms this, it can set the refresh interval to K-1 and include the refresh interval in the bitstream so that the decoding end knows the refresh cycle of the video.

[0284] For example, multiple sets of values ​​N can be preset. The encoding end uses the data in each set of N to truncate the error and calculate the loss. Then, the N value with the smallest loss can be selected and entered into the bitstream for the decoding end to use.

[0285] If the number of test frames is long, there can be multiple N values. For example, the refresh parameter N value can include 32 and 48. Then, when the frame number of the current frame is 32 or 48 or an integer multiple thereof, the current frame is a refreshed frame; otherwise, the current frame is a non-refreshed frame.

[0286] For example, the specific encoding and decoding process can refer to common video encoding and decoding technologies. At the encoding end, the current video frame to be encoded is encoded (e.g., entropy encoding) to obtain the encoding result (in the form of a data bitstream) corresponding to the current video frame to be encoded. The encoding result is then sent to the decoding end so that the decoding end can perform decoding and reconstruction processing on the encoding result to obtain the reconstructed video frame corresponding to the current video frame to be encoded.

[0287] Based on the above embodiments, this embodiment should be noted that neural networks (e.g., in the video encoding and decoding process of this application) can be used. Figure 2 or Figure 3 The video compression framework shown performs encoding and decoding processing, for example, using a neural network to encode the current video frame to be encoded. Specific neural network architectures may include, but are not limited to, those of common deep learning-based video encoding and decoding models, which will not be elaborated upon here.

[0288] When determining that the current frame is a refresh frame, some types of buffered information can be disabled, so that only the buffered information of the remaining types is used for encoding and decoding the current frame. For example, for Figure 3 The video compression framework can disable residual features, motion information, and motion context information in the buffer while maintaining the reference frame reconstruction features during frame encoding and decoding. This allows for feature extraction using the reconstructed features of the reference frame to obtain the reference features of the current frame during frame decoding. Alternatively, without referencing residual features (i.e., the temporal residual features in the above implementation) and motion features, the motion information bitstream can be processed using a motion vector entropy module and a motion vector decoder to obtain motion information. Then, the motion compensation and temporal prediction modules can be used to perform motion compensation on the reference features based on the motion information output by the motion vector encoder. The residual information of the current frame is obtained by processing the residual information bitstream using the residual information entropy module and the residual information decoder without referring to the residual features (i.e., the temporal residual features in the above embodiments). The residual information of the current frame is obtained by processing the residual information and motion information of the current frame. When the residual features of the buffer are turned off, the output of the residual information encoder and / or the residual information entropy module can be optimized without using the optimization module. That is, the output of the residual information encoder can be directly input to the residual information entropy module, and the output of the residual information entropy module can be directly input to the residual information decoder.

[0289] When the current frame is not a refresh frame, i.e., when the current frame is a non-refresh frame, the available buffer information can be used to encode and decode the current frame. For example, for Figure 3The video compression framework, when decoding non-refreshed frames, can utilize the reconstructed features of the reference frame and / or the reconstructed image to extract features and obtain the reference features of the current frame; it can also, in the case of reference residual features (i.e., the temporal residual features in the above embodiments) and motion features, use the motion vector entropy module and motion vector decoder to process the motion information bitstream to obtain motion information, and then use the motion compensation and temporal prediction module to perform motion compensation on the reference features based on the motion information output by the motion vector encoder to obtain the motion information of the current frame; and it can also utilize the reference residual features (i.e., the temporal residual features in the above embodiments) to extract features and obtain the reference features and / or ... In the case of residual features, the residual information bitstream is processed by the residual information entropy module and the residual information decoder to obtain the residual information of the current frame. Based on the residual information and motion information of the current frame, the reconstructed image of the current frame is obtained. In addition, the output of the residual information encoder and / or the residual information entropy module can be optimized by the optimization module (i.e. the first optimization module and / or the second optimization module mentioned above). That is, the output of the residual information encoder optimized by the optimization module can be input into the residual information entropy module, and the output of the residual information entropy module optimized by the optimization module can be input into the residual information decoder.

[0290] The above reference relationships were discovered during the testing process of this application. Therefore, various reference relationships, including but not limited to those mentioned above, can be introduced during model training to ensure that the training and testing processes are aligned.

[0291] Furthermore, for refresh frames, the image of a reference frame can be used as reference information, and feature extraction can be performed on the image of the reference frame to obtain reference features for the refresh frame. These reference features are then used to encode and decode the current frame. For the first preset number of frames following the refresh frame, the image and features of the reference frame can be used as reference information, and feature extraction can be performed on the image and features of the reference frame to obtain reference features for the first preset number of frames. Then, for the second preset number of frames following the first preset number of frames, the features of the reference frame can be used as reference information, and feature extraction can be performed on the features of the reference frame to obtain reference features for the second preset number of frames; alternatively, for the remaining frames in the refresh cycle, i.e., frames other than the refresh frame and the first preset number of frames following it, the features of the reference frame can be used as reference information, and feature extraction can be performed on the features of the reference frame to obtain reference features for the remaining frames.

[0292] Regarding the above embodiments, this application provides a computer device; please refer to [link / reference]. Figure 20 , Figure 20This is a schematic diagram of a computer device according to an embodiment of the present application. The computer device 200 includes a memory 201 and a processor 202, wherein the memory 201 and the processor 202 are coupled to each other. The memory 201 stores program data, and the processor 202 executes the program data to implement the steps of any embodiment of the above-described filtering processing method. The computer device 200 can serve as the encoding end and / or decoding end in the image encoding / decoding system of the above embodiments, executing the steps of any embodiment of the above-described image decoding method and image encoding method.

[0293] In this embodiment, processor 202 can also be referred to as CPU (Central Processing Unit). Processor 202 may be an integrated circuit chip with signal processing capabilities. Processor 202 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. The general-purpose processor can be a microprocessor, or processor 202 can be any conventional processor.

[0294] The methods described in the above embodiments can be implemented as computer programs; therefore, this application proposes a computer-readable storage medium. Please refer to [link to relevant documentation]. Figure 21 , Figure 21 This is a schematic diagram of a computer-readable storage medium according to an embodiment of the present application. The computer-readable storage medium 300 stores program data 301 that can be executed by a processor. The program data 301 can be executed by the processor to implement the steps of any of the above-described image decoding method and image encoding method embodiments.

[0295] In this embodiment, the computer-readable storage medium 300 can be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or a medium that can store program data 301. Alternatively, it can be a server that stores the program data 301, which can send the stored program data 301 to other devices for execution, or it can self-run the stored program data 301.

[0296] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments. For the sake of brevity, this application will not repeat the details here.

[0297] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to. For the sake of brevity, the present application will not repeat them here.

[0298] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. An image decoding method characterized by, The method comprises: in response to the current frame not being a refresh frame, opening all kinds of cache information; in response to the current frame being a refresh frame, only opening part of the cache information, the part of the cache information at least including reference frame features; decoding the current frame by using the opened cache information.

2. The image decoding method according to claim 1, characterized by, The response to the current frame being a refresh frame only opening part of the cache information comprises: in response to the current frame being a refresh frame, closing the residual information and motion information in the cache information, and only opening the reconstructed features of the reference frame P frame model.

3. The image decoding method according to claim 1, characterized by, The method comprises: in response to the frame number of the current frame not reaching a refresh period, the current frame not being a refresh frame; in response to the frame number of the current frame reaching a refresh period, the current frame being a refresh frame; wherein the refresh period is parsed from a bitstream or preset.

4. The image decoding method according to claim 1, characterized by, The decoding of the current frame by using the opened cache information comprises: if the opened cache information includes residual information, fusing the hyper-prior decoded motion features and residual features of the current frame to obtain first fused features; decoding the motion feature code stream of the current frame based on the first fused features to obtain the motion information of the current frame; decoding the current frame based on the motion information of the current frame.

5. The image decoding method of claim 1, wherein, The decoding of the current frame by using the opened cache information comprises: if the opened cache information includes residual information, fusing the hyper-prior decoded residual features and first alignment features of the current frame to obtain second fused features, the first alignment features being obtained by motion compensation on first temporal residual features, the first temporal residual features being residual information decoded from a decoded frame before the current frame; decoding the residual feature code stream of the current frame based on the second fused features to obtain the residual information of the current frame; decoding the current frame based on the residual information of the current frame.

6. The image decoding method of claim 1, wherein, The decoding of the current frame by using the opened cache information comprises: if the opened cache information includes residual information, optimizing the decoded residual features of the current frame by using second temporal residual features and / or prediction features to obtain first optimized features, the second temporal residual features being residual information decoded from a decoded frame before the current frame; reconstructing the residual information from the first optimized features to obtain the residual information of the current frame; decoding the current frame based on the residual information of the current frame.

7. An image coding method characterized by, The method comprises: in response to the current frame not being a refresh frame, opening all kinds of cache information; in response to the current frame being a refresh frame, closing part of the cache information encoding the current frame by using the opened cache information.

8. The image coding method according to claim 7, characterized by, A plurality of candidate refresh period parameters are set, and the method comprises: for the same to-be-encoded video stream, respectively based on each refresh period parameter, correspondingly encoding and decoding the to-be-encoded video stream to obtain loss information in the encoding and decoding process corresponding to each refresh period parameter; selecting the refresh period parameter corresponding to the encoding and decoding process with the minimum loss information to send to a decoding end, so that the decoding end and the encoding end judge a refresh frame based on the same refresh period parameter.

9. An image decoding method, characterized by, The method comprises: determining a hyper-prior decoded motion feature of a current frame; fusing the hyper-prior decoded motion feature and a residual feature to obtain a first fused feature; decoding a motion feature code stream of the current frame based on the first fused feature to obtain motion information of the current frame; decoding the current frame based on the motion information of the current frame.

10. The image decoding method of claim 9, wherein, The fusing the hyper-prior decoded motion feature and a residual feature to obtain a first fused feature comprises: fusing the hyper-prior decoded motion feature, a reference feature and the residual feature to obtain the first fused feature.

11. The image decoding method according to claim 9, characterized in that, The fusing the hyper-prior decoded motion feature and a residual feature to obtain a first fused feature comprises: scaling the residual feature by using a first scaling factor to obtain a scaled residual feature; fusing the hyper-prior decoded motion feature and the scaled residual feature to obtain the first fused feature.

12. An image coding method characterized by comprising: The method comprises: determining a hyper-prior decoded motion feature of a current frame; fusing the hyper-prior decoded motion feature and a residual feature to obtain a first fused feature; decoding a motion feature code stream of the current frame based on the first fused feature to obtain motion information of the current frame.

13. The image coding method of claim 12, wherein, The determining a hyper-prior decoded motion feature of a current frame comprises: hyper-prior encoding processing a to-be-encoded motion feature of the current frame based on a current frame feature and / or a residual feature to obtain a hyper-prior motion code stream of the current frame; analyzing the hyper-prior motion code stream of the current frame to obtain the hyper-prior decoded motion feature of the current frame.

14. The image coding method of claim 13, wherein, The hyper-prior encoding processing a to-be-encoded motion feature of the current frame based on a current frame feature and / or a residual feature comprises: scaling the residual feature by using a first scaling factor to obtain a scaled residual feature; hyper-prior encoding processing the to-be-encoded motion feature of the current frame based on the current frame feature and / or the scaled residual feature.

15. An image decoding method, characterized by, The method comprises: determining a hyper-prior decoded residual feature of a current frame; fusing the hyper-prior decoded residual feature of the current frame and a first aligned feature to obtain a second fused feature; the first aligned feature is obtained by motion compensation on a first temporal residual feature, and the first temporal residual feature is residual information generated by decoding a decoded frame before the current frame; decoding a residual feature code stream of the current frame based on the second fused feature to obtain residual information of the current frame; decoding the current frame based on the residual information of the current frame.

16. The image decoding method of claim 15, wherein, The method comprises: motion compensating the first temporal residual feature based on first motion information and / or the hyper-prior decoded residual feature.

17. The image decoding method of claim 16, wherein, The motion compensating the first temporal residual feature based on first motion information and / or the hyper-prior decoded residual feature comprises: transforming the first motion information to obtain second motion information; or determine third motion information by using the first temporal residual feature and the hyper-prior decoded residual feature; and perform motion compensation on the first temporal residual feature by using the third motion information to obtain the first aligned feature; or, input one of the first motion information and the second motion information, the first temporal residual feature and the hyper-prior decoded residual feature into a fine-tuning module to obtain third motion information output by the fine-tuning module; and perform motion compensation on the first temporal residual feature by using the third motion information to obtain the first aligned feature.

18. The image decoding method of claim 15, wherein, The fusing the hyper-prior decoded residual feature and the first aligned feature of the current frame to obtain the second fused feature comprises: fusing the hyper-prior decoded residual feature, the predicted feature and the first aligned feature to obtain the second fused feature.

19. The image decoding method of claim 16, wherein, The motion compensation on the first temporal residual feature based on the first motion information and / or the hyper-prior decoded residual feature comprises: scaling the first temporal residual feature by using a second scaling factor to obtain a scaled first temporal residual feature; performing motion compensation on the scaled first temporal residual feature based on the first motion information and / or the hyper-prior decoded residual feature.

20. An image coding method characterized by, The method comprises: determining a hyper-prior decoded residual feature of a current frame; fusing the hyper-prior decoded residual feature and a first aligned feature of the current frame to obtain a second fused feature; the first aligned feature is obtained by performing motion compensation on a first temporal residual feature; the first temporal residual feature is residual information decoded from a decoded frame before the current frame; performing decoding processing on a residual feature code stream of the current frame based on the second fused feature to obtain residual information of the current frame.

21. The image coding method according to claim 20, wherein The determination of the hyper-prior decoded residual feature of the current frame comprises: performing hyper-prior encoding processing on a to-be-encoded residual feature of the current frame based on a second aligned feature and / or a current frame feature to obtain a hyper-prior residual code stream of the current frame; the second aligned feature is obtained by performing motion compensation on a first temporal residual feature; performing analysis on the hyper-prior residual code stream of the current frame to obtain the hyper-prior decoded residual feature of the current frame.

22. The image coding method of claim 21, wherein, The motion compensation on the first temporal residual feature based on the first motion information and / or the to-be-encoded residual feature comprises: scaling the first temporal residual feature by using a second scaling factor to obtain a scaled first temporal residual feature; 23. The image coding method of claim 22, wherein, performing motion compensation on the scaled first temporal residual feature based on the first motion information and / or the to-be-encoded residual feature. The motion compensation on the first temporal residual feature based on the first motion information and / or the to-be-encoded residual feature comprises: performing transformation processing on the first motion information to obtain second motion information; and performing motion compensation on the first temporal residual feature by using the second motion information to obtain the second aligned feature; or, 24. The image coding method of claim 22, wherein, ​ ​ determining fourth motion information by using the first temporal residual feature and the to-be-encoded residual feature; and performing motion compensation on the first temporal residual feature by using the fourth motion information to obtain the second aligned feature; or inputting one of the first motion information and the second motion information, the first temporal residual feature and the to-be-encoded residual feature into a fine-tuning module to obtain fourth motion information output by the fine-tuning module; and performing motion compensation on the first temporal residual feature by using the fourth motion information to obtain the second aligned feature.

25. The image coding method of claim 21, wherein, The hyper-prior encoding processing of the to-be-encoded residual feature of the current frame based on the second aligned feature and / or the current frame feature comprises: inputting the second aligned feature, the current frame feature and the to-be-encoded residual feature into a hyper-prior encoding module to obtain the hyper-prior residual code stream.

26. An image decoding method, comprising: The method comprises: optimizing a decoded residual feature of a current frame by using a second temporal residual feature and / or a prediction feature, to obtain a first optimized feature, wherein the second temporal residual feature is residual information decoded from a decoded frame before the current frame; reconstructing residual information from the first optimized feature to obtain residual information of the current frame; decoding the current frame based on the residual information of the current frame.

27. The image decoding method of claim 26, wherein the optimizing the decoded residual feature of the current frame by using the second temporal residual feature and / or the prediction feature comprises: a first optimization module optimizing the decoded residual feature output by a residual information entropy module by using the second temporal residual feature and / or the prediction feature, to obtain the first optimized feature; and the reconstructing residual information from the first optimized feature to obtain the residual information of the current frame comprises: a residual information decoder reconstructing residual information from the first optimized feature output by the first optimization module, to obtain the residual information of the current frame.

28. The image decoding method of claim 26, wherein, The method comprises: fusing one of the second temporal residual feature and the third aligned feature, and / or one of the prediction feature and the transform prediction feature into the decoded residual feature, to obtain the first optimized feature; and The transform prediction feature is obtained by transforming the prediction feature, and the third aligned feature is obtained by performing motion compensation on the second temporal residual feature.

29. The image decoding method of claim 28, wherein, The method further comprises: performing motion compensation on the second temporal residual feature based on the first motion information and / or the decoded residual feature.

30. The image decoding method of claim 29, wherein, The performing motion compensation on the second temporal residual feature based on the first motion information and / or the decoded residual feature comprises: transforming the first motion information to obtain second motion information; performing motion compensation on the second temporal residual feature by using the second motion information to obtain the third aligned feature; or performing motion compensation on the second temporal residual feature by using the first motion information to obtain the third aligned feature. determining fifth motion information by using the second temporal residual feature and the decoded residual feature; and performing motion compensation on the second temporal residual feature by using the fifth motion information to obtain the third aligned feature; or inputting one of the first motion information and the second motion information, the second temporal residual feature and the decoded residual feature into a fine-tuning module to obtain fifth motion information output by the fine-tuning module; and performing motion compensation on the second temporal residual feature by using the fifth motion information to obtain the third aligned feature.

31. An image coding method characterized by, The method comprises: optimizing a decoded residual feature of a current frame by using a second temporal residual feature and / or a prediction feature, to obtain a first optimized feature, the second temporal residual feature being residual information decoded from a decoded frame before the current frame; reconstructing residual information from the first optimized feature to obtain residual information of the current frame.

32. The image coding method of claim 31, wherein, The method further comprises: optimizing a to-be-encoded residual feature by using the second temporal residual feature and / or the prediction feature, to obtain a second optimized feature; quantizing and / or entropy encoding the second optimized feature to obtain a residual feature code stream of the current frame; parsing the residual feature code stream to obtain the decoded residual feature.

33. The image coding method of claim 32, wherein the step of optimizing the to-be-encoded residual feature by using the second temporal residual feature and / or the prediction feature comprises: a second optimization module optimizing the to-be-encoded residual feature output by a residual information encoder by using the second temporal residual feature and / or the prediction feature, to obtain the second optimized feature; and the step of quantizing and / or entropy encoding the second optimized feature to obtain the residual feature code stream of the current frame comprises: a residual information entropy module quantizing and / or entropy encoding the second optimized feature output by the second optimization module to obtain the residual feature code stream of the current frame.

34. The image coding method of claim 32, wherein, the step of optimizing the to-be-encoded residual feature by using the second temporal residual feature and / or the prediction feature to obtain the second optimized feature comprises: fusing one of the second temporal residual feature and a fourth aligned feature, and / or one of the prediction feature and a transformed prediction feature, into the to-be-encoded residual feature to obtain the second optimized feature; the transformed prediction feature being obtained by transforming the prediction feature; and the fourth aligned feature being obtained by performing motion compensation on the second temporal residual feature. The method further comprises:

35. The image coding method of claim 34, wherein, performing motion compensation on the second temporal residual feature based on the first motion information and / or the to-be-encoded residual feature to obtain the fourth aligned feature. the step of performing motion compensation on the second temporal residual feature based on the first motion information and / or the to-be-encoded residual feature to obtain the fourth aligned feature comprises:

36. The image coding method of claim 35, wherein, transforming the first motion information to obtain second motion information; and performing motion compensation on the second temporal residual feature by using the second motion information to obtain the fourth aligned feature; or performing motion compensation on the second temporal residual feature by using the first motion information to obtain the fourth aligned feature. determining sixth motion information by using the second temporal residual feature and the to-be-encoded residual feature; performing motion compensation on the second temporal residual feature by using the sixth motion information to obtain the fourth aligned feature; or, inputting one of the first motion information and the second motion information, the second temporal residual feature and the to-be-encoded residual feature into a fine-tuning module to obtain sixth motion information output by the fine-tuning module; performing motion compensation on the second temporal residual feature by using the sixth motion information to obtain the fourth aligned feature.

37. An image decoding method, comprising: The image decoding method is applied to video decoding, and at least three continuous frame number intervals are set for a to-be-decoded video, the three continuous frame number intervals are a first interval, a second interval and a third interval in sequence, and the method comprises: in response to the frame number of a current frame being located in one of the first interval and the third interval, performing feature extraction on an image of a reference frame of the current frame to obtain reference features of the current frame; in response to the frame number of the current frame being located in the second interval, performing feature fusion extraction on the image and the features of the reference frame to obtain the reference features of the current frame; in response to the frame number of the current frame being located in the remaining one of the first interval and the third interval, performing feature extraction on the features of the reference frame to obtain the reference features of the current frame; performing motion compensation based on the reference features to obtain prediction information of the current frame; decoding the reconstructed image of the current frame based on the prediction information.

38. The image decoding method of claim 37, wherein: the feature fusion extraction on the image and the features of the reference frame comprises: performing feature fusion extraction on the image and the features of the reference frame of the current frame by using a first feature extraction network to obtain first intermediate features of the current frame; and performing feature extraction on the first intermediate features by using a unified feature extraction network to obtain the reference features of the current frame; the feature extraction on the image of the reference frame of the current frame comprises: performing feature extraction on the image of the reference frame of the current frame by using a second feature extraction network to obtain second intermediate features of the current frame; and performing feature extraction on the second intermediate features by using the unified feature extraction network to obtain the reference features of the current frame; the feature extraction on the features of the reference frame comprises: performing feature extraction on the features of the reference frame of the current frame by using a third feature extraction network to obtain third intermediate features of the current frame; and performing feature extraction on the third intermediate features by using the unified feature extraction network to obtain the reference features of the current frame.

39. An image coding method characterized by, The image encoding method is applied to video encoding, and at least three continuous frame number intervals are set for a to-be-encoded video, the three continuous frame number intervals are a first interval, a second interval and a third interval in sequence, and the method comprises: in response to the frame number of a current frame being located in one of the first interval and the third interval, performing feature extraction on an image of a reference frame of the current frame to obtain reference features of the current frame; in response to the frame number of the current frame being located in the second interval, performing feature fusion extraction on the image and the features of the reference frame to obtain the reference features of the current frame; in response to the frame number of the current frame being located in the remaining one of the first interval and the third interval, performing feature extraction on the features of the reference frame to obtain the reference features of the current frame; performing motion compensation based on the reference features to obtain prediction information of the current frame; decoding the reconstructed image of the current frame based on the prediction information. in response to the frame number of the current frame being located in the second interval, performing feature fusion extraction on the image and the feature of the reference frame to obtain a reference feature of the current frame; in response to the frame number of the current frame being located in the remaining one of the first interval and the third interval, performing feature extraction on the feature of the reference frame to obtain a reference feature of the current frame; performing motion compensation based on the reference feature to obtain prediction information of the current frame.

40. A computer device, comprising: A computer readable storage medium storing program data, which when accessed by a machine is capable of causing the machine to perform the steps of any one of claims 1 to 39.

41. A computer-readable storage medium, comprising: A computer readable storage medium storing program data, which when accessed by a machine is capable of causing the machine to perform the steps of any one of claims 1 to 39.