Image decoding method, image coding method and related equipment
By generating and optimizing dependent reference features during image encoding and decoding, the problem of insufficient decoding performance in existing technologies is solved, achieving more efficient inter-frame information transmission and improved decoding performance.
Patent Information
- Application Number
- CN202411104598.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-12
- Publication Date
- 2026-02-24
AI Technical Summary
Existing image encoding and decoding methods have low decoding performance and are difficult to effectively utilize long-term dependent feature information for inter-frame information transmission, resulting in insufficient encoding and decoding efficiency and accuracy.
By acquiring the current frame bitstream for decoding, determining the dependency reference features, and passing them to subsequent frame bitstreams for decoding, the long-term dependency generation module generates long-term reference features, and the feature fusion and refresh modules optimize the feature information to improve the accuracy of inter-frame information transmission.
It strengthens the transmission of long-reliant feature information, improves image decoding performance, and enhances the accuracy and efficiency of encoding and decoding.
Smart Images

Figure CN121567883A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image encoding and decoding, and in particular to an image decoding method, an image encoding method, and related equipment. Background Technology
[0002] Because video images are large in size, they typically need to be encoded and compressed. The compressed video image data is called a video stream. The video stream can be transmitted to a decoding end via wired or wireless network for decoding and viewing. The entire image encoding and compression process can include prediction, transformation, quantization, and encoding. This reduces the amount of video data, thereby reducing network bandwidth usage during transmission and minimizing storage space.
[0003] However, current image encoding and decoding methods suffer from problems such as low decoding performance. Summary of the Invention
[0004] The main technical problem addressed by this application is to provide an image decoding method, an image encoding method, a computer device, and a computer-readable storage medium that can improve the decoding performance of images.
[0005] To address the aforementioned problems, the first aspect of this application provides an image decoding method, which includes: acquiring a current frame bitstream; decoding the current frame bitstream to obtain at least one decoding information; and determining a dependency reference feature based on the at least one decoding information; wherein the dependency reference feature is used to pass to a subsequent frame bitstream to decode the subsequent frame bitstream using the dependency reference feature.
[0006] To address the aforementioned issues, a second aspect of this application provides an image decoding method, comprising: acquiring a current frame bitstream; decoding the current frame bitstream to obtain initial reconstructed features; and processing the initial reconstructed features using dependent reference features to obtain reconstructed information of the current frame; wherein the dependent reference features are obtained by decoding the preceding frame bitstream using the aforementioned image decoding method.
[0007] To address the aforementioned issues, a third aspect of this application provides an image decoding method, comprising: acquiring a current frame bitstream; decoding the current frame bitstream to obtain motion information; and using the motion information to perform motion compensation on dependent reference features to obtain prediction information for the current frame; wherein the dependent reference features are obtained by decoding the preceding frame bitstream using the aforementioned image decoding method.
[0008] To address the aforementioned issues, a fourth aspect of this application provides an image encoding method, comprising: acquiring a current frame image to be encoded; encoding the current frame image to obtain a current frame bitstream; and providing the current frame bitstream to a decoding end so that the decoding end can decode the current frame bitstream using any of the image decoding methods described above.
[0009] To address the aforementioned problems, a fifth aspect of this application provides a computer device comprising a memory and a processor coupled to each other, wherein the memory stores program data and the processor executes the program data to implement any step of any of the methods described above.
[0010] To address the aforementioned problems, a sixth aspect of this application provides a computer-readable storage medium storing program data executable by a processor, the program data being used to implement any step of any of the methods described above.
[0011] The above scheme obtains the current frame bitstream, decodes it to get at least one type of decoding information, and determines a dependency reference feature based on the at least one type of decoding information. The dependency reference feature is used to pass to the subsequent frame bitstream so that the subsequent frame bitstream can be decoded using the dependency reference feature. This can strengthen the transmission of long-term dependency feature information and make full use of the transmitted dependency reference feature for decoding. This can strengthen the long-term information transmission between frames, improve the accuracy of prediction, and enhance the decoding performance of the image.
[0012] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in this application, the accompanying drawings required in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Among them:
[0014] Figure 1 This is a schematic diagram of the structure of an embodiment of the video encoding and decoding system of this application;
[0015] Figure 2 This is a flowchart illustrating the first embodiment of the image decoding method of this application;
[0016] Figure 3 This is a schematic diagram of the framework of an embodiment of the image encoding and decoding system of this application;
[0017] Figure 4 This is an example schematic diagram of an embodiment of the long-term dependency generation module of this application;
[0018] Figure 5 This is an example schematic diagram of an embodiment of the feature fusion module of this application;
[0019] Figure 6 This is an example schematic diagram of an embodiment of the feature refresh module of this application;
[0020] Figure 7 This is a flowchart illustrating the second embodiment of the image decoding method of this application;
[0021] Figure 8 This application Figure 7 A flowchart illustrating an embodiment of step S23;
[0022] Figure 9 This is an example schematic diagram of an embodiment of motion compensation in this application;
[0023] Figure 10 This is a flowchart illustrating the third embodiment of the image decoding method of this application;
[0024] Figure 11 This is an example schematic diagram of an embodiment of the loop network method of this application;
[0025] Figure 12 This is an example schematic diagram of an embodiment of the time-domain correlation method of this application;
[0026] Figure 13 This is an example schematic diagram of an embodiment of non-local attention processing in this application;
[0027] Figure 14 This is a flowchart illustrating an embodiment of the image encoding method of this application;
[0028] Figure 15 This is a schematic diagram of the structure of the first embodiment of the decoding end of this application;
[0029] Figure 16 This is a schematic diagram of the structure of the second embodiment of the decoding end of this application;
[0030] Figure 17 This is a schematic diagram of the structure of the third embodiment of the decoding end of this application;
[0031] Figure 18 This is a schematic diagram of the structure of an embodiment of the encoding end of this application;
[0032] Figure 19 This is a schematic diagram of the structure of an embodiment of the computer device of this application;
[0033] Figure 20 This is a schematic diagram of the structure of an embodiment of the computer-readable storage medium of this application. Detailed Implementation
[0034] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0035] The terms "first" and "second" in this application are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.
[0036] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0037] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this document means two or more. Moreover, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of objects. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0038] This application provides the following embodiments, and each embodiment is described in detail below.
[0039] Please see Figure 1 , Figure 1 This is a schematic diagram of the structure of an embodiment of the image encoding and decoding system of this application.
[0040] The image encoding / decoding system 100 includes an encoding end 101 and a decoding end 102. The encoding end 101 and the decoding end 102 can be computer equipment, electronic equipment, etc., and can be any device with processing capabilities, such as a computer, server, mobile phone, tablet, etc. This application does not impose any limitations on this. The encoding end 101 and the decoding end 102 can communicate with each other and can be used to perform encoding and / or decoding operations on images / videos.
[0041] The encoding end 101 can be used to perform preprocessing and encoding / compression steps for images / videos to obtain bitstream data. The encoding end 101 can transmit the bitstream data to the decoding end 102. The decoding end 102 can receive the bitstream data from the encoding end 101 and perform decoding and other related steps involving the bitstream data, as well as steps related to backend vision tasks, such as image / video processing and classification.
[0042] This application provides an image decoding method, an image encoding method, and related equipment to improve image encoding and decoding performance and accuracy. In some embodiments, the encoding end 101 and the decoding end 102 of this embodiment can be used to implement any step of the following embodiments.
[0043] Please see Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the image decoding method of this application. The specific steps of this embodiment can be executed using the decoding end described above. The method may include the following steps:
[0044] S11: Get the current frame bitstream.
[0045] The encoding end can encode the current frame image to obtain the current frame bitstream, and the decoding end can obtain the current frame bitstream for decoding. The current frame bitstream includes at least a motion bitstream and a context bitstream. The motion bitstream is obtained by the encoding end encoding motion information, and the context bitstream is obtained by the encoding end encoding context information. Specifically, the encoding end performs motion estimation on the current frame image and a reference frame image to obtain motion information. The motion information can be motion vectors or motion context, etc., and this application does not limit the type of motion information. Then, the motion information is encoded to obtain the motion bitstream. Motion compensation is performed using the reconstructed motion information to predict the predicted information. The context information of the current frame image and the motion information is obtained, and the context information is encoded to obtain the context bitstream. The decoding end can obtain both the motion bitstream and the context bitstream.
[0046] S12: Decode the current frame bitstream to obtain at least one type of decoding information.
[0047] Decoding the current frame bitstream can also decode the motion bitstream and the context bitstream to obtain at least one type of decoding information.
[0048] In some embodiments, at least one type of decoding information includes at least one of the following: prediction features, reconstruction features, decoded intermediate features, the current reconstructed frame, the predicted image frame, motion information, etc. It is understood that the decoding information may include any decoding information obtainable from the current frame that has been decoded during the decoding process. This application does not limit the decoding information.
[0049] For example, decoding a motion bitstream yields motion information, and using this motion information for motion compensation provides prediction information, which includes prediction features and / or predicted image frames. Similarly, decoding a context bitstream yields context information, which is then used for frame reconstruction to obtain reconstructed information, including reconstructed features and / or the current reconstructed frame. The reconstructed features can be used as reference features, and the current reconstructed frame as a reference frame image for subsequent frame encoding and decoding processes. This application does not limit the method of obtaining decoding information.
[0050] In some implementations, during the decoding of the current frame bitstream, the current frame bitstream can be decoded with reference to the dependency reference features obtained from the decoding of the preceding frame bitstream, in order to obtain at least one type of decoding information. The preceding frame can be represented as an image frame that precedes the current frame in the encoding / decoding order.
[0051] S13: Determine a dependency reference feature based on at least one decoding information; wherein the dependency reference feature is used to pass to the subsequent frame bitstream so as to decode the subsequent frame bitstream using the dependency reference feature.
[0052] During encoding and decoding, at least one type of decoded information can be used as reference information so that subsequent frames can be encoded and decoded using this reference information as reference frames. In this embodiment, dependent reference features can be generated for at least one type of decoded information, and these dependent reference features can be used as reference information and passed to subsequent frames for encoding and decoding. During decoding, the dependent reference features are passed to the subsequent frame bitstream so that the subsequent frame bitstream can be decoded with reference to these dependent reference features. For example, predicted features and / or reconstructed features can be used as reference features to generate dependent reference features.
[0053] In end-to-end video encoding and decoding, it is necessary to refer to previously decoded frames to remove temporal redundancy and improve compression ratio. When a reference frame close to the current frame is occluded or suffers information loss, information from a more distant reference frame can be referenced. However, as the number of frames being encoded and decoded increases, the reference information that can be passed from distant reference frames to the current frame becomes increasingly scarce. To address this, this embodiment employs a method of generating dependent reference features for at least one type of decoded information. This allows the acquisition of features from distant frames (reference frames) as reference information to be passed to the next frame for encoding and decoding, achieving long-term information dependency and improving encoding and decoding performance.
[0054] The above scheme obtains the current frame bitstream, decodes it to get at least one type of decoding information, and determines a dependency reference feature based on the at least one type of decoding information. The dependency reference feature is used to pass to the subsequent frame bitstream so that the subsequent frame bitstream can be decoded using the dependency reference feature. This can strengthen the transmission of long-term dependency feature information and make full use of the transmitted dependency reference feature for decoding. This can strengthen the long-term information transmission between frames, improve the accuracy of prediction, and enhance the decoding performance of the image.
[0055] Please see Figure 3 , Figure 3 This is a schematic diagram of the framework of an embodiment of the image encoding and decoding system of this application. The image encoding and decoding system includes a long-term dependency generation module, a reference information module, a feature fusion module (optional), a feature refresh module, a feature adjustment module (optional), a motion estimation module, a motion information encoding module (i.e., a motion information encoder), a motion information entropy model, a motion information decoding module (i.e., a motion information decoder), a motion compensation module (such as motion compensation and temporal prediction), a context encoding module (i.e., a context encoder), a context entropy model (such as a residual information entropy model), a context decoding module (i.e., a context decoder), and a frame reconstruction module, etc. This application is not limited to this.
[0056] The system comprises the following modules: a long-term dependency generation module for generating dependency reference features; a reference information module for caching reference information (such as dependency reference features); a feature fusion module for fusing features from dependency reference features; a feature refresh module for updating dependency reference features; and a feature adjustment module for performing feature adjustments, such as fine-tuning.
[0057] The motion estimation module estimates motion between the current frame image (or current frame features) and a reference frame (or reference information or dependent reference features) to obtain motion information (such as motion context or motion vectors). The motion information encoding module (i.e., the motion information encoder) reduces the dimensionality of the motion information (or motion information residuals) and extracts compact information to obtain the motion information to be encoded. The motion information entropy model obtains the probability of each character appearing in the quantized motion information to be encoded and performs arithmetic encoding to output the motion bitstream. During decoding, low-dimensional motion features can be obtained by decoding the motion bitstream. The motion information decoding module (i.e., the motion information decoder) increases the dimensionality of the decoded low-dimensional motion features and reconstructs the motion information.
[0058] The motion compensation module is used to perform motion compensation based on motion information to obtain prediction information. For example, it performs motion compensation operations such as warping on reference information using motion information, extracts relevant information from the reference information, and obtains the prediction information for the current frame.
[0059] The context encoding module (i.e., the context encoder) is used to reduce the dimensionality and compress the context information between the current frame image (or current frame features) and the reference frame (or reference information or dependent reference features) to obtain the context information to be encoded. Context information includes, but is not limited to, the difference information or concatenated information between the two frames. Its purpose is to remove temporal correlations between frames and spatial correlations within frames. The context entropy model is used to encode the context information, obtaining the probability of each character appearing in the quantized context information to be encoded, and performing arithmetic encoding to obtain the context bitstream. The context decoding module (i.e., the context decoder) is used to increase the dimensionality of the low-dimensional context features obtained from the decoded bitstream and reconstruct the context information. In this process, the context information is combined with prediction information to obtain the initial reconstructed information. For example, if the context information at the encoding end is represented as a difference, the decoding end can add the context information to the prediction information to obtain the initial reconstructed information of the current frame.
[0060] The frame reconstruction module is used to further process the preliminary reconstruction information to obtain the reconstruction information of the current frame, such as the current reconstructed frame and reconstruction features. Both are stored in the reference information module and can be used as reference information for the next frame and subsequent frames.
[0061] The following combination Figure 3 Each module of this application is described.
[0062] In some embodiments, a long-term dependency generation module can be used to generate dependency reference features. The long-term dependency generation module includes a long-term dependency generation network, which processes the decoded information and the dependency reference features of the preceding frame to obtain the dependency reference features corresponding to the decoded information. When the current frame is the first frame, the dependency reference features of the preceding frame can be the initialized reference features. The long-term dependency generation network can process temporal information and extract long-distance feature dependencies. For example, the long-term dependency generation network may include recurrent neural networks (RNNs), long short-term memory (LSTMs), or convLSTMs, or other convolutional networks; this application is not limited to these.
[0063] In some implementations, the generated dependency reference features include long-term reference features and / or short-term reference features.
[0064] Please see Figure 4 In the long-term dependency generation module, for the first frame, the initialized reference features and the first frame's features can be input to the long-term dependency generation network. That is, the first frame's feature input can be the decoding information of the first frame (such as predicted features, reconstructed features, etc.), and the first frame's feature output can be dependency reference features, such as long-term reference features and / or short-term reference features (here denoted as long / short-term reference features), which can be used as reference information for long-term references and passed to the next frame. For the second frame, the dependency reference features of the first frame and the decoding information of the second frame are input, and the dependency reference features of the second frame can be output. In this way, the dependency reference features of N frames can be obtained. N is a positive integer. The dependency reference features include long-term reference features and / or short-term reference features.
[0065] For example, when the long-term dependency generation network is an RNN, it can include either long-term reference features or short-term reference features. For example, when the long-term dependency generation network is an LSTM or ConvLSTM, it can include both long-term and short-term reference features. This embodiment uses a ConvLSTM long-term dependency generation network as an example, which can transmit two reference features, that is, both long-term and short-term reference features.
[0066] In some implementations, the long-term dependency generation network (i.e., the long-term dependency generation module) can be positioned before, during, or after the decoding module. The decoding module is used to decode the current frame bitstream to obtain decoding information. The decoding module can refer to any module in the framework of the aforementioned image encoding and decoding system. For example, at least one type of decoding information includes at least one of the following: prediction features, reconstruction features, decoded intermediate features, the current reconstructed frame, the predicted image frame, motion information, etc. It is understood that the decoding information can include any decoding information that can be obtained from the current frame that has been decoded during the decoding process. This application does not limit the decoding information. The long-term dependency generation module can be positioned at least one of the following positions: before, during, or after any module in the framework of the image encoding and decoding system, to enhance the transmission of long / short-term information between frames during the encoding and decoding process.
[0067] The decoding module includes at least one of the following: a frame reconstruction module and a motion compensation module. The frame reconstruction module is used to obtain reconstructed features, and the motion compensation module is used to obtain predicted features.
[0068] For example, see [link to relevant documentation]. Figure 3 The decoded information includes at least one of the following: predicted features, reconstructed features, etc., which can be used as reference features. The long-term dependency generation module is set after the frame reconstruction module (optionally) and the motion compensation module, so that the long-term dependency generation module generates dependency reference feature 1 (such as long-term reference feature 1, short-term reference feature 1) and dependency reference feature 2 (such as long-term reference feature 2, short-term reference feature 2) for the predicted features and reconstructed features, respectively. That is, the long-term dependency generated from the predicted features and reconstructed features can be used as reference information to pass to the next frame to achieve long-term information dependency, obtain long-distance features, and provide long-term dependencies for subsequent image frames.
[0069] The long-term dependency generation module is connected to the reference information module, and dependency reference features can be saved to the reference information module.
[0070] In some embodiments, continue reading Figure 3After determining the dependent reference features, a feature fusion module can be used to fuse several long-term reference features to obtain fused long-term reference features. These dependent reference features include several long-term reference features, which may correspond to the same or different decoding information. That is, the several long-term reference features can all be generated from corresponding predicted features, or all of them can be generated from corresponding reconstructed features, or they can be generated from both predicted and reconstructed features. For example, the several long-term reference features may include dependent reference feature 1 and dependent reference feature 2. For example, the several long-term reference features may include several dependent reference feature 1 or several dependent reference feature 2, etc. For instance, several long-term reference features 1 and 2 can be fused to obtain fused long-term reference features. This application does not limit the fused long-term reference features.
[0071] In some implementations, please refer to Figure 5 In the feature fusion module, a feature fusion network can be used to fuse N (N greater than or equal to 1) long-term reference features to obtain a single fused long-term reference feature. When N is 1, a single long-term reference feature can be fused to obtain a single fused long-term reference feature; this process can optimize the features. The feature fusion network includes neural networks such as concatenation networks + convolutional networks, attention networks, or residual block networks, or combinations thereof; this application does not limit the feature fusion network. The above method can achieve the fusion of multiple long-term reference features and optimize the long-term reference features.
[0072] For example, with N=2, long-term reference feature 1 and long-term reference feature 2 of the current frame are fused to obtain a fused long-term reference feature. The feature fusion network includes a channel attention network and a convolutional network. After the channel attention network performs attention processing on long-term reference feature 1 and long-term reference feature 2, the convolutional network convolves the attention processing result to obtain a single fused long-term reference feature.
[0073] The above approach proposes an effective method for fusing long-term reference features. This method effectively fuses multiple long-term reference features and can adaptively extract effective information as reference information for subsequent frame bitstreams during decoding.
[0074] In some embodiments, continue reading Figure 3After determining the dependent reference features, or after the feature fusion module fuses several long-term reference features to obtain fused long-term reference features, the feature refresh module can be used to update the long-term reference features. For example, prediction information can be obtained in the current frame, i.e., after using the decoded motion information to perform motion compensation on the corresponding dependent reference features of the previous frame, to obtain the long-term / short-term reference features of the current frame, and then the feature refresh module can be used to update the long-term reference features.
[0075] The feature refresh module updates the long-term reference features using short-term reference features corresponding to the reference frame image, resulting in updated long-term reference features. The short-term reference features are obtained through feature extraction from the reference frame image, which includes the current reconstructed frame. Specifically, the reference information module caches reference information, including reconstructed features, the reference frame image (such as the current reconstructed frame), predicted features, prediction information, decoded intermediate features, and any other available decoding information. Feature extraction can be performed on the reference frame image to obtain short-term reference features, which are then used to update the long-term reference features or fused long-term reference features, resulting in updated long-term reference features. The updated long-term reference features can be used as reference information for encoding / decoding the next frame. Using short-term reference features from the reference frame image branch to update the long-term reference features allows for optimization of the long-term reference features.
[0076] In some implementations, a preset update method can be used to update the long-term reference features using the short-term reference features corresponding to the reference frame image, resulting in updated long-term reference features. The preset update method includes any of the following: direct stitching, weighted update, etc. This application does not limit the preset update method.
[0077] Optionally, when using the direct concatenation method, the short-term reference features and long-term reference features can be concatenated and fused. For example, a concatenation network can be used to concatenate the short-term reference features and long-term reference features, and then a convolutional network can be used to fuse them to obtain the updated long-term reference features.
[0078] Optionally, when using the weight update method, the weight factor of each input branch can be learned based on the short-term reference features and the long-term reference features (a series of input branches), the weight value can be obtained, and each input branch can be adjusted separately and then fused to obtain the updated long-term reference features.
[0079] Please see Figure 6In the feature refresh module, long-term and short-term reference features can be input into the weight network. The weight network processes the long-term and short-term reference features to obtain their respective weight values. That is, weight value 1 for the long-term reference features and weight value 2 for the short-term reference features. The weight network includes weight calculation methods such as channel attention networks, spatial attention networks, and non-local attention networks. For example, the weight network includes a convolutional network + residual block network. This application does not impose any restrictions on the weight network.
[0080] Then, the long-term reference feature and the short-term reference feature are modulated using their respective weight values (weight value 1, weight value 2) to obtain the short-term modulation feature and the long-term modulation feature, respectively. The modulation methods include: channel-level multiplication, spatial-level multiplication, channel + spatial-level multiplication, and spatial feature transformation combining multiplication and addition. For example, multiplication operations can be used for adjustment; this application does not limit the modulation method.
[0081] Then, the short-term and long-term modulation features are fused to obtain updated long-term reference features, enabling feature fine-tuning / adjustment of the reference-dependent features. The fusion process can be performed using neural networks (such as convolutional networks + residual block networks). This application does not impose any restrictions on the fusion process.
[0082] As the number of frames increases, the transmitted long-term reference features may accumulate errors. Therefore, this application prioritizes using a weighted update method to adaptively select effective long / short-term reference information for fusion, thereby improving the accuracy of the updated long-term reference features. The proposed method effectively refreshes features by fully utilizing short-term reference features from the reference frame image branch to update and optimize long-term reference features, reducing long-term error accumulation.
[0083] In some embodiments, the dependency reference features obtained above can be used for encoding and decoding of subsequent frames. For example, reference information such as long-term reference features (e.g., long-term reference feature 1 / long-term reference feature 2), short-term reference features (e.g., short-term reference feature 1 / short-term reference feature 2), updated long-term reference features, fused long-term reference features, and reference frame images (e.g., the current reconstructed frame and the reconstructed frame of the preceding frame) can be used to encode and decode the current frame. Alternatively, the dependency reference features obtained during the encoding and decoding process of the preceding frame can be used to encode and decode the current frame.
[0084] For example, during the encoding and decoding of the current frame, the encoding and decoding of the current frame can be performed with reference to dependent reference features obtained from the encoding and decoding process of the previous frame, such as the encoding and decoding process of motion information, the motion compensation process, and the frame reconstruction process. The following embodiments illustrate this.
[0085] Please see Figure 7 , Figure 7 This is a flowchart illustrating a second embodiment of the image decoding method of this application. The specific steps of this embodiment can be executed using the decoding end described above. The method may include the following steps:
[0086] S21: Get the current frame bitstream.
[0087] The current frame bitstream includes a motion bitstream and a context bitstream. The motion bitstream is obtained by encoding motion information at the encoder. The context bitstream is obtained by encoding context information at the encoder. The encoder transmits the motion bitstream and context bitstream to the decoder so that the decoder can access them.
[0088] S22: Decode the current frame bitstream to obtain motion information.
[0089] Continue reading Figure 3 The motion information entropy model can decode motion bitstreams to obtain low-dimensional motion features. The motion information decoding module (i.e., the motion information decoder) upscales the decoded low-dimensional motion features and reconstructs the motion information.
[0090] In some implementations, the feature adjustment module adjusts the dependent reference features and then decodes the motion stream based on the adjusted dependent reference features.
[0091] In some implementations, the dependent reference features can be input into the feature adjustment module, which then performs feature adjustment on the dependent reference features to obtain a second reference adjusted feature. Feature adjustment includes scale adjustment, dimension adjustment, information adjustment, feature optimization, etc., and this application does not limit the scope of feature adjustment.
[0092] Then, the second reference adjustment feature and the motion bitstream are input into the motion information entropy model. The motion information entropy model processes the second reference adjustment feature and the motion bitstream to obtain the second probability parameter. This second probability parameter can be used for subsequent decoding to obtain motion information. The second probability parameter can include multiple probability parameters, and the second reference adjustment feature can be used to participate in obtaining at least one of the probability parameters.
[0093] For example, the second probability parameter includes at least one of the following: mean, variance, and quantization parameter. The second reference adjustment feature can be used to participate in obtaining at least one of the mean, variance, and quantization parameter. For example, it can participate in generating the mean, or in generating the variance, or in generating both the mean and variance, etc. This application does not impose any limitations on this.
[0094] The second probability parameter can be used to decode the motion stream to obtain motion information. Specifically, the second probability parameter and the motion stream are input into the motion information decoder to decode the motion stream and obtain the motion information.
[0095] In some implementations, the dependent reference features of this embodiment include: long-term reference features and / or short-term reference features. When the dependent reference features include both long-term and short-term reference features, the long-term and short-term reference features obtained by fusing them can be used as the dependent reference features; alternatively, the long-term and short-term reference features can be used as dependent reference features separately. When the long-term and short-term reference features are used as dependent reference features separately, motion information can be obtained from the long-term and short-term reference features respectively, and then the motion information obtained from the long-term and short-term reference features can be combined to obtain the final motion information.
[0096] For example, when relying on reference features including short-term reference features, during the decoding of the motion bitstream using the motion information entropy model, short-term reference features corresponding to a reference frame image are obtained. These short-term reference features are obtained through feature extraction of the reference frame image, which includes the reconstructed frame of the preceding frame. Subsequently, a feature adjustment module adjusts the short-term reference features corresponding to the reference frame image to obtain second reference adjusted features. Then, the second reference adjusted features and the motion bitstream are input into the motion information entropy model for processing to obtain second probability parameters. These second probability parameters are used for subsequent decoding to obtain motion information; for example, the second probability parameters include at least one of the following: mean, variance, quantization parameters, etc. In some application scenarios, during the encoding process, the motion information and the second reference adjusted features can be input into the motion information entropy model, and the second probability parameters can be obtained as described above. These second probability parameters are used to encode the motion information to obtain the motion bitstream.
[0097] Specifically, for example, regarding feature adjustment, if the scale of the input short-term reference features is relatively large, while the scale of the features required by the motion information entropy model is relatively small, a feature adjustment module can be used to scale the short-term reference features to the scale required by the motion information entropy model, thus obtaining short-term adjusted features. For example, it could be adjusted to the same dimension as the motion bitstream or motion information, thereby potentially adjusting the feature information of the short-term reference features. The feature adjustment module can employ neural networks such as convolutional networks, residual block networks, and attention networks. For example, the feature adjustment module can use a convolutional network + residual block network. This application does not impose any restrictions on the feature adjustment module. Then, the short-term adjusted features can be used as temporal prior information input into the motion information entropy model, and the motion bitstream can also be input into the motion information entropy model to decode and obtain the motion information. Feature adjustment through the feature adjustment module can improve the accuracy of parameter prediction.
[0098] The above scheme, considering the need to ensure cross-platform decoding consistency during subsequent motion entropy model quantization, sets the input of the feature adjustment module to short-term reference features. This way, during subsequent quantization, only the feature extraction modules of the motion entropy model and the reference frame image branch need to be quantized. In other words, only the short-term reference features from the reference frame image are processed and passed to the motion entropy model, reducing link dependencies and facilitating subsequent cross-platform decoding consistency.
[0099] S23: Using motion information, motion compensation is performed on the dependent reference features to obtain the prediction information of the current frame; wherein, the dependent reference features are obtained by decoding the previous frame bitstream using the image decoding method described above.
[0100] A motion compensation module is employed to perform motion compensation on dependent reference features based on decoded motion information, thereby obtaining prediction information for the current frame. This prediction information includes the predicted image frame and / or predicted features. For example, motion compensation operations such as warping are performed on the reference information using motion information, and relevant information is extracted from the reference information to obtain the prediction information for the current frame.
[0101] In some implementations, the dependent reference features may include short-term reference features and long-term reference features (long-term reference feature 1 / long-term reference feature 2 / updated long-term reference feature / fused long-term reference feature) corresponding to the reference frame image. This application does not limit the dependent reference features.
[0102] By performing motion compensation on long / short-term reference features based on decoded motion information, prediction information for the current frame and reference features to be passed to the next frame can be obtained.
[0103] In some implementations, the prediction information of the current frame includes prediction information and / or prediction features for reference. The prediction features for reference are also the reference features passed to the next frame. The prediction features for reference are used as a kind of decoding information of the current frame bitstream and can be used to generate the dependency reference features corresponding to the current frame. That is, they can be passed to the long-term dependency generation module to generate dependency reference features.
[0104] In some implementations, the prediction information and the prediction features used for reference may be the same or different. For example, this embodiment may employ an implementation where the prediction information and the prediction features used for reference are different.
[0105] In some embodiments, please refer to Figure 8 This embodiment can further extend step S23 of the above embodiment. Using motion information to perform motion compensation on the reference features to obtain the prediction information for the current frame, this embodiment may include the following steps:
[0106] S231: Using motion information, the dependent reference features of multiple branches are aligned to obtain multiple reference compensation results.
[0107] Multiple branch dependency reference features can be used as reference information, allowing for alignment / motion compensation of these features using motion information to obtain multiple reference compensation results. Each branch's dependency reference feature corresponds to a specific motion information, and the motion information for each branch can be the same or different.
[0108] In some implementations, the dependent reference features can be transformed to change their scale so that they correspond to motion information. After transformation, each branch's dependent reference feature corresponds to a piece of motion information, and the motion information corresponding to each branch can be the same or different. Then, the transformed motion information is used to align the dependent reference features of multiple branches to obtain multiple corresponding reference compensation results.
[0109] In some implementations, please refer to Figure 9 The dependent reference features of multiple branches include long-term reference features and short-term reference features. Optionally, a transformation network can be used to transform the motion information to obtain first motion information corresponding to the short-term reference feature and second motion information corresponding to the long-term reference feature, so that each branch's dependent reference feature corresponds to one motion information, and the motion information corresponding to each branch can be the same or different. Optionally, no transformation processing of the motion information is required, that is, long-term reference features and short-term reference features can share the same motion information. The transformation network can include neural networks, such as convolutional networks, etc., and this application does not limit the transformation network.
[0110] Then, using the first motion information, the short-term reference features are aligned to obtain a short-term compensation result; and using the second motion information, the long-term reference features are aligned to obtain a long-term compensation result. The alignment process can include motion compensation methods such as interpolation-based warp, deformable convolution, warp + deformable convolution, or combinations thereof with convolutional networks; this application does not limit the scope of these methods. Specifically, the warp operation transforms the corresponding macroblocks of the reference information by translation, rotation, etc., based on the motion information to simulate their position in the current frame image. In other words, motion compensation is a method for describing the difference between the reference information and the current frame image; specifically, it describes how each small block of the reference information moves to a certain position in the current frame image, thereby reducing spatial redundancy in the image sequence. Of course, in other embodiments, the motion compensation operation can also be a traditional convolution operation, etc.; this application does not limit the alignment processing method. This embodiment uses an interpolation-based warp operation as an example for illustration.
[0111] In some implementations, the method can be extended to multi-scale structures to obtain short-term / long-term reference features at different scales. Motion information is then used to perform motion compensation on the short-term / long-term reference features at different scales to obtain compensation results at different scales.
[0112] S232: Fuse multiple reference compensation results to obtain the prediction information for the current frame.
[0113] The reference compensation results from multiple branches are fused to obtain the prediction information for the current frame. This prediction information includes prediction information and / or prediction features used for reference. The prediction features used for reference represent the reference features passed to the next frame and can be stored in the reference information module for encoding and decoding in the next frame. The fusion processing methods include convolutional networks, attention networks, and convolutional networks combined with residual block networks. For example, this application can use a convolutional network combined with a residual block network to fuse multiple reference compensation results to obtain corresponding prediction information and reference features, which are not identical. This application does not limit the fusion processing method.
[0114] In some implementations, the short-term compensation results and long-term compensation results can be fused to obtain the prediction information for the current frame.
[0115] In some implementations, please refer to [the relevant documentation]. Figure 9 Extending to multi-scale structures, multi-scale short-term compensation results and multi-scale long-term compensation results can be obtained. The multi-scale short-term compensation result is obtained by aligning the multi-scale short-term reference features using motion information; the multi-scale short-term reference features are obtained by scaling the short-term reference features. Similarly, the multi-scale long-term compensation result is obtained by aligning the multi-scale long-term reference features using motion information; the multi-scale long-term reference features are obtained by scaling the long-term reference features. For example, scaling transformations include upsampling or downsampling at different factors; this application does not impose restrictions on scaling transformations.
[0116] After aligning the short-term and long-term reference features at different scales, short-term and long-term compensation results at different scales can be obtained respectively. For example, motion compensation is performed at the low-scale stage, and the low-scale compensation result can be passed to the high-scale stage for fusion. That is, compensation results at different scales can be input for fusion.
[0117] In some implementations, short-term compensation results, multi-scale short-term compensation results, long-term compensation results, and multi-scale long-term compensation results can be fused to obtain the prediction information for the current frame. For example, the multi-scale fusion processing method may include fusing short-term and long-term compensation results at different scales, such as fusing short-term and long-term compensation results at first-level, second-level, and third-level scales to obtain the prediction information for the current frame. For example, the multi-scale fusion processing method may include upsampling the third-level scale to the same scale as the second-level scale, then concatenating and using convolutional fusion to obtain the fusion features of the current scale, and so on, to obtain the prediction information for the current frame by performing convolutional fusion on the first-level scale. For example, the multi-scale fusion processing method may include upsampling the third-level scale to the same scale as the second-level scale, using convolutional fusion to obtain the fusion features, and then fusing them with the previous scales to obtain the fusion features of the current scale, thereby obtaining the prediction information for the current frame. This application does not limit the fusion method of multi-scale compensation results.
[0118] The aforementioned long-term reference features are obtained by fusing short-term and long-term reference features through feature updates. The feature refresh module described above is a primary fusion. In this embodiment, the short-term compensation results corresponding to the short-term and long-term reference features can be fused, which is a secondary fusion. This allows both the prediction information and the reference features to reference rich long / short-term information, improving the accuracy of inter-frame prediction, and thus improving the accuracy of the acquired prediction information. It can adapt to motion compensation using long / short-term reference features as reference information, improving the flexibility of motion compensation and the accuracy of prediction information.
[0119] Please see Figure 10 , Figure 10 This is a flowchart illustrating a third embodiment of the image decoding method of this application. The specific steps of this embodiment can be executed using the decoding end described above. The method may include the following steps:
[0120] S31: Get the current frame bitstream.
[0121] The current frame bitstream includes a motion bitstream and a context bitstream. The motion bitstream is obtained by encoding motion information at the encoder. The context bitstream is obtained by encoding context information at the encoder. The encoder transmits the motion bitstream and context bitstream to the decoder so that the decoder can access them.
[0122] S32: Decode the current frame bitstream to obtain the initial reconstructed features.
[0123] The decoding end decodes the current frame bitstream to obtain the initial reconstructed features.
[0124] Continue reading Figure 3 By performing decoding and motion compensation on the motion stream through the above embodiments, the prediction information of the current frame can be obtained. The prediction information includes prediction features and / or prediction image frames. Then, a context decoding module (context decoder) decodes the context stream to obtain context information. Specifically, the low-dimensional context features obtained from decoding the context stream can be up-dimensioned and reconstructed to obtain the context information. The frame reconstruction module uses the context information and prediction information to obtain the initial reconstruction information, i.e., the initial reconstruction features / current reconstructed frame. During this process, if the context information obtained by the encoder is represented as a difference, i.e., the difference between the current frame image and the prediction information, the decoder adds the prediction information to the context information to obtain the initial reconstruction information.
[0125] In some embodiments, during the decoding of the current frame bitstream, feature adjustment can be performed using dependent reference features / prediction information, and then decoding can be performed to obtain initial reconstructed information.
[0126] In some implementations, the dependent reference features and / or prediction information can be input into the feature adjustment module, which then performs feature adjustment on the dependent reference features and / or prediction information to obtain a first reference adjusted feature. Feature adjustment includes scale adjustment, dimension adjustment, information adjustment, feature optimization, etc., and this application does not limit the feature adjustment. The prediction information is obtained by motion compensation using motion information decoded from the motion bitstream. The process of obtaining the prediction information can be referred to the specific implementation process of the above embodiments, and will not be repeated here.
[0127] The context bitstream and the first reference adjustment feature are input into the context entropy model for processing to obtain the first probability parameter. The first probability parameter is used for decoding to obtain the initial reconstructed features; wherein, the first probability parameter may include multiple probability parameters, and the first reference adjustment feature may be used to participate in obtaining at least one of the first probability parameters. In some application scenarios, the first probability parameter may represent the probability parameter of the undecoded bitstream in the context bitstream.
[0128] For example, the first probability parameter includes at least one of the following: mean, variance, and quantization parameter; the first reference adjustment feature can be used to participate in obtaining at least one of the mean, variance, and quantization parameter, for example, to participate in generating the mean, or to participate in generating the variance, or to participate in generating both the mean and variance, etc. This application does not impose any limitations on this.
[0129] The first probability parameter can be used to decode the context bitstream to obtain context information. Then, the context information and prediction information are used to reconstruct the initial reconstructed features. Specifically, the first probability parameter and the context bitstream can be input into the context decoder to decode and obtain context information. Thus, the frame reconstruction module can use the prediction information and context information to obtain the initial reconstructed features. In some application scenarios, during the encoding process, the context information and the first reference adjustment feature can be input into the context entropy model, and the first probability parameter can be obtained as described above. The first probability parameter is used to encode the context information to obtain the context bitstream.
[0130] In some implementations, the dependent reference features of this embodiment include: long-term reference features and / or short-term reference features. When the dependent reference features include both long-term and short-term reference features, the long-term and short-term reference features obtained by fusing them can be used as the dependent reference features; alternatively, the long-term and short-term reference features can be used as dependent reference features separately. When the long-term and short-term reference features are used as dependent reference features separately, initial reconstruction features can be obtained from the long-term and short-term reference features respectively, and then the initial reconstruction features obtained from the long-term and short-term reference features can be combined to obtain the final initial reconstruction features.
[0131] S33: Using the dependent reference features, the initial reconstruction features are processed to obtain the reconstruction information of the current frame; wherein, the dependent reference features are obtained by decoding the previous frame bitstream using the image decoding method described above.
[0132] To improve reconstruction quality, dependent reference features (long / short-term reference features) are introduced during the frame reconstruction stage. The initial reconstruction features are processed using dependent reference features to obtain the reconstruction information of the current frame. This allows for full utilization of temporal information in frame reconstruction. Since dependent reference features are temporal information, there are motion changes in the initial reconstruction features of the current frame output by the context decoder, except for static image frames. By processing the initial reconstruction features using dependent reference features, more accurate reconstruction information of the current frame can be obtained.
[0133] In some implementations, the dependent reference features of this embodiment include: long-term reference features and / or short-term reference features. When the dependent reference features include both long-term and short-term reference features, the long-term and short-term reference features obtained by fusing them can be used as the dependent reference features; alternatively, the long-term and short-term reference features can be used as dependent reference features separately. When the long-term and short-term reference features are used as dependent reference features separately, the reconstruction information of the current frame can be obtained from the long-term and short-term reference features respectively, and then the reconstruction information of the current frame obtained from the long-term and short-term reference features can be combined to obtain the final reconstruction information of the current frame.
[0134] The reconstruction information of the current frame includes reconstruction features and / or a reference frame image (the current reconstructed frame can be used as a reference frame for subsequent frames).
[0135] The initial or processed reconstructed features are used as decoding information for the current frame bitstream and passed to the long-term dependency generation module to generate dependency reference features.
[0136] In some implementations, a preset processing method can be used to process the initial reconstruction features using dependent reference features to obtain the reconstruction information of the current frame. Examples include convolutional network methods, recurrent network methods, temporal correlation methods, etc., and this application is not limited to these.
[0137] Convolutional Network Method: This method concatenates initial reconstructed features with dependent reference features to obtain reconstructed spliced features. The input is then fed into a pre-defined convolutional network, which processes the reconstructed spliced features to obtain reconstruction weights. These weights are then used to weight the initial reconstructed features, resulting in the reconstructed information for the current frame. This method can automatically learn useful long / short-term information to compensate for the initial reconstructed information, leading to better reconstructed information, i.e., better reconstructed features (reference features) and / or the current frame image (reference frame image).
[0138] Recurrent network method: This method uses a pre-set recurrent network to process the initial reconstruction features and dependent reference features to obtain the reconstruction information of the current frame. For example, the pre-set recurrent network includes recurrent networks such as RNN and LSTM. Recurrent networks have a gating mechanism similar to attention, which can adaptively utilize important long / short-term reference features to enhance the initial reconstruction features, and can also pass long-distance information as reference information.
[0139] Please see Figure 11The pre-defined recurrent network can incorporate an LSTM recurrent network structure. For the current frame (frame N), the inputs can be the initial reconstructed features of frame N, the dependent reference features of frame N-1, and the network hidden state (i.e., the hidden state of the LSTM) of frame N-1, where N is an integer greater than 1. The hidden state of the LSTM is the set of hidden states at each time step when the model processes sequential data. The hidden state plays a crucial role in the LSTM network; it depends not only on the current input but also on the hidden state of the previous time step, thus enabling it to update and carry information to the next time step at each time step. In this embodiment, the dependent reference features include long-term reference features and / or short-term reference features. The long-short-term reference features obtained by fusing the long-term and short-term reference features can be used as dependent reference features input to the pre-defined recurrent network. The fusion process can be performed using a convolutional network.
[0140] The initial reconstructed features and long-term and short-term reference features are processed based on the network's hidden states using a pre-defined recurrent network, outputting the reconstructed features and network hidden states to be passed to the next frame. Furthermore, frame reconstruction is performed on the output reconstructed features to obtain the reconstruction information of the current frame, i.e., the reconstructed features and the reference frame image are obtained. The frame reconstruction method includes using a convolutional network + residual block network + upsampling network + residual block network. This application does not limit this approach. For example, the convolutional network + residual block network outputs the reconstructed features, and the upsampling network + residual block network outputs the reference frame image.
[0141] Temporal correlation method: It can process the initial reconstruction features based on the regional correlation between the initial reconstruction features and the dependent reference features to obtain the reconstruction information of the current frame.
[0142] In some implementations, the regional correlation between the initial reconstructed features and the dependent reference features is obtained. Regional correlation can represent the correlation between corresponding regions, the correlation between corresponding feature points, etc. Using the regional correlation, the dependent reference features are weighted to obtain the predicted values of the reconstructed information, that is, the predicted values of the corresponding regions and feature points of the initial reconstructed features. Then, the predicted values of the initial reconstructed features and the reconstructed information are fused to obtain the reconstructed information of the current frame.
[0143] Optionally, this embodiment takes the dependent reference features including long-term reference features and / or short-term reference features as an example. The long-term and short-term reference features can be fused to obtain long-term and short-term reference features, which are then used as dependent reference features for calculating regional correlation. The fusion process can employ neural networks such as convolutional networks and attention networks for feature fusion, dimensionality reduction, and other processing. This embodiment uses a concatenation network + convolutional network approach for fusion; that is, a concatenation network is used to concatenate the long-term and short-term reference features, and then a convolutional network is used for fusion to obtain the long-term and short-term reference features. This application does not limit the fusion method.
[0144] Please see Figure 12 According to a preset segmentation method, the initial reconstructed features and dependent reference features are segmented into blocks, resulting in several reconstructed feature blocks and several dependent reference blocks. The preset segmentation method can be uniform or non-uniform, with the same number of blocks, which can be divided into N blocks (N>=1). Blocks at the same location have the same size, meaning the initial reconstructed features and dependent reference features can be segmented into blocks using the same preset segmentation method. The initial reconstructed features are divided into several N reconstructed feature blocks (denoted as A1, A2, A3, ..., AN). Dependent reference features (e.g., long-term and short-term reference features) are divided into several N dependent reference blocks (denoted as B1, B2, B3, ..., BN).
[0145] For example, with N=4, the initial reconstructed features can be divided into N blocks (A1, A2, A3, A4), and the long-term and short-term features can be divided into N blocks (B1, B2, B3, B4).
[0146] The correlation between each reconstructed feature block and its corresponding dependent reference block is calculated separately to obtain the regional correlation between them. Then, using the regional correlations of each dependent reference block, a weighted processing is performed on each dependent reference block to obtain the predicted value of the reconstructed information. Specifically, a non-local attention network can be used to perform the above steps to obtain the predicted value of the reconstructed information; that is, a non-local attention network is used to perform non-local attention processing on each reconstructed feature block and its corresponding dependent reference block to obtain the predicted value of the reconstructed information. The weighted processing can include a weighted sum method, where each feature point in the predicted value of the reconstructed information is obtained by weighting the regional correlations of all feature points in the dependent reference block. The predicted value of the reconstructed information can be a predicted point or a predicted value of the reconstructed information. This embodiment uses the predicted value of the reconstructed information as an example for illustration; for ease of description, the following example is referred to as the predicted value.
[0147] For example, a non-local attention network is used to perform non-local attention processing on blocks A1 and B1 to obtain the predicted value C1 corresponding to blocks A1 and B1; a non-local attention network is used to perform non-local attention processing on blocks A2 and B2 to obtain the predicted value C2 corresponding to blocks A2 and B2; similarly, a non-local attention network is used to perform non-local attention processing on blocks AN and BN to obtain the predicted value CN corresponding to blocks AN and BN.
[0148] In some implementations, please refer to Figure 13 This paper takes the example of using a non-local attention network to process blocks A1 and B1 with non-local attention to obtain the predicted value C1 for blocks A1 and B1. Each block can contain multiple feature points. The correlation between the reconstructed feature block A1 and the corresponding long-term and short-term reference feature block B1 is calculated to obtain the regional correlation W1 between the reconstructed feature block A1 and the corresponding long-term and short-term reference feature block B1. The regional correlation W1 can represent the weight of each feature point in the block. The long-term and short-term reference feature block B1 is weighted using the regional correlation W1 to obtain the predicted value C1. That is, in the predicted value C1, each feature point is obtained by the weighted sum of all feature points in the long-term and short-term reference features. In this way, the predicted value for each corresponding block can be obtained.
[0149] Then, the predicted values of the initial reconstructed features and reconstructed information are input into the reconstruction fusion network for fusion processing to obtain the reconstructed information of the current frame. The output can be the reconstructed features and the reference frame image. The reconstruction fusion network includes, but is not limited to, neural networks such as convolutional networks, attention networks, and residual block networks.
[0150] For example, the reconstruction fusion network uses a combination of convolutional network + residual block network and upsampling network + residual block network in sequence. The first combination, convolutional network + residual block network, can output reconstructed features, which are used as reference features. The second combination, upsampling network + residual block network, outputs the current reconstructed frame, which is used as the reference frame image.
[0151] The above method can introduce long-term and short-term reference features into the frame reconstruction stage, which can improve the accuracy of reconstruction information, improve reconstruction effect, and enhance decoding performance.
[0152] It is understood that the modules and networks of the above embodiments of this application can be combined with each other to form multiple solutions, which will not be described in detail here.
[0153] Please see Figure 14 , Figure 14 This is a flowchart illustrating an embodiment of the image encoding method of this application. The specific steps of this embodiment can be executed using the encoding terminal described above. The method may include the following steps:
[0154] S41: Obtain the current frame image to be encoded.
[0155] The current frame image can be the image that needs to be encoded, and this application does not impose any restrictions on this.
[0156] S42: Encode the current frame image to obtain the current frame bitstream; the current frame bitstream is provided to the decoding end so that the decoding end can use any of the above image decoding methods to decode the current frame bitstream.
[0157] Continue reading Figure 3The encoding / decoding system can execute the steps of each module on the current frame image according to its framework to obtain the current frame bitstream. The current frame bitstream includes a motion bitstream and a context bitstream. During encoding, the reconstructed features and predicted features obtained from decoding the current frame bitstream can be used as reference information. For example, the generation of this reference information depends on the reference features being passed to subsequent frames for encoding. Unlike the decoding process at the decoding end, during encoding, a motion estimation module performs motion estimation on the reference frame image and the current frame image to obtain motion information. The motion information encoding module can obtain the cached motion context and short-term reference features extracted from the reference frame image to encode the motion information. Additionally, the feature adjustment module adjusts the short-term reference features and inputs them into the motion information entropy model to encode the motion information, obtaining the motion bitstream. Then, the motion bitstream is decoded to obtain the decoded motion information, and motion compensation is performed to obtain the prediction information for the current frame. The context encoding module obtains the context information of the current frame image and the prediction information of the current frame, and then encodes the context information to obtain the context bitstream. Furthermore, the encoding and decoding of other modules involved in the encoding process can be specifically referred to in the description of the above embodiments, and will not be repeated here. By passing dependent reference features during the encoding and decoding process, the performance and accuracy of encoding and decoding can be improved.
[0158] Therefore, the encoding end provides the current frame bitstream to the decoding end, which can then use any of the aforementioned image decoding methods to decode the current frame bitstream. The specific implementation of this step, where the decoding end uses any of the aforementioned image decoding methods to decode the current frame bitstream, can be found in the specific implementation process of the decoding end described above, and will not be elaborated upon here.
[0159] It is understood that in the above method of specific implementation, the order in which each step is written does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0160] For the above embodiments, this application provides a decoding end and an encoding end, used to implement the steps of any embodiment of the image decoding method and image encoding method described above.
[0161] Please see Figure 15 , Figure 15 This is a schematic diagram of the structure of the first embodiment of the decoding end of this application. The decoding end 50 includes a bitstream module 51, a decoding module 52, and a generation module 53.
[0162] The bitstream module 51 is used to obtain the current frame bitstream.
[0163] The decoding module 52 is used to decode the current frame bitstream to obtain at least one type of decoding information.
[0164] The generation module 53 is used to determine a dependency reference feature based on at least one decoding information; wherein the dependency reference feature is used to pass to the subsequent frame bitstream so as to decode the subsequent frame bitstream using the dependency reference feature.
[0165] Please see Figure 16 , Figure 16 This is a schematic diagram of the structure of the second embodiment of the decoding end of this application. The decoding end 60 includes a bitstream module 61, a decoding module 62, and a reconstruction module 63.
[0166] The bitstream module 61 is used to obtain the current frame bitstream.
[0167] The decoding module 62 is used to decode the current frame bitstream to obtain the initial reconstructed features.
[0168] The reconstruction module 63 is used to process the initial reconstruction features using the dependent reference features to obtain the reconstruction information of the current frame; wherein, the dependent reference features are obtained by decoding the previous frame bitstream using the image decoding method described above.
[0169] Please see Figure 17 , Figure 17 This is a schematic diagram of the structure of the third embodiment of the decoding end of this application. The decoding end 70 includes a bitstream module 71, a decoding module 72, and a compensation module 73.
[0170] The bitstream module 71 is used to obtain the current frame bitstream.
[0171] The decoding module 72 is used to decode the current frame bitstream to obtain motion information.
[0172] The compensation module 73 is used to perform motion compensation on the dependent reference features using motion information to obtain the prediction information of the current frame; wherein, the dependent reference features are obtained by decoding the previous frame bitstream using the image decoding method described above.
[0173] Please see Figure 18 , Figure 18 This is a schematic diagram of the structure of an embodiment of the encoding terminal of this application. The encoding terminal 80 includes an acquisition module 81 and an encoding module 82.
[0174] The acquisition module 81 is used to acquire the current frame image to be encoded.
[0175] The encoding module 82 is used to encode the current frame image to obtain the current frame bitstream; the current frame bitstream is used to provide the decoding end so that the decoding end can use any of the above image decoding methods to decode the current frame bitstream.
[0176] It should be noted that the decoding end and encoding end provided in the above embodiments belong to the same concept as the image decoding method and image encoding method provided in the corresponding embodiments. The specific ways in which each module and unit performs operations have been described in detail in the method embodiments and will not be repeated here. In practical applications, the decoding end and encoding end provided in the above embodiments can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. This application does not impose any limitations on this. The specific implementation methods of the decoding end and encoding end in the above embodiments can be referred to the implementation process of the above embodiments and will not be repeated here.
[0177] Regarding the above embodiments, this application provides a computer device; please refer to [link / reference]. Figure 19 , Figure 19 This is a schematic diagram of the structure of a computer device according to an embodiment of this application. The computer device 200 includes a memory 201 and a processor 202, wherein the memory 201 and the processor 202 are coupled to each other. The memory 201 stores program data, and the processor 202 executes the program data to implement the steps of any embodiment of the image decoding method and image encoding method described above. The computer device 200 can serve as the encoding end and / or decoding end in the video encoding and decoding system of the above embodiments, executing the steps of any embodiment of the image decoding method and image encoding method, as well as any reasonable combination thereof.
[0178] In this embodiment, processor 202 can also be referred to as CPU (Central Processing Unit). Processor 202 may be an integrated circuit chip with signal processing capabilities. Processor 202 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. The general-purpose processor can be a microprocessor, or processor 202 can be any conventional processor.
[0179] The methods described in the above embodiments can be implemented as computer programs; therefore, this application proposes a computer-readable storage medium. Please refer to [link to relevant documentation]. Figure 20 , Figure 20This is a schematic diagram of a computer-readable storage medium according to an embodiment of the present application. The computer-readable storage medium 300 stores program data 301 that can be executed by a processor. The program data 301 can be executed by the processor to implement the steps of any embodiment of the image decoding method and image encoding method described above, as well as the methods provided by any reasonable combination thereof.
[0180] In one embodiment, the program data 301 can be stored in the aforementioned computer-readable storage medium 300 as a program file in the form of a software product, so that a computer device (which may be a personal computer, server, or network device, etc.) or processor can execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned computer-readable storage medium 300 may include various media capable of storing program data 301, such as a USB flash drive, portable hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk; or it may be a terminal device such as a computer, server, mobile phone, or tablet storing the program data 301. This terminal device can send the stored program data 301 to other devices for execution, or it may self-execute the stored program data 301. This application does not limit this.
[0181] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments. For the sake of brevity, this application will not repeat the details here.
[0182] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to each other. For the sake of brevity, the present application will not repeat them here.
[0183] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. An image decoding method, characterized in that, include: Get the current frame bitstream; The current frame bitstream is decoded to obtain at least one type of decoding information; Based on the at least one decoding information, a dependency reference feature is determined; wherein the dependency reference feature is used to pass to the subsequent frame bitstream so as to decode the subsequent frame bitstream using the dependency reference feature.
2. The method according to claim 1, characterized in that, The at least one type of decoding information includes at least one of the following: predicted features and reconstructed features; And / or, determining the dependent reference feature based on the at least one decoding information includes: The long-term dependency generation network is used to process the decoded information and the dependency reference features of the preceding frame to obtain the dependency reference features corresponding to the decoded information; wherein, the dependency reference features include long-term reference features and / or short-term reference features.
3. The method according to claim 2, characterized in that, The decoding information is obtained by decoding the current frame bitstream using the decoding module, and the long-term dependency generation network is set before the decoding module, in the decoding module, or after the decoding module. The decoding module includes at least one of the following: a frame reconstruction module and a motion compensation module, wherein the frame reconstruction module is used to obtain reconstructed features and the motion compensation module is used to obtain predicted features.
4. The method according to claim 2, characterized in that, The dependency reference feature includes several long-term reference features, which correspond to the same decoding information or different decoding information; after determining the dependency reference feature based on the at least one decoding information, the process includes: The aforementioned long-term reference features are fused to obtain fused long-term reference features.
5. The method according to claim 1 or 4, characterized in that, The method further includes: The long-term reference features are updated using the short-term reference features corresponding to the reference frame image to obtain the updated long-term reference features, wherein the short-term reference features are obtained by feature extraction from the reference frame image.
6. The method according to claim 5, characterized in that, The step of updating the long-term reference features using the short-term reference features corresponding to the reference frame image to obtain the updated long-term reference features includes: The short-term reference features and the long-term reference features are processed using a weighted network to obtain their respective weight values; The short-term reference feature and the long-term reference feature are modulated using their respective weight values to obtain short-term modulation features and long-term modulation features. The short-term modulation features and the long-term modulation features are fused to obtain the updated long-term reference features; or, The short-term reference features and the long-term reference features are spliced and fused to obtain the updated long-term reference features.
7. An image decoding method, characterized in that, include: Get the current frame bitstream; The current frame bitstream is decoded to obtain the initial reconstructed features; The initial reconstruction features are processed using the dependent reference features to obtain the reconstruction information of the current frame; wherein the dependent reference features are obtained by decoding the preceding frame bitstream using the method described in any one of claims 1 to 6.
8. The method according to claim 7, characterized in that, The process of using dependent reference features to process the initial reconstruction features to obtain the reconstruction information of the current frame includes: Obtain the regional correlation between the initial reconstructed features and the dependent reference features; By utilizing the regional correlation, the dependent reference features are weighted to obtain the predicted value of the reconstructed information; The initial reconstruction features and the predicted values of the reconstruction information are fused together to obtain the reconstruction information of the current frame.
9. The method according to claim 8, characterized in that, The step of obtaining the regional correlation between the initial reconstructed features and the dependent reference features includes: According to the preset block division method, the initial reconstruction features and the dependent reference features are divided into blocks respectively to obtain several reconstruction feature blocks and several dependent reference blocks; The correlation between each reconstructed feature block and its corresponding dependent reference block is calculated to obtain the regional correlation between each reconstructed feature block and its corresponding dependent reference block. The step of using the regional correlation to weight the dependent reference features to obtain the predicted value of the reconstructed information includes: By utilizing the regional correlations corresponding to each dependent reference block, weighted processing is performed on each dependent reference block to obtain the predicted value of the reconstructed information.
10. The method according to claim 7, characterized in that, The process of using dependent reference features to process the initial reconstruction features to obtain the reconstruction information of the current frame includes: The initial reconstruction features and the dependent reference features are processed using a preset recurrent network to obtain the reconstruction information of the current frame.
11. The method according to claim 7, characterized in that, The process of using dependent reference features to process the initial reconstruction features to obtain the reconstruction information of the current frame includes: The initial reconstructed features and the dependent reference features are spliced together to obtain the reconstructed spliced features; The reconstructed splicing features are processed using a pre-defined convolutional network to obtain reconstruction weight values; The initial reconstruction features are weighted using the reconstruction weight values to obtain the reconstruction information of the current frame.
12. The method according to claim 7, characterized in that, The current frame bitstream includes a context bitstream and a motion bitstream; decoding the current frame bitstream to obtain initial reconstructed features includes: The dependent reference features and / or prediction information are adjusted to obtain a first reference adjusted feature; wherein the prediction information is obtained by motion compensation using the motion information decoded from the motion bitstream. The context code stream and the first reference adjustment feature are input into the context entropy model to obtain the first probability parameter, wherein the first probability parameter is used to decode to obtain the initial reconstruction feature, and the first reference adjustment feature is used to participate in obtaining at least one probability parameter among the first probability parameters.
13. The method according to claim 7 or 12, characterized in that, The dependent reference features include: long-term reference features and / or short-term reference features. When the dependent reference features include long-term reference features and short-term reference features, the long-term reference features and the short-term reference features are fused together to obtain the long-term and short-term reference features as dependent reference features, or the long-term reference features and the short-term reference features are used as dependent reference features respectively. And / or, the reconstruction information includes reconstruction features and / or reference frame images; And / or, the initial reconstruction features or the processed reconstruction features are used as decoding information for the current frame bitstream.
14. An image decoding method, characterized in that, include: Get the current frame bitstream; Decode the current frame bitstream to obtain motion information; Using the motion information, motion compensation is performed on the dependent reference features to obtain the prediction information of the current frame; wherein, the dependent reference features are obtained by decoding the preceding frame bitstream using the method described in any one of claims 1 to 6.
15. The method according to claim 14, characterized in that, The step of using the motion information to perform motion compensation on the reference features to obtain the prediction information for the current frame includes: Using the motion information, the dependent reference features of multiple branches are aligned to obtain multiple reference compensation results; The multiple reference compensation results are fused to obtain the prediction information for the current frame.
16. The method according to claim 15, characterized in that, The dependent reference features of the multiple branches include long-term reference features and short-term reference features; before aligning the dependent reference features of the multiple branches using the motion information to obtain multiple reference compensation results, the process includes: The motion information is transformed to obtain first motion information corresponding to short-term reference features and second motion information corresponding to long-term reference features; The motion information is used to align the dependent reference features of multiple branches to obtain multiple reference compensation results, including: Using the first motion information, the short-term reference features are aligned to obtain a short-term compensation result; and Using the second motion information, the long-term reference features are aligned to obtain the long-term compensation result.
17. The method according to claim 16, characterized in that, The process of fusing the multiple reference compensation results to obtain the prediction information of the current frame includes: The short-term compensation result and the long-term compensation result are fused to obtain the prediction information of the current frame.
18. The method according to claim 15 or 16, characterized in that, Before fusing the multiple reference compensation results to obtain the prediction information of the current frame, the process includes: Obtain multi-scale short-term compensation results and multi-scale long-term compensation results; wherein, the multi-scale short-term compensation results are obtained by aligning multi-scale short-term reference features using the motion information, and the multi-scale short-term reference features are obtained by scaling the short-term reference features; the multi-scale long-term compensation results are obtained by aligning multi-scale long-term reference features using the motion information, and the multi-scale long-term reference features are obtained by scaling the long-term reference features. The process of fusing the multiple reference compensation results to obtain the prediction information of the current frame includes: The short-term compensation result, the multi-scale short-term compensation result, the long-term compensation result, and the multi-scale long-term compensation result are fused to obtain the prediction information of the current frame.
19. The method according to claim 14, characterized in that, The current frame bitstream includes a motion bitstream; decoding the current frame bitstream to obtain motion information includes: The dependent reference features are adjusted to obtain the second reference adjusted features; The second reference adjustment feature and the motion bitstream are input into the motion information entropy model for processing to obtain a second probability parameter, wherein the second probability parameter is used for decoding to obtain the motion information; the second reference adjustment feature is used to participate in obtaining at least one probability parameter among the second probability parameters.
20. The method according to claim 14 or 19, characterized in that, The dependent reference features include: long-term reference features and / or short-term reference features. When the dependent reference features include long-term reference features and short-term reference features, the long-term reference features and the short-term reference features are fused together to obtain the long-term and short-term reference features as dependent reference features, or the long-term reference features and the short-term reference features are used as dependent reference features respectively. And / or, the prediction information of the current frame includes prediction information and / or prediction features for reference; wherein the prediction features for reference are used as decoding information of the current frame bitstream; And / or, the prediction information and the prediction features used for reference are different or the same.
21. An image encoding method, characterized in that, include: Obtain the current frame image to be encoded; The current frame image is encoded to obtain the current frame bitstream; The current frame bitstream is provided to the decoding end so that the decoding end can decode the current frame bitstream using the method of any one of claims 1 to 20.
22. A computer device, characterized in that, It includes a memory and a processor coupled to each other, the memory storing program data, and the processor executing the program data to implement the steps of the method of any one of claims 1 to 20, and / or, to implement the steps of the method of claim 21.
23. A computer-readable storage medium, characterized in that, The system stores program data that can be executed by a processor, the program data being used to implement the steps of the method according to any one of claims 1 to 20, and / or to implement the steps of the method according to claim 21.