Image decoding method, image coding method and related equipment
By acquiring a target reference frame with a preset number of frames, the motion stream is decoded and motion compensated, which solves the problem of low image decoding accuracy and improves image encoding and decoding performance.
Patent Information
- Application Number
- CN202411061130.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-02
- Publication Date
- 2026-02-10
AI Technical Summary
Existing image encoding and decoding methods have low image decoding accuracy.
By acquiring a preset number of target reference frames, decoding the motion stream to obtain motion information, and performing motion compensation on the preset number of target reference frames based on the motion information to obtain target prediction information, the accuracy of image decoding is improved.
It improves the accuracy of image encoding and decoding and enhances encoding and decoding performance.
Smart Images

Figure CN121509679A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image encoding and decoding, and in particular to an image decoding method, an image encoding method, and related equipment. Background Technology
[0002] Because video images are large in size, they typically need to be encoded and compressed. The compressed video image data is called a video stream. The video stream can be transmitted to a decoding end via wired or wireless network for decoding and viewing. The entire image encoding and compression process can include prediction, transformation, quantization, and encoding. This reduces the amount of video data, thereby reducing network bandwidth usage during transmission and minimizing storage space.
[0003] However, current image encoding and decoding methods suffer from problems such as low decoding accuracy. Summary of the Invention
[0004] The main technical problem addressed by this application is to provide an image decoding method, an image encoding method, and related equipment that can improve the accuracy of image decoding.
[0005] To address the aforementioned issues, the first aspect of this application provides an image decoding method, comprising: acquiring a target reference frame of a preset number of frames; decoding a motion stream to obtain motion information; wherein the motion stream is obtained by encoding the motion information at an encoding end, and the motion information is obtained by motion estimation based on the current image frame and the target reference frame of the preset number of frames; performing motion compensation on the target reference frame of the preset number of frames based on the motion information to obtain target prediction information; wherein the target prediction information is used to decode the current reconstructed frame of the current image frame.
[0006] To address the aforementioned issues, a second aspect of this application provides an image encoding method. This method includes: acquiring a current image frame and determining a target reference frame of a preset number of frames; performing motion estimation on the current image frame and the target reference frame of the preset number of frames to obtain motion information; and performing motion compensation on the target reference frame of the preset number of frames based on the motion information to obtain target prediction information, wherein the motion information and the target prediction information are used to encode the current frame bitstream of the current image frame.
[0007] To address the aforementioned problems, a third aspect of this application provides a computer device comprising a memory and a processor coupled to each other, the memory storing program data and the processor executing the program data to implement any step of any of the methods described above.
[0008] To address the aforementioned problems, a fourth aspect of this application provides a computer-readable storage medium storing program data executable by a processor, the program data being used to implement any step of any of the methods described above.
[0009] In the above scheme, this application obtains a preset number of target reference frames, decodes the motion bitstream to obtain motion information. Since the motion bitstream is obtained by encoding motion information at the encoding end, and the motion information is obtained by motion estimation based on the current image frame and the preset number of target reference frames, it can adapt to the preset number of target reference frames to decode motion information. Then, based on the motion information, motion compensation is performed on the preset number of target reference frames to obtain target prediction information. By referring to the preset number of target reference frames, it can adapt to the preset number of target reference frames for prediction, improving the accuracy of prediction. The predicted target prediction information is used to decode the current image frame to obtain the current reconstructed frame, which can improve the accuracy of image encoding and decoding, thereby improving encoding and decoding performance.
[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in this application, the accompanying drawings required in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Among them:
[0012] Figure 1 This is a schematic diagram of the structure of an embodiment of the image encoding and decoding system of this application;
[0013] Figure 2 This is a flowchart illustrating the first embodiment of the image decoding method of this application;
[0014] Figure 3 This is a schematic diagram of the framework of an embodiment of the image encoding and decoding system of this application;
[0015] Figure 4 This is a flowchart illustrating the second embodiment of the image decoding method of this application;
[0016] Figure 5 This is an example schematic diagram of an embodiment of a candidate reference frame of this application;
[0017] Figure 6 This is an example schematic diagram of another embodiment of the candidate reference frame of this application;
[0018] Figure 7This is an example schematic diagram of an embodiment of the target reference frame of this application;
[0019] Figure 8 This is an example schematic diagram of another embodiment of the target reference frame of this application;
[0020] Figure 9 This is an example schematic diagram of one embodiment of the quality assessment method of this application;
[0021] Figure 10 This is an example schematic diagram of one embodiment of the cost assessment method of this application;
[0022] Figure 11 This is a flowchart illustrating the third embodiment of the image decoding method of this application;
[0023] Figure 12 This is an example schematic diagram of an embodiment of the motion information encoding / decoding process of this application;
[0024] Figure 13 This is an example schematic diagram of another embodiment of the motion information encoding / decoding process of this application;
[0025] Figure 14 This is an example schematic diagram of another embodiment of the encoding / decoding process of motion information in this application;
[0026] Figure 15 This is a flowchart illustrating the fourth embodiment of the image decoding method of this application;
[0027] Figure 16 This is an example schematic diagram of an embodiment of motion compensation in this application;
[0028] Figure 17 This is an example schematic diagram of another embodiment of motion compensation in this application;
[0029] Figure 18 This is a flowchart illustrating the fifth embodiment of the image decoding method of this application;
[0030] Figure 19 This is an example schematic diagram of an embodiment of the fusion processing of this application;
[0031] Figure 20 This is a flowchart illustrating the sixth embodiment of the image decoding method of this application;
[0032] Figure 21 This is a flowchart illustrating an embodiment of the image encoding method of this application;
[0033] Figure 22 This is a schematic diagram of the structure of an embodiment of the decoding end of this application;
[0034] Figure 23 This is a schematic diagram of the structure of an embodiment of the encoding end of this application;
[0035] Figure 24 This is a schematic diagram of the structure of an embodiment of the computer device of this application;
[0036] Figure 25 This is a schematic diagram of the structure of an embodiment of the computer-readable storage medium of this application. Detailed Implementation
[0037] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0038] The terms "first" and "second" in this application are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.
[0039] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0040] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this document means two or more. Moreover, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of objects. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0041] This application provides the following embodiments, and each embodiment is described in detail below.
[0042] Please see Figure 1 , Figure 1 This is a schematic diagram of the structure of an embodiment of the image encoding and decoding system of this application.
[0043] The image encoding / decoding system 100 includes an encoding end 101 and a decoding end 102. The encoding end 101 and the decoding end 102 can be computer equipment, electronic equipment, etc., and can be any device with processing capabilities, such as a computer, server, mobile phone, tablet, etc. This application does not impose any limitations on this. The encoding end 101 and the decoding end 102 can communicate with each other and can be used to perform encoding and / or decoding operations on images / videos.
[0044] The encoding end 101 can be used to perform preprocessing and encoding / compression steps for images / videos to obtain bitstream data. The encoding end 101 can transmit the bitstream data to the decoding end 102. The decoding end 102 can receive the bitstream data from the encoding end 101 and perform decoding and other related steps involving the bitstream data, as well as steps related to backend vision tasks, such as image / video processing and classification.
[0045] This application provides an image decoding method, an image encoding method, and related equipment to improve the accuracy of image decoding during the encoding and decoding process. In some embodiments, the encoding end 101 and the decoding end 102 of this embodiment can be used to implement any step of the following embodiments.
[0046] Please see Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the image decoding method of this application. The specific steps of this embodiment can be executed using the decoding end described above. The method may include the following steps:
[0047] S11: Obtain the target reference frame of the preset frame number.
[0048] The encoding end can encode the current image frame to obtain the current frame bitstream, and the decoding end can obtain the current frame bitstream and decode it to obtain the current reconstructed frame of the current image frame. The current frame bitstream includes at least a motion bitstream and a context bitstream, which can be used to decode the current image frame to obtain the reconstructed frame, that is, the decoded image.
[0049] The method involves acquiring a preset number of target reference frames. These target reference frames can be unidirectional or bidirectional. The preset number of frames can be one, two, or more, and can be pre-set or adaptively set. For example, the preset number of frames can be set during the encoding process or before encoding. This application does not limit the specific value or setting method of the preset number of frames.
[0050] In some implementations, a reference information buffer may be used to cache reference information. Decoded reference information is retrieved, and a target reference frame is determined based on this information. The target reference frame includes multiple decoded target reference frames located in a preset direction of the current image frame. This preset direction may be unidirectional or bidirectional. The target reference frame / reference information includes at least one of the following: a decoded reconstructed frame, reconstructed features of the reconstructed frame, decoded intermediate features, motion-compensated prediction information, etc. The decoded intermediate features may be, for example, cached information generated from multiple reconstructed information. The motion-compensated prediction information may include prediction information obtained after motion compensation, such as a prediction frame or prediction features. For example, the prediction information here may be the prediction information corresponding to a decoded reconstructed frame. For instance, it may be the prediction information obtained by decoding the motion stream of the previous frame before decoding the current image frame, and then using the motion information to perform motion compensation on the target reference frame of the previous frame. When there are multiple target reference frames, the prediction information may be the fused target prediction information obtained through motion compensation or multiple target prediction information before fusion. It is understood that the target reference frame may include any reconstruction information that can be obtained during the decoding process of the decoded reconstructed frame. This application does not impose any restrictions on the target reference frame.
[0051] In some implementations, the target reference frame can be any frame. For example, the target reference frame can include any number of image frames of the current image frame, either unidirectionally or bidirectionally, up to a preset number. This application does not limit the number of target reference frames, the preset direction, etc. It is understood that since adjacent frames have the greatest correlation, adjacent frames can be given priority, that is, the adjacent frames of the current image frame can be used as target reference frames.
[0052] S12: Decode the motion stream to obtain motion information; wherein, the motion stream is obtained by encoding motion information at the encoding end, and the motion information is obtained by motion estimation based on the current image frame and a target reference frame with a preset number of frames.
[0053] During the encoding process of the current image frame, the encoder can perform motion estimation on the current image frame and a preset number of target reference frames to obtain motion information. The basic idea of motion estimation is to divide the current image frame into many non-overlapping image blocks (e.g., each image block could be 16×16 pixels). Then, for each image block within the search range of the target reference frames, according to a preset matching rule, the image block that best matches the current block (the current image block) in the target reference frames is identified as the matching block. The relative displacement between the matching block and the current block is the motion information. In other words, by analyzing the pixel changes between the current image frame and the target reference frames, the motion information between the current image frame and the preset number of target reference frames can be determined, i.e., motion vectors. These motion vectors can be used for prediction and motion compensation in subsequent processes to reduce redundant information in the image sequence. In one application scenario, a motion estimation network can be used to perform motion estimation on the current image frame and the preset number of target reference frames to obtain motion information. For example, the motion estimation network can be an optical flow network, a MEMC (Motion Estimation and Motion Compensation Driven Neural) network, or a flow network, etc. This application does not specify the method of obtaining motion information.
[0054] In some implementations, when the preset number of target reference frames is no more than 1, motion estimation can be directly performed on the current image frame and the preset number of target reference frames to obtain motion information. When the preset number of target reference frames is greater than 1, multiple target reference frames can be fused into a single target reference frame, and then motion estimation can be performed on the current image frame and the fused single target reference frame to obtain motion information. Alternatively, motion estimation can be performed separately on the current image frame and the preset number of target reference frames to obtain a preset amount of motion information. This application does not limit the motion estimation. The fusion method for multiple target reference frames is not limited.
[0055] The encoder encodes the motion information to obtain a motion stream, which is then transmitted to the decoder. The decoder then acquires the motion stream, decodes it, and obtains the decoded motion information.
[0056] S13: Based on motion information, perform motion compensation on a preset number of target reference frames to obtain target prediction information; wherein, the target prediction information is used to decode and obtain the current reconstructed frame of the current image frame.
[0057] Based on the decoded motion information, motion compensation is performed on a preset number of target reference frames to obtain target prediction information. Motion compensation includes warping operations, global motion compensation, block motion compensation, or its variant variable block motion compensation. Among these, the warping operation, based on the motion information, performs translation, rotation, and other transformations on the corresponding macroblocks of the target reference frame to simulate their position in the current image frame. In other words, motion compensation is a method for describing the difference between the target reference frame and the current image frame; specifically, it describes how each small block of the target reference frame moves to a certain position in the current image frame, thereby reducing spatial redundancy in the image sequence. Of course, in other embodiments, the motion compensation operation can also be a traditional convolution operation, etc. This application does not limit the method of motion compensation.
[0058] Then, the target prediction information can be used to obtain context information to decode the current reconstructed frame of the current image frame, that is, the decoded image of the current image frame.
[0059] In the above scheme, this application obtains a preset number of target reference frames, decodes the motion bitstream to obtain motion information. Since the motion bitstream is obtained by encoding the motion information at the encoding end, the motion information is obtained by motion estimation of the current image frame and the preset number of target reference frames. It can adapt to the preset number of target reference frames to decode the motion information. Then, based on the motion information, motion compensation is performed on the preset number of target reference frames to obtain target prediction information. By referring to the preset number of target reference frames, it can adapt to the preset number of target reference frames to make predictions, thereby improving the accuracy of prediction. The predicted target prediction information is used to decode the current image frame to obtain the current reconstructed frame, which can improve the accuracy of image encoding and decoding, thereby improving the image encoding and decoding performance.
[0060] Please see Figure 3 , Figure 3 This is a schematic diagram of the framework of an embodiment of the image encoding and decoding system of this application. The image encoding and decoding system includes a reference information module, a preset frame selection module (optional), a motion estimation module, a motion encoding / decoding module, a motion compensation module, a prediction fusion module (optional), a context encoding module (i.e., a context encoder), an entropy model, and a context decoding module (i.e., a context decoder).
[0061] The system includes the following modules: a reference information module for caching reference information; a preset frame selection module (optional) for selecting a target reference frame from the cached reference information; a motion estimation module for estimating motion between the current image frame and the target reference frame to obtain motion information; a motion encoding / decoding module for encoding the motion information to obtain a motion bitstream, or decoding the motion bitstream to obtain decoded motion information; a motion compensation module for performing motion compensation on the target reference frame based on the motion information to obtain target prediction information; an optional prediction fusion module for fusing the target prediction information; a context encoding module for reducing and compressing the context information between the current image frame and the target reference frame to obtain context information; an entropy model for encoding the context information, obtaining the probability of each character appearing in the quantized context information to be encoded, and performing arithmetic encoding to obtain the context bitstream; and a context decoding module for decoding the context bitstream to obtain the context information. The following section combines... Figure 3 The various modules will be explained.
[0062] In some embodiments, step S11 of the above embodiments can be further extended. The step of obtaining a target reference frame of a preset number of frames can employ a preset frame selection module to determine the target reference frame based on frame selection related information, so as to adaptively select frames to obtain the target reference frame.
[0063] In some embodiments, please refer to Figure 4 , Figure 4 This is a flowchart illustrating a second embodiment of the image decoding method of this application. The specific steps of this embodiment can be executed using the decoding end described above. The method may include the following steps:
[0064] S21: Obtain the decoded reference information, which contains several candidate reference frames.
[0065] Combination Figure 3 The reference information module can be used to cache reference information. Any reconstruction information that can be obtained during the decoding process of a decoded reconstructed frame can be stored in the reference information module, serving as several candidate reference frames so that the reference information can be selected as the target reference frame.
[0066] In some embodiments, the plurality of candidate reference frames includes: multiple candidate reference frames located in a preset direction of the current image frame, wherein the preset direction includes unidirectional or bidirectional directions, and the candidate reference frames include at least one of the following: a decoded reconstructed frame, reconstructed features of the reconstructed frame, decoded intermediate features, motion-compensated prediction information, etc. It is understood that the plurality of candidate reference frames may include any reconstructed information that can be obtained during the decoding process of the decoded reconstructed frame. This application does not limit the reference information.
[0067] Optionally, a unidirectional candidate reference frame may include several candidate reference frames in either the forward or backward direction. For example, see [link to relevant documentation]. Figure 5 Let the current frame be time t. A one-way candidate reference frame can include candidate reference frames of M frames before time t (denoted as the temporal forward frame). For example, a one-way candidate reference frame can include candidate reference frames of N frames after time t (denoted as the temporal backward frame). M and N are positive integers.
[0068] Optionally, bidirectional candidate reference frames can include a preset number of candidate reference frames for both forward and backward directions. For example, see [link to relevant documentation]. Figure 6 Let the current frame be time t. Bidirectional candidate reference frames can include candidate reference frames from M frames preceding time t. For example, unidirectional candidate reference frames can include candidate reference frames from N frames following time t. M and N are positive integers.
[0069] S22: Using frame selection information, select a preset number of candidate reference frames from several candidate reference frames and determine them as the target reference frames.
[0070] Among them, the frame selection related information includes at least one of the decoding performance conditions of the decoding end and the preset frame selection information, and the frame selection related information is used to determine at least the preset number of frames.
[0071] Decoding performance conditions include the operating performance, processing performance, and memory performance of the decoding end, and this application is not limited to these. The preset frame selection information includes at least one of the preset number of target reference frames and the frame index of the target reference frames; the preset frame selection information is obtained by decoding the frame selection information bitstream, which is encoded by the encoding end using a preset frame selection method based on the preset number of target reference frames and the frame index determined from several candidate reference frames.
[0072] In the above manner, the decoding end can adaptively determine the preset number of frames based on the preset frame selection information and the decoding performance conditions of the decoding end, so as to adaptively select the target reference frames of the preset number of frames. For example, the number of target reference frames can be reduced, increased, or the same number of frames can be used compared to the encoding end. When the number of target reference frames is increased, the target reference frames can be obtained based on the reference information and motion information can be transmitted based on the decoded reconstructed frames. This application does not impose any restrictions on this.
[0073] In some implementations, the current frame bitstream also includes a frame selection information bitstream. The frame selection information bitstream is decoded to obtain preset frame selection information. This preset frame selection information includes at least one of the preset number of target reference frames and the frame index of the target reference frames. In some application scenarios, since the target number of target reference frames and the set of frame indices (frame IDs) L are not fixed during encoding, the frame index L of the target reference frames can be encoded to obtain the frame selection information bitstream, which is then transmitted to the decoding end. The decoding end decodes the frame selection information bitstream to obtain the preset frame selection information. In some application scenarios, where the number of target reference frames and the frame index are different for different image frames, the encoding end can encode at least one of the preset number of target reference frames and the frame index to obtain the frame selection information bitstream. This application does not limit the frame selection information bitstream.
[0074] Combination Figure 3 The preset frame selection module can be used to select target reference frames. Specifically, it can utilize preset frame selection information, such as selecting a preset number of candidate reference frames from several candidate reference frames based on the frame index, and determining these as target reference frames to achieve an adaptive frame selection mechanism. The target reference frames include multiple decoded target reference frames located in a preset direction of the current image frame; the preset direction can be unidirectional or bidirectional.
[0075] Optionally, a predetermined number of target reference frames can be selected from a plurality of candidate reference frames in a unidirectional manner. For example, please refer to Figure 7 Let the current frame be time t. We can select two adjacent frames (before time t) from the candidate reference frames of the M frames in the forward direction as the target reference frames.
[0076] Optionally, a predetermined number of target reference frames can be selected from a plurality of bidirectional candidate reference frames. For example, see [link to relevant documentation]. Figure 8 Let the current frame be time t. From several candidate reference frames in the forward and backward directions, select the frame adjacent to the reference frame in the forward direction (before time t) and the frame adjacent to the reference frame in the backward direction (after time t) as the target reference frame.
[0077] In the aforementioned method, the reconstruction quality of each frame fluctuates within the reference information module. If the selected target reference frame is of poor quality or has poor correlation, errors may accumulate when the number of frames to be encoded and decoded is long, affecting the encoding and decoding performance of subsequent frames. To address this, this application employs a preset frame selection module to adaptively select target reference frames based on the reference information, thereby improving image encoding and decoding performance and accuracy.
[0078] In some implementations, the frame selection information bitstream is obtained by the encoder using a preset frame selection method, based on a preset number of target reference frames and their frame indices determined from a plurality of candidate reference frames. The preset frame selection method includes any of the following: a quality evaluation method, a cost evaluation method, or a joint learning method. This allows for the comprehensive selection of target reference frames that are more relevant to the current image frame and have better quality, thereby improving image encoding and decoding performance and accuracy.
[0079] This application proposes three frame selection mechanisms, as follows:
[0080] (1) Quality assessment method: The quality assessment method can be used to select the target reference frame based on the ranking of the quality assessment values of each candidate reference frame.
[0081] The preset number of target reference frames is K, and the number of candidate reference frames is P, where K and P are positive integers. Quality evaluation can be performed on each candidate reference frame among the P frames to obtain a quality evaluation value for each candidate reference frame. Quality evaluation includes PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity), Mean Square Error (MSE), Mean Absolute Error (MAE), or Signal-to-Noise Ratio (SNR), etc. It is understood that other quality evaluation methods can also be used in this application, and this application does not limit the methods used for quality evaluation.
[0082] Then, the candidate reference frames of the P-frame are sorted according to their quality assessment values. The candidate reference frames ranked first by a preset number of frames are selected, that is, the K candidate reference frames with the best quality are selected and determined as the target reference frames. In some implementations, after the encoding end selects the target reference frame using a preset frame selection module, it can record the frame index of the target reference frame, encode the frame index into a frame selection information bitstream, and transmit it to the decoding end, thereby decoding to obtain the frame index of the target reference frame.
[0083] Please see Figure 9For example, the preset number of target reference frames is K=3, and the number of candidate reference frames is P=6. The frame IDs of each candidate reference frame are [1,2,3,4,5,6], with the nearest frame ID being 1, recorded sequentially, and the furthest frame ID being 6. Quality evaluation is performed on each candidate reference frame, and the PSNR is calculated, with quality evaluation values of [38.6dB, 38.0dB, 37.8dB, 38.9dB, 37.7dB, 38.4dB]. The six quality evaluation values are then sorted from largest to smallest. The three with the best PSNR can be selected as target reference frames, with frame IDs of 4, 1, and 6.
[0084] (2) Cost evaluation method: The cost evaluation method can be used to select candidate groups whose cost evaluation values meet the preset cost conditions as target reference frames. The cost evaluation value is obtained by evaluating the cost of each candidate group and the current image frame respectively. Each candidate group is obtained by grouping several candidate reference frames according to the preset grouping method. The number of candidate groups is the preset number of groups. Each candidate group contains a preset number of candidate reference frames.
[0085] The preset number of target reference frames is K, and the number of candidate reference frames is P, where K and P are positive integers. For example, P reconstructed frames from neighboring frames closest to the current image frame can be selected as candidate reference frames. It is understood that the selected candidate reference frames are not limited to decoded frames close to the current frame; they can be any frame, unidirectional or multidirectional image frames, etc. This application does not restrict the selection of candidate reference frames. During the encoding process, candidate reference frames can be uncoded candidate reference frames.
[0086] Obtain the preset number of groups Z, where Z is a positive integer. Group the candidate reference frames according to the preset grouping method to obtain the preset number of candidate groups; each candidate group contains the preset number of candidate reference frames. For example, for P candidate reference frames, select Z candidate groups, with each candidate group containing K candidate reference frames. The candidate reference frames contained in each candidate group do not have to be completely identical. In some application scenarios, candidate groups within a preset close-range range may contain identical candidate reference frames.
[0087] Please see Figure 10 For example, the preset grouping method includes: according to the order of each candidate reference frame, taking K candidate reference frames as candidate groups at preset interval frames A (1<=A<=P). The number of overlapping frames in each candidate group is (KA). In some application scenarios, when the number of groups does not meet Z, candidate groups within a preset proximity range (such as relatively close, closest) to the current image frame can contain the same candidate reference frames.
[0088] Then, cost evaluation is performed on each candidate group and the current image frame separately to obtain the cost evaluation value for each candidate group. For example, the rate-distortion loss of the candidate reference frame and the current image frame for each candidate group can be calculated separately, and the rate-distortion loss can be used as the cost evaluation value.
[0089] Finally, candidate groups whose cost evaluation values meet preset cost conditions are selected and determined as target reference frames. The preset cost condition can be either the minimum cost evaluation value or a ranking of cost evaluation values from smallest to largest, with the highest-ranked value. In this way, candidate group Z with the minimum rate-distortion loss can be selected as the target reference frames. The frame indices of candidate group Z can be encoded to obtain the frame selection information bitstream, which is then fed into the current frame bitstream and transmitted to the decoding end for use. The decoding end decodes the frame selection information bitstream to obtain candidate group Z and determine K target reference frames for subsequent decoding operations.
[0090] (3) Joint learning method: The joint learning method can be used to select candidate reference frames that meet the nearest neighbor condition from several candidate reference frames as the filtered candidate reference frames; the filtered candidate reference frames are processed by a preset gating network to obtain the target reference frame.
[0091] The joint learning approach employs a pre-defined gating network to filter, fuse, and compensate candidate reference frames (such as a series of cached reference information, cached information, etc.), adaptively selecting candidate reference frames beneficial to subsequent encoding and decoding performance. The fusion process can be weighted fusion to combine multiple frames into one or more target reference frames. For example, multiple candidate reference frames can be weighted and fused into one or more target reference frames. The pre-defined gating network of the pre-selection frame module can be pre-trained, and it can be jointly trained with other modules to obtain a trained pre-defined gating network that can adaptively select useful information. The pre-defined gating network can be a neural network, such as an attention network or a 3D convolutional network; this application does not limit the pre-defined gating network.
[0092] In some implementations, during training, a preset number of candidate reference frames that meet the nearest neighbor condition can be selected from a plurality of candidate reference frames as filtered candidate reference frames. The preset nearest neighbor condition can be candidate reference frames that are close to or adjacent to the current image frame. The current image frame and each filtered candidate reference frame can be stitched together to obtain a stitched reference frame. Then, the stitched reference frame is input into a preset gating network for processing. Each candidate reference frame may have different weights; through training, the preset gating network can automatically learn the weight distribution of each candidate reference frame.
[0093] In some implementations, a preset number of candidate reference frames that meet the nearest neighbor condition can be selected from a plurality of candidate reference frames as filtered candidate reference frames. The preset nearest neighbor condition can be candidate reference frames that are close to or adjacent to the current image frame. Then, a preset gating network is used to process the filtered candidate reference frames to obtain the target reference frame. For example, the weights of each candidate reference frame can be obtained, and the candidate reference frame with the highest weight (a preset number of frames) can be selected as the target reference frame. Alternatively, the target reference frame can be obtained by fusing the candidate reference frames according to their weights. This method eliminates the need to encode the frame selection information bitstream. Both the decoding and encoding ends can process the filtered candidate reference frames through the preset gating network during their respective encoding / decoding processes to obtain the target reference frame.
[0094] Through the above method, this application adopts a preset frame selection module, which can adaptively select target reference frames based on reference information. The selected target reference frames are more relevant to the current image frame or have better quality, which can improve image encoding and decoding performance and improve the accuracy of image encoding and decoding.
[0095] In some embodiments, it can be combined Figure 3 Step S12 of the above embodiment is further extended. Figure 3 The motion encoding / decoding module can be used to encode / decode motion information. In this embodiment, the motion encoding / decoding module can be used to encode / decode motion information based on a preset number of target reference frames. A preset number of L target reference frames are obtained. The encoding end uses a motion estimation module to perform motion estimation on the current image frame and each of the L target reference frames, resulting in L pieces of motion information, with each target reference frame corresponding to one piece of motion information.
[0096] When the preset frame number L of the target reference frame is 1, there is only one motion information, and subsequent steps can be performed directly, such as motion information encoding / decoding, motion compensation, context encoding / decoding, or frame reconstruction.
[0097] When the preset number of frames L of the target reference frame is greater than 1, there are a preset number of motion information L, and the motion encoding / decoding module needs to encode / decode the L motion information.
[0098] In some implementations, when the preset frame number L is greater than 1, the encoding end performs motion estimation on the current image frame and the target reference frame of the preset frame number to obtain a preset number of motion information. Then, the encoding end can encode the preset number L of motion information using a preset encoding method to obtain a motion bitstream. That is, the motion bitstream obtained at the decoding end is obtained by the encoding end encoding the preset number of motion information using a preset encoding method, where the preset frame number is greater than 1. In this case, the decoding end can correspondingly use a preset decoding method to decode the motion bitstream to obtain the preset number L of decoded motion information.
[0099] For a target reference frame with a preset frame number L greater than 1, the process of encoding / decoding L motion information is illustrated in the following embodiments.
[0100] In some embodiments, please refer to Figure 11 , Figure 11 This is a flowchart illustrating a third embodiment of the image decoding method of this application. The specific steps of this embodiment can be executed using the decoding end described above. The method may include any of the following steps:
[0101] S31: In response to the fact that the motion bitstream is obtained by encoding a preset number of motion information separately at the encoding end, and decoding the preset number of motion bitstreams separately to obtain the preset number of motion information.
[0102] The preset encoding / decoding method can include any of the following: individual encoding / decoding method, merged encoding / decoding method, or reference encoding / decoding method. In other words, the preset encoding method can include any of the following: individual encoding method, merged encoding method, or reference encoding method. Correspondingly, the preset decoding method can include any of the following: individual decoding method, split decoding method, or reference decoding method.
[0103] When the encoding / decoding process of motion information uses a separate encoding / decoding method, a one-to-one multi-stream approach can be achieved. Each piece of motion information can be encoded / decoded independently, with each piece of motion information corresponding to a motion stream. Decoding each motion stream yields the corresponding decoded motion information. The decoding end can respond to the fact that the motion streams are obtained by the encoding end from individually encoding a preset number of pieces of motion information, and then individually decode the preset number of motion streams to obtain the preset number of decoded motion information.
[0104] Combination Figure 3 The motion encoding / decoding module can include an encoding network and a decoding network. The encoding network is used to encode motion information to obtain a motion stream. After the motion stream is transmitted to the decoding end, the decoding end can use the decoding network to decode the motion stream to obtain the decoded motion information.
[0105] Please see Figure 12 At the encoding end, encoding networks (encoding network 1, encoding network 2, encoding network 3, ..., encoding network L) can be used to encode a preset number of L motion information items individually, resulting in a preset number of L motion bitstreams. At the decoding end, decoding networks (decoding network 1, decoding network 2, decoding network 3, ..., decoding network L) can be used to decode a preset number of motion bitstreams individually, resulting in a preset number of L decoded motion information items. The encoding and decoding networks may or may not share network parameters.
[0106] S32: In response to the fact that the motion bitstream is obtained by encoding the preset number of motion information into a merged motion information, the motion bitstream is decoded to obtain the merged motion information, the merged motion information is split to obtain the preset number of motion information, or the merged motion information is used as motion information.
[0107] When the motion information encoding / decoding process uses a merged encoding / decoding method, a many-to-one single bitstream can be represented. The encoding end can merge all motion information into a single motion bitstream, and the decoding end can recover a single motion information after decoding this motion bitstream. This single motion information can be used as the decoded motion information, or, as needed, the single motion information can be split into L decoded motion information pieces.
[0108] Please see Figure 13 At the encoding end, a preset number L motion information items can be merged to obtain merged motion information, i.e., merged into one motion information item. Then, an encoding network is used to encode the merged motion information to obtain a single motion bitstream. At the decoding end, a decoding network is used to decode this single motion bitstream to obtain a single decoded motion information item, i.e., the decoded merged motion information. This decoded merged motion information can be directly used as the motion information for decoding a single motion bitstream. Alternatively, the merged motion information can be split to obtain a preset number L decoded motion information items. The splitting process of the merged motion information is optional and can be determined based on the specific application scenario; this application does not impose any restrictions on this.
[0109] S33: In response to the motion bitstream being obtained by the encoder encoding a preset number of motion information by referring to the corresponding encoding reference motion information, and by referring to the corresponding decoding reference motion information, decoding a preset number of motion bitstreams to obtain a preset number of motion information.
[0110] The encoded reference motion information is at least one of a preset number of motion information, and the decoded reference motion information includes at least one of the decoded motion information corresponding to the preset number of motion bitstreams.
[0111] When a reference encoding / decoding method is used in the motion information encoding / decoding process, a many-to-one multi-stream approach can be represented. Each piece of motion information can be encoded and decoded separately. During the encoding of each piece of motion information, it can be encoded with reference to corresponding encoding reference motion information (at least one of the other motion information) to obtain the corresponding motion stream; each piece of motion information can correspond to one motion stream. During the decoding of each motion stream, it can be decoded with reference to decoding reference motion information (the other decoded motion information) to obtain the corresponding decoded motion information.
[0112] Please see Figure 14 At the encoding end, during the encoding of a preset number L motion information items, encoding networks (encoding network 1, encoding network 2, ..., encoding network L) can be used to encode the preset number L motion information items respectively. During the encoding of the current motion information, other motion information can be referenced for encoding, resulting in a motion bitstream of the current motion information. Encoding each motion information item in this way yields a preset number L motion bitstreams. The identifier or frame index of the corresponding encoded reference motion information for each motion information item can be recorded and encoded into the motion bitstream, allowing the decoding end to determine the corresponding decoding reference motion information based on the identifier or frame index.
[0113] At the decoding end, during the encoding of a preset number L motion streams, decoding networks (decoding network 1, decoding network 2, ..., decoding network L) can be used to decode the preset number L motion streams respectively. During the decoding of the current motion stream, the current motion stream can be encoded with reference to other previously decoded motion information to obtain the decoded motion information of the current motion stream. Decoding each motion stream in this way yields the preset number L decoded motion information. The encoding and decoding networks may or may not share network parameters.
[0114] The encoding reference motion information for each motion information can be determined according to the reference encoding method. For example, it can be motion information of at least one frame in one or two directions. For example, the encoding reference motion information can include backward multi-frame motion information, forward multi-frame motion information, forward and backward multi-frame motion information, etc. The number of frames and the reference motion information corresponding to the encoding reference motion information can be different in the encoding process of each motion information. It is understood that when the motion information encoding / decoding process adopts the reference encoding / decoding method, the encoding and decoding process of motion information can be specifically referred to the process of encoding and decoding the current image frame with reference to other image frames in the image encoding and decoding process. This application does not limit this.
[0115] The above scheme encodes the motion information corresponding to a preset number of target reference frames into a single or multiple motion streams by using a preset decoding method, and decodes the motion streams by using the preset decoding method. This allows the encoding / decoding process to adapt to the motion information of the preset number of target reference frames, and subsequent processes can use the motion information of the preset number of frames, thereby improving the accuracy and flexibility of image encoding and decoding.
[0116] Combination Figure 3 After the motion encoding / decoding module decodes the motion information, it can use the motion compensation module to perform motion compensation on a preset number of target reference frames based on the decoded motion information to obtain target prediction information. When the preset number of target reference frames L = 1, there is only one motion information, and subsequent steps can be performed directly, that is, motion compensation, context encoding / decoding, or frame reconstruction can be performed directly using the decoded motion information. When the preset number of target reference frames L is greater than 1, there are a preset number L motion information, and motion compensation may need to be performed on L motion information.
[0117] In some embodiments, the encoding / decoding system of this application may further include a feature extraction module, which is optional. The feature extraction module includes feature extraction 1 and feature extraction 2. Feature extraction 1 is used to extract features from the current image frame during encoding to obtain first feature information of the current image frame. Feature extraction 2 is used to extract features from a target reference frame to obtain second feature information of the target reference frame. The second feature information can then be used for subsequent motion compensation to obtain target prediction information, and to obtain context information based on the first feature information and the target prediction features. During decoding, feature extraction 2 can be used to extract features from the target reference frame to obtain second feature information of the target reference frame. The second feature information can then be used for subsequent motion compensation to obtain target prediction information.
[0118] In some implementations, when the preset frame number L is greater than 1, motion compensation can be performed on the target reference frames of the preset frame number using a preset compensation method based on the decoded motion information to obtain target prediction information. The motion information is obtained by the encoding end through motion estimation of the current image frame to be encoded and the target reference frames of the preset frame number, respectively.
[0119] In some embodiments, please refer to Figure 15 , Figure 15 This is a flowchart illustrating the fourth embodiment of the image decoding method of this application. The specific steps of this embodiment can be executed using the decoding end described above. The method may include any of the following steps:
[0120] S41: Based on motion information, perform motion compensation on a preset number of target reference frames to obtain a preset number of target prediction information.
[0121] When there are multiple preset frames, motion compensation can be performed on the target reference frames of the preset number of frames based on motion information using a preset compensation method to obtain target prediction information. The preset compensation method includes any of the following: independent compensation method and merged compensation method.
[0122] When the motion compensation module adopts an independent compensation method, each target reference frame corresponds to a motion information. Each target reference frame is subjected to motion compensation based on its corresponding motion information, and the target prediction information corresponding to each target reference frame can be obtained.
[0123] Optionally, motion compensation can be performed after dimensional transformation of the motion information.
[0124] In some implementations, please refer to Figure 16 For L target reference frames, each corresponds to L motion information. A first-dimensional transformation can be performed on a preset number of L motion information to obtain the transformed preset number of L motion information. This first-dimensional transformation can be performed using a convolutional network; this application does not impose restrictions on the first-dimensional transformation. Then, based on the transformed preset number of L motion information, motion compensation is performed on each preset number of L target reference frames to obtain the preset number of L target prediction information. A motion compensation network can be used to perform motion compensation on each target reference frame, and the network parameters of each motion compensation network may or may not be shared.
[0125] S42: Perform fusion processing on the target reference frames of a preset number of frames to obtain fused reference frames. Based on motion information, perform motion compensation on the fused reference frames to obtain prediction information for a single target.
[0126] When the motion compensation module uses a merged compensation method, a preset number of target reference frames can be fused to obtain a fused reference frame. In other words, a preset number of target reference frames can be merged into a single target reference frame. Then, based on the motion information, motion compensation is performed on the fused reference frame to obtain the prediction information for a single target.
[0127] In some implementations, the number of motion information is a preset number, which can be L motion information, where L is greater than 1 or equal to 1. For example, in the process of encoding and decoding motion information using a merged encoding / decoding method, merged motion information can be used as the decoded motion information. In this case, the decoded motion information can be a single piece.
[0128] Optionally, motion compensation can be performed after dimensional transformation of the motion information.
[0129] In some implementations, please refer to Figure 17 For L target reference frames, each corresponds to L motion information. A second-dimensional transformation is performed on a preset number of L motion information to obtain a transformed single motion information. This second-dimensional transformation can be performed using a convolutional network; this application does not impose any restrictions on the second-dimensional transformation. The preset number of L target reference frames are then fused to obtain a fused reference frame, which is also a single target reference frame. Finally, based on the transformed single motion information, motion compensation is performed on the fused reference frame to obtain single target prediction information.
[0130] When using merged motion information as the decoded motion information, a preset number L target reference frames can be fused to obtain a fused reference frame. Then, based on the merged motion information, motion compensation is performed on the fused reference frame to obtain the prediction information for a single target.
[0131] The above scheme can adapt to the target reference frames of a preset number of frames for motion compensation, which can improve the flexibility and accuracy of motion compensation and improve the accuracy of obtaining target prediction information.
[0132] Combination Figure 3 After the motion compensation module acquires the target prediction information, a prediction fusion module can be used to fuse the target prediction information to obtain fused target prediction information, which is then input into subsequent modules. Alternatively, the target prediction information can be directly input into subsequent modules, for example, by using a context decoding module to decode the context bitstream, so that the current reconstructed frame can be obtained based on the decoded context information and the target prediction information.
[0133] In some implementations, the number of target prediction information items is a preset number, which can be a single item or a number greater than 1. In the prediction fusion module, a fusion network can be used to fuse the target prediction information to obtain fused target prediction information. In this way, the target prediction information can be fused, reduced in dimensionality, or optimized.
[0134] In some embodiments, please refer to Figure 18 , Figure 18 This is a flowchart illustrating the fifth embodiment of the image decoding method of this application. The specific steps of this embodiment can be executed using the decoding end described above. The method may include the following steps:
[0135] S51: Obtain reference auxiliary information.
[0136] Reference auxiliary information can be introduced for fusion processing to obtain reference auxiliary information, which includes at least one of the following: decoded reconstructed frames, reconstructed features, decoded intermediate features, target reference frames, prediction information after motion compensation, and other cached reference information. This application does not impose any restrictions on the reference auxiliary information.
[0137] S52: A fusion network is used to fuse the reference auxiliary information and the target prediction information to obtain fused target prediction information.
[0138] The fusion network may include at least one of a residual network, a recurrent network, or an attention network. For example, the fusion network may include a convolutional network and a residual network; the convolutional network can be used for channel dimensionality reduction, and the residual network can be used to obtain residuals, i.e., difference information. It is understood that the fusion network of this application may also be other network structures, and this application does not limit the fusion network.
[0139] Then, the reference auxiliary information and target prediction information are input into the fusion network. The fusion network is used to fuse the reference auxiliary information and target prediction information to obtain fused target prediction information, which is a better overall target prediction information, so that the context encoding / decoding process can be further performed.
[0140] In some implementations, the number of target prediction information items is a preset number L. When L is greater than 1, the L target prediction information items generated from the L target reference frames can be fused and dimensionality reduced to obtain a single, overall fused target prediction information. When L equals 1, subsequent steps can be executed directly, or the fusion process can be performed; this prediction fusion module is optional. If a prediction fusion module is used to fuse individual target prediction information, the addition of reference auxiliary information can optimize the target prediction information, achieving feature tuning.
[0141] For example, please refer to Figure 19 The input contains a preset number L target prediction information, for example, L=5. The fusion network can include convolutional networks and residual networks. The fusion network fuses the preset number L target prediction information and reference auxiliary information to obtain a single fused target prediction information.
[0142] The above scheme can incorporate reference auxiliary information into the target prediction information for fusion processing, and can perform fusion, dimensionality reduction or information optimization on the target prediction information to further improve the accuracy of the target prediction information.
[0143] Combination Figure 3After the motion compensation module obtains the target prediction information, or after the prediction fusion module fuses the target prediction information, a subsequent context encoding / decoding process can be performed. In this embodiment, the context decoding module performs context decoding.
[0144] In some embodiments, please refer to Figure 20 , Figure 20 This is a flowchart illustrating the sixth embodiment of the image decoding method of this application. The specific steps of this embodiment can be executed using the decoding end described above. The method may include the following steps:
[0145] S61: Get the context stream.
[0146] The context stream is obtained by encoding context information at the encoder. The encoder transmits the context stream to the decoder so that the decoder can obtain the context stream.
[0147] S62: Decode the context stream to obtain context information.
[0148] A context decoding module (context decoder) is used to decode the context bitstream to obtain context information. Specifically, the low-dimensional context features obtained from decoding the context bitstream can be upscaled and reconstructed to obtain the context information.
[0149] S63: Using contextual information and target prediction information, obtain the current reconstructed frame of the current image frame.
[0150] After acquiring the context information, the context decoding module uses the context information in conjunction with the target prediction information to obtain a preliminary current reconstructed frame of the current image frame. For example, if the context information acquired at the encoding end is represented as a difference, that is, the difference between the current image frame and the target prediction information, then the context decoding module at the decoding end adds the target prediction information to the context information to obtain a preliminary current reconstructed frame.
[0151] Then, the preliminary current reconstructed frame is further processed to obtain information such as the current reconstructed frame and reconstruction features of the current image frame. The current reconstructed frame and reconstruction features can be saved in the reference information so that it can be used as the target reference frame for the next frame.
[0152] The above scheme, by obtaining more accurate target prediction information, can improve the accuracy of obtaining the current reconstructed frame of the current image frame and enhance image encoding and decoding performance by using the context information and target prediction information after decoding the upper and lower bitstreams to obtain the context information and target prediction information.
[0153] In conjunction with the above embodiments, this application also provides an image encoding method. The specific steps of this image encoding method can be executed using the encoding terminal described above.
[0154] Please see Figure 21 , Figure 21 This is a schematic flowchart of an embodiment of the image encoding method of this application. The method may include the following steps:
[0155] S71: Obtain the current image frame and determine the target reference frame of the preset frame number.
[0156] The current image frame is the image frame to be encoded.
[0157] In some implementations, combined Figure 3 A reference information module can be used to obtain several candidate reference frames. These candidate reference frames include multiple candidate reference frames located in a preset direction of the current image frame. The preset direction can be unidirectional or bidirectional. Each candidate reference frame includes at least one of the following: uncoded image frames, image features, or other reference information.
[0158] Then, the preset frame selection module uses a preset frame selection method to select a preset number of candidate reference frames from several candidate reference frames to obtain the preset number of target reference frames. After determining the preset number of target reference frames, the preset frame selection information is determined based on the preset number of target reference frames and the frame index; wherein, the preset frame selection information includes at least one of the preset number of target reference frames and the frame index of target reference frames; the preset frame selection information is encoded to obtain the frame selection information bitstream.
[0159] In some implementations, the preset frame selection module uses a preset frame selection method that includes any of the following: quality assessment method, cost assessment method, or joint learning method.
[0160] In some implementations, the preset frame selection method includes a quality assessment method. A quality assessment is performed on several candidate reference frames to obtain a quality assessment value for each candidate reference frame; the candidate reference frames are then sorted according to their quality assessment values, and the candidate reference frame ranked at the top preset frame number is selected as the target reference frame.
[0161] In some implementations, the preset frame selection method includes a cost evaluation method. A number of candidate reference frames are grouped according to a preset grouping method to obtain a preset number of candidate groups; each candidate group contains a preset number of candidate reference frames; cost evaluation is performed on each candidate group and the current image frame to obtain a cost evaluation value for each candidate group; candidate groups whose cost evaluation values meet preset cost conditions are selected and determined as target reference frames.
[0162] In some implementations, the preset frame selection method includes a joint learning method. From a number of candidate reference frames, a preset number of candidate reference frames that meet the nearest neighbor condition are selected as the filtered candidate reference frames; a preset gating network is used to process the filtered candidate reference frames to obtain the target reference frame.
[0163] S72: Perform motion estimation on the current image frame and the target reference frame of a preset number of frames to obtain motion information.
[0164] In this step, a motion estimation module can be used to perform motion estimation on the current image frame and a preset number of target reference frames to obtain motion information. The number of motion information items is a preset number; these preset number of motion information items are obtained by performing motion estimation on the current image frame to be encoded and the preset number of target reference frames respectively.
[0165] After acquiring motion information, the motion estimation module can input it into the motion encoding / decoding module to encode a preset number of motion information items, resulting in a motion stream. Specifically, when there are multiple preset frames and preset quantities, a preset encoding method is used to encode the preset number of motion information items to obtain the motion stream. In some application scenarios, the motion encoding / decoding module includes an encoding network, which is used to encode the preset number of motion information items to obtain the motion stream.
[0166] In some implementations, the preset encoding method includes any of the following:
[0167] (1) Encode a preset number of motion information separately to obtain a preset number of motion bitstreams.
[0168] (2) Merge a preset number of motion information to obtain merged motion information, encode the merged motion information to obtain a single motion stream.
[0169] (3) Referencing the corresponding coded reference motion information, encode a preset number of motion information pieces to obtain multiple motion streams. The coded reference motion information is at least one of the preset number of motion information pieces. That is, the coded reference motion information can be at least one of other motion information pieces.
[0170] S73: Based on motion information, perform motion compensation on a preset number of target reference frames to obtain target prediction information. The motion information and target prediction information are used to encode the current frame bitstream of the current image frame.
[0171] A motion compensation module can be used to perform motion compensation on a preset number of target reference frames based on motion information to obtain target prediction information. The motion information and target prediction information are used to encode the current frame bitstream of the current image frame.
[0172] The current frame bitstream can include a motion bitstream and a context bitstream. The motion bitstream is obtained by encoding motion information, and the context bitstream is obtained by encoding context information. Context information can be obtained using target prediction information and the current image frame.
[0173] In some implementations, the motion compensation module can perform motion compensation on a preset number of target reference frames based on motion information using a preset compensation method to obtain target prediction information.
[0174] In some implementations, there are multiple preset number of target reference frames, that is, when the preset number of frames is greater than 1, motion compensation can be performed on the preset number of target reference frames based on motion information using a preset compensation method to obtain target prediction information.
[0175] The preset compensation methods include any of the following:
[0176] Individual compensation method: Based on motion information, motion compensation is performed on a preset number of target reference frames to obtain a preset number of target prediction information.
[0177] Merging compensation method: The target reference frames of a preset number of frames are fused to obtain fused reference frames. Based on motion information, motion compensation is performed on the fused reference frames to obtain prediction information for a single target.
[0178] In some implementations, the number of motion information items is a preset number. In the separate compensation method, a first-dimensional transformation can be performed on the preset number of motion information items to obtain the transformed preset number of motion information items; based on the transformed preset number of motion information items, motion compensation is performed on a preset number of target reference frames to obtain the preset number of target prediction information items.
[0179] In some implementations, the number of motion information items is a preset number. In the merging compensation method, a second-dimensional transformation is performed on the preset number of motion information items to obtain the transformed single motion information; a preset number of target reference frames are fused to obtain fused reference frames; based on the transformed single motion information, motion compensation is performed on the fused reference frames to obtain single target prediction information.
[0180] In some implementations, after obtaining the target prediction information, the process includes: using a fusion network to fuse the target prediction information to obtain fused target prediction information; wherein the number of target prediction information is a preset number.
[0181] In some implementations, reference auxiliary information is obtained; a fusion network is used to fuse the reference auxiliary information and the target prediction information to obtain fused target prediction information.
[0182] After obtaining the target prediction information, it can be input into the context encoding module (i.e., context encoder) for context encoding.
[0183] In some implementations, a context encoding module can be used to obtain context information using the current image frame and target prediction information. This context encoding module can acquire the context information between the current image frame and the target prediction information, and then perform dimensionality reduction and compression on this context information. The context information includes differences or concatenation information, etc., to remove inter-frame temporal and intra-frame spatial correlations. Then, an entropy model is used to encode the context information to obtain a context bitstream. The entropy model can obtain the probability of each character appearing in the quantized context information to be encoded, perform arithmetic encoding, and output the context bitstream.
[0184] In some implementations, the target prediction information includes at least one of a target prediction frame and target prediction features. The target prediction features represent the feature information of the target prediction frame. During the process of obtaining context information using the current image frame and the target prediction information, the context information can be obtained from the current image frame and the target prediction frame, or it can be obtained using the first feature information extracted from the current image frame and the target prediction features. This application does not impose any limitations on this.
[0185] To obtain context information using the first feature information extracted from the current image frame and the target prediction features, the encoding / decoding system of this application may further include a feature extraction module. This feature extraction module is optional and includes Feature Extraction 1 and Feature Extraction 2. Feature Extraction 1 is used to extract features from the current image frame during encoding to obtain the first feature information of the current image frame. Feature Extraction 2 is used to extract features from the target reference frame to obtain the second feature information of the target reference frame. The second feature information can then be used for subsequent motion compensation to obtain target prediction information, and to obtain context information based on the first feature information and the target prediction features. During decoding, Feature Extraction 2 can be used to extract features from the target reference frame to obtain the second feature information of the target reference frame. The second feature information can then be used for subsequent motion compensation to obtain target prediction information.
[0186] The specific implementation of this step can be found in the specific implementation process of the encoding end in the above embodiments, and will not be repeated here.
[0187] In some embodiments, for the image encoding and decoding system described above, each module can cooperate or combine to implement any of the above embodiments, and the implementation methods of each module can be combined with each other to form multiple implementation methods. This application does not limit the implementation methods adopted by each module. It is understood that in the above methods of specific embodiments, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0188] In accordance with the above embodiments, this application provides a decoding terminal for implementing the steps of any embodiment of the above image decoding method.
[0189] Please see Figure 22 , Figure 22 This is a schematic diagram of the structure of one embodiment of the decoding end of this application. The decoding end 80 includes a reference frame module 81, a motion decoding module 82, and a motion compensation module 83.
[0190] The reference frame module 81 is used to obtain a target reference frame of a preset number of frames.
[0191] The motion decoding module 82 is used to decode the motion bitstream to obtain motion information; wherein, the motion bitstream is obtained by encoding the motion information at the encoding end, and the motion information is obtained by motion estimation based on the current image frame and a target reference frame of a preset number of frames.
[0192] The motion compensation module 83 is used to perform motion compensation on a preset number of target reference frames based on motion information to obtain target prediction information; wherein, the target prediction information is used to decode and obtain the current reconstructed frame of the current image frame.
[0193] In connection with the above embodiments, this application provides an encoding end for implementing the steps of any embodiment of the above image encoding method.
[0194] Please see Figure 23 , Figure 23 This is a schematic diagram of the structure of an embodiment of the encoding terminal of this application. The encoding terminal 90 includes a reference determination module 91, a motion estimation module 92, and a motion compensation module 93.
[0195] The reference determination module 91 is used to acquire the current image frame and determine the target reference frame of a preset number of frames.
[0196] The motion estimation module 92 is used to perform motion estimation on the current image frame and the target reference frame of a preset number of frames to obtain motion information.
[0197] The motion compensation module 93 performs motion compensation on a preset number of target reference frames based on motion information to obtain target prediction information. The motion information and target prediction information are used to encode the current frame bitstream of the current image frame.
[0198] It should be noted that the decoding end and encoding end provided in the above embodiments belong to the same concept as the image decoding method and image encoding method provided in the corresponding embodiments. The specific way in which each module and unit performs operations has been described in detail in the method embodiments and will not be repeated here. In practical applications, the decoding end and encoding end provided in the above embodiments can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. This application does not impose any limitations on this.
[0199] Regarding the above embodiments, this application provides a computer device; please refer to [link / reference]. Figure 24 , Figure 24 This is a schematic diagram of a computer device according to an embodiment of the present application. The computer device 200 includes a memory 201 and a processor 202, wherein the memory 201 and the processor 202 are coupled to each other. The memory 201 stores program data, and the processor 202 executes the program data to implement the steps of any embodiment of the above-described filtering processing method. The computer device 200 can serve as the encoding end and / or decoding end in the image encoding / decoding system of the above embodiments, executing the steps of any embodiment of the above-described image decoding method and image encoding method.
[0200] In this embodiment, processor 202 can also be referred to as CPU (Central Processing Unit). Processor 202 may be an integrated circuit chip with signal processing capabilities. Processor 202 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. The general-purpose processor can be a microprocessor, or processor 202 can be any conventional processor.
[0201] The methods described in the above embodiments can be implemented as computer programs; therefore, this application proposes a computer-readable storage medium. Please refer to [link to relevant documentation]. Figure 25 , Figure 25This is a schematic diagram of a computer-readable storage medium according to an embodiment of the present application. The computer-readable storage medium 300 stores program data 301 that can be executed by a processor. The program data 301 can be executed by the processor to implement the steps of any of the above-described image decoding method and image encoding method embodiments.
[0202] In this embodiment, the computer-readable storage medium 300 can be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or a medium that can store program data 301. Alternatively, it can be a server that stores the program data 301, which can send the stored program data 301 to other devices for execution, or it can self-run the stored program data 301.
[0203] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments. For the sake of brevity, this application will not repeat the details here.
[0204] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to. For the sake of brevity, the present application will not repeat them here.
[0205] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. An image decoding method, characterized in that, include: Obtain the target reference frame with a preset frame number; The motion stream is decoded to obtain motion information; wherein the motion stream is obtained by encoding motion information at the encoding end, and the motion information is obtained by motion estimation based on the current image frame and the target reference frame of the preset number of frames; Based on the motion information, motion compensation is performed on the target reference frames of the preset number of frames to obtain target prediction information; wherein, the target prediction information is used to decode the current reconstructed frame of the current image frame.
2. The method according to claim 1, characterized in that, The acquisition of the target reference frame with a preset number of frames includes: Obtain decoded reference information, which includes several candidate reference frames; Using frame selection information, a preset number of candidate reference frames are selected from the plurality of candidate reference frames and determined as the target reference frame.
3. The method according to claim 2, characterized in that, The plurality of candidate reference frames include: multiple candidate reference frames located in a preset direction of the current image frame, wherein the preset direction includes unidirectional or bidirectional, and the candidate reference frames include at least one of the following: a decoded reconstructed frame, reconstructed features of the reconstructed frame, decoded intermediate features, and motion-compensated prediction information. And / or, the frame selection related information includes at least one of the decoding performance conditions of the decoding end and preset frame selection information, and the frame selection related information is used at least to determine the preset number of frames; wherein, the preset frame selection information includes at least one of the preset number of target reference frames and the frame index of the target reference frames; the preset frame selection information is obtained by decoding the frame selection information bitstream, and the frame selection information bitstream is obtained by the encoding end using a preset frame selection method, based on the preset number of target reference frames and the frame index determined from the plurality of candidate reference frames.
4. The method according to claim 3, characterized in that, The preset frame selection method includes any one of the following: quality assessment method, cost assessment method, and joint learning method; The quality assessment method is used to select the target reference frame based on the ranking of the quality assessment values of each candidate reference frame. The cost evaluation method is used to: select candidate groups whose cost evaluation values meet preset cost conditions as the target reference frames, wherein the cost evaluation values are obtained by evaluating the cost of each candidate group and the current image frame respectively, each candidate group is obtained by grouping the several candidate reference frames according to a preset grouping method, the number of candidate groups is a preset number of groups, and each candidate group contains the preset number of candidate reference frames. The joint learning method is used to: select a preset number of candidate reference frames that meet the nearest neighbor condition from the plurality of candidate reference frames, and use them as the filtered candidate reference frames; process the filtered candidate reference frames using a preset gating network to obtain the target reference frame.
5. The method according to claim 1, characterized in that, The motion bitstream is obtained by encoding a preset number of motion information items using a preset encoding method at the encoding end. The preset number of motion information items is obtained by the encoding end performing motion estimation based on the current image frame and the preset number of target reference frames. Decoding the motion stream to obtain motion information includes: The motion stream is decoded using a preset decoding method to obtain the preset number of motion information items.
6. The method according to claim 5, characterized in that, Decoding the motion stream using a preset decoding method to obtain the preset number of motion information items includes any of the following steps: In response to the fact that the motion bitstream is obtained by the encoding end by encoding a preset number of motion information separately, the preset number of motion bitstreams are decoded separately to obtain the preset number of motion information; In response to the fact that the motion stream is obtained by encoding the preset number of motion information into a merged motion information by the encoding end, the motion stream is decoded to obtain the merged motion information, the merged motion information is split to obtain the preset number of motion information, or the merged motion information is used as the motion information; In response to the fact that the motion bitstream is obtained by the encoding end encoding a preset number of motion information with reference to corresponding encoding reference motion information, and decoding the preset number of motion bitstreams with reference to corresponding decoding reference motion information, the preset number of motion information is obtained; wherein, the encoding reference motion information includes at least one of the preset number of motion information, and the decoding reference motion information includes at least one of the decoded motion information referenced by the preset number of motion bitstreams.
7. The method according to claim 1, characterized in that, The step of performing motion compensation on the target reference frames of the preset number of frames based on the motion information to obtain target prediction information includes: Based on the motion information, motion compensation is performed on the target reference frames of the preset number of frames using a preset compensation method to obtain the target prediction information.
8. The method according to claim 7, characterized in that, The preset number of frames is multiple; the step of performing motion compensation on the target reference frames of the preset number of frames based on the motion information using a preset compensation method to obtain the target prediction information includes any of the following steps: Based on the motion information, motion compensation is performed on the preset number of target reference frames to obtain a preset number of target prediction information; The target reference frames of the preset number of frames are fused to obtain fused reference frames. Based on the motion information, motion compensation is performed on the fused reference frames to obtain single target prediction information.
9. The method according to claim 8, characterized in that, The number of motion information items is a preset number; The step of performing motion compensation on the preset number of target reference frames based on the motion information to obtain a preset number of target prediction information includes: The preset number of motion information pieces are transformed in the first dimension to obtain the preset number of transformed motion information pieces. Based on the transformed preset number of motion information, motion compensation is performed on the preset number of target reference frames to obtain the preset number of target prediction information; And / or, the step of fusing the target reference frames of the preset number of frames to obtain fused reference frames, and performing motion compensation on the fused reference frames based on the motion information to obtain single target prediction information includes: The preset number of motion information pieces are transformed in a second dimension to obtain the transformed single motion information. The target reference frames of the preset number of frames are fused to obtain the fused reference frames; Based on the transformed single motion information, motion compensation is performed on the fused reference frame to obtain the single target prediction information.
10. The method according to claim 7, characterized in that, After obtaining the target prediction information by performing motion compensation on the target reference frames of the preset number of frames based on the motion information using a preset compensation method, the process includes: The target prediction information is fused using a fusion network to obtain fused target prediction information.
11. The method according to claim 10, characterized in that, The step of fusing the target prediction information using a fusion network to obtain fused target prediction information includes: Obtain reference auxiliary information; wherein the reference auxiliary information includes at least one of the following: decoded reconstructed frame, reconstructed features, decoded intermediate features, target reference frame, and motion-compensated prediction information; The reference auxiliary information and the target prediction information are fused using the fusion network to obtain the fused target prediction information.
12. The method according to claim 1, characterized in that, After performing motion compensation on the target reference frames of the preset number of frames based on the motion information to obtain target prediction information, the process includes: Obtain the context stream; The context stream is decoded to obtain context information; Using the context information and the target prediction information, the current reconstructed frame of the current image frame is obtained.
13. An image encoding method, characterized in that, include: Obtain the current image frame, and determine the target reference frame with a preset number of frames; Motion estimation is performed on the current image frame and the target reference frame of the preset number of frames to obtain motion information; Based on the motion information, motion compensation is performed on the target reference frames of the preset number of frames to obtain target prediction information, wherein the motion information and the target prediction information are used to encode the current frame bitstream of the current image frame.
14. The method according to claim 13, characterized in that, The target reference frame for determining the preset number of frames includes: Acquire several candidate reference frames; Using a preset frame selection method, a preset number of candidate reference frames are selected from the plurality of candidate reference frames to obtain the preset number of target reference frames; And / or, after determining the target reference frame of the preset number of frames, the process includes: Based on the preset number of frames and the frame index of the target reference frame, preset frame selection information is determined; wherein, the preset frame selection information includes at least one of the preset number of frames of the target reference frame and the frame index of the target reference frame; The preset frame selection information is encoded to obtain the frame selection information bitstream.
15. The method according to claim 14, characterized in that, The preset frame selection method includes: a quality assessment method; the step of selecting a preset number of candidate reference frames from the plurality of candidate reference frames using the preset frame selection method to obtain the preset number of target reference frames includes: The quality of the candidate reference frames is evaluated to obtain the quality evaluation value of each candidate reference frame. The candidate reference frames are sorted according to the quality assessment value, and the candidate reference frame with the preset number of frames in the order is selected as the target reference frame.
16. The method according to claim 14, characterized in that, The preset frame selection method includes: a cost evaluation method; the step of selecting a preset number of candidate reference frames from the plurality of candidate reference frames using the preset frame selection method to obtain the preset number of target reference frames includes: The candidate reference frames are grouped according to a preset grouping method to obtain a preset number of candidate groups; wherein each candidate group contains the preset number of candidate reference frames. Cost evaluation is performed on each candidate group and the current image frame respectively to obtain the cost evaluation value of each candidate group; Candidate groups whose cost evaluation values meet preset cost conditions are selected and determined as the target reference frames.
17. The method according to claim 14, characterized in that, The preset frame selection method includes: a joint learning method; the step of selecting a preset number of candidate reference frames from the plurality of candidate reference frames using the preset frame selection method to obtain the target reference frame includes: From the plurality of candidate reference frames, select the preset number of candidate reference frames that meet the nearest neighbor condition, and use them as the filtered candidate reference frames; The filtered candidate reference frames are processed using a preset gating network to obtain the target reference frame.
18. The method according to claim 13, characterized in that, The number of motion information items is a preset number; after performing motion estimation on the current image frame and the preset number of target reference frames to obtain motion information, the process includes: The preset number of motion information items are encoded to obtain a motion bitstream; In cases where there are multiple preset frames, a preset encoding method is used to encode the preset number of motion information to obtain a motion bitstream.
19. The method according to claim 17, characterized in that, The step of encoding the preset number of motion information items using a preset encoding method to obtain a motion bitstream includes any of the following steps: Each of the preset number of motion information items is individually encoded to obtain a preset number of motion bitstreams; The preset number of motion information items are merged to obtain merged motion information, and the merged motion information is encoded to obtain a single motion stream; By referring to the corresponding encoded reference motion information, the preset number of motion information are encoded to obtain multiple motion streams; wherein, the encoded reference motion information is at least one of the preset number of motion information.
20. The method according to claim 13, characterized in that, The current frame bitstream includes a context bitstream; after performing motion compensation on the target reference frames of the preset number of frames based on the motion information to obtain target prediction information, the process includes: Context information is obtained using the current image frame and the target prediction information; The context information is encoded to obtain the context bitstream.
21. A computer device, characterized in that, The method includes a memory and a processor coupled to each other, the memory storing program data, and the processor executing the program data to implement the steps of the method according to any one of claims 1 to 12, and / or to implement the steps of the method according to any one of claims 13 to 20.
22. A computer-readable storage medium, characterized in that, The system stores program data that can be executed by a processor, the program data being used to implement the steps of the method according to any one of claims 1 to 12, and / or to implement the steps of the method according to any one of claims 13 to 20.