End-to-end video coding and decoding method, system and device and storage medium

By introducing low-resolution video encoder and iterative enhancement methods in end-to-end video encoding and decoding technology, the problems of large motion information and emerging object descriptions are solved, and higher encoding compression ratios and performance are achieved.

CN120186360AActive Publication Date: 2025-06-20UNIV OF SCI & TECH OF CHINA
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510453087.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-06-20
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

Existing end-to-end video encoding and decoding technologies are difficult to effectively model large motion information and describe new objects in the current frame, resulting in low encoding compression ratio.

Method used

By introducing a low-resolution video encoder, iterative enhancement is used to use decoded low-resolution reconstruction motion information and low-resolution reconstruction frames to obtain higher quality motion information and context information, and upsample during the encoding and decoding process to ensure that the resolution of the final output context information is the same as the current frame.

Benefits of technology

The characterization ability of motion information and the quality of predictive context are significantly improved, thereby improving the compression ratio and encoding performance of video encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120186360A_ABST
    Figure CN120186360A_ABST
Patent Text Reader

Abstract

The invention discloses an end-to-end video coding and decoding method, system and device and a storage medium, which are corresponding schemes, and in the scheme, additional spatial domain reference is introduced for end-to-end video coding; a new motion information coding and decoding and motion alignment method is designed to solve the problems encountered by the end-to-end video in large motion information representation and description of a newly appearing object; specifically, a low-resolution video encoder is introduced, decoded low-resolution reconstruction motion information and low-resolution reconstruction frames are used for iterative enhancement at a decoding end to obtain motion information with higher quality, and meanwhile, the low-resolution reconstruction frames are continuously enhanced in the iterative enhancement to provide more additional spatial domain reference frames. Based on higher-quality motion information and richer reference frames, the quality of the prediction context is enhanced, so that the coding and decoding performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video coding and decoding, and particularly relates to an end-to-end video coding and decoding method, system, device and storage medium. Background Art

[0002] In the storage and transmission of videos, it is usually necessary to perform encoding operations on videos to reduce storage capacity and transmission bandwidth. The process of video coding and decoding usually includes motion information encoding and decoding, context prediction, and frame encoding and decoding. In the encoding process, first, the motion information is encoded into a binary bitstream. Secondly, the reconstructed motion information is used to perform motion alignment on the reference frame to obtain the prediction context. Finally, the prediction context is used to encode the current frame into a binary bitstream. In the decoding process, first, the binary bitstream of the motion information is decoded into the reconstructed motion information. Secondly, the reconstructed motion information is used to perform motion alignment on the reference frame to obtain the prediction context. Finally, the binary bitstream of the current frame is decoded into the reconstructed frame using the prediction context. Among them, the higher the quality of the prediction context during the prediction process, the higher the compression ratio that can be achieved under the same reconstruction quality.

[0003] In recent years, with the rapid development of deep learning technology and end-to-end image coding technology based on deep learning, the end-to-end video coding technology has also been significantly improved. End-to-end video coding replaces all modules in the video coding system with learnable modules based on neural networks and is optimized based on the rate-distortion loss function. A better module combination and utilization method are the keys to improving the compression ratio of end-to-end video coding. Among them, enhancing the quality of the reconstructed motion information and optimizing the motion alignment technology can significantly improve the quality of the prediction context, thereby obtaining a higher coding compression ratio.

[0004] Prior Art One: A motion information coding and decoding method for enhancing the representation ability of motion information.

[0005] Motion information encoding and decoding will obtain the reconstructed motion information at the decoding end, and the quality of the reconstructed motion information affects the quality of context prediction in motion alignment. Generally speaking, at the encoding end, the reference frame and the current frame x t are used to estimate the motion information m t , and this motion information will be encoded into a binary bitstream by the motion information encoder, and the decoder will decode this binary bitstream to obtain the reconstructed motion information The representation ability and quality of the reconstructed motion information will affect the subsequent motion alignment process. The following introduces the existing motion information coding and decoding methods for enhancing the representation ability of motion information.

[0006] One implementation enhances the representation ability of motion information at the decoding end. The reconstructed motion information is obtained by decoding, and the reconstructed motion information, the reference frame, and the reference features are further fused to generate multiple motion information offset maps The motion information offset map is added to the reconstructed motion information to obtain multiple enhanced motion information. This enhancement method cannot perceive the content of the current frame, and the enhancement effect is limited; the encoded motion information is simple and has a large distortion, and it cannot effectively process large motion information.

[0007] Another implementation method enhances the motion information representation ability during the encoding and decoding process. When estimating the motion information at the encoding end, multiple motion information offset maps are estimated and directly encode the motion information offset map instead of encoding a motion information m t . Multiple reconstructed motion information offset maps are obtained at the decoding end for the subsequent motion alignment process. This method enhances the motion information representation ability through multiple motion information offset maps estimated at the encoding end, can perceive the content of the current frame, and reduces the distortion of motion information encoding. However, this method needs to encode multiple motion information offset maps, which will increase the bitstream of motion information encoding; the motion information offset is limited by the local receptive field of the convolutional network, and it is still difficult to effectively model large motion information, and the representation ability is still limited.

[0008] Articles related to the prior art one are as follows:

[0009] Article 1: Li J, Li B, Lu Y. Neural video compression with diverse contexts[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2023:22616-22626.

[0010] Article 2: Hu Z, Lu G, Xu D. FVC: A new framework towards deep video compression in feature space[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2021:1502-1511.

[0011] Prior art two: A motion alignment method that enhances context prediction.

[0012] After obtaining the reconstructed motion information, effective motion alignment can obtain a high-quality prediction context. Motion alignment aligns the content of the reference frame to the current temporal position through motion. The accuracy of this operation is affected by the quality of the motion information and the quality of the reference frame. The higher the quality and the stronger the representation ability of the motion information and the higher the quality and the more diverse the content of the reference frame can both effectively improve the quality of the prediction context. A higher-quality prediction context directly affects the coding compression ratio of the current frame. The following introduces how to perform enhanced motion alignment context prediction after obtaining the reconstructed motion information.

[0013] In one implementation, multiple-scale features are extracted from the reference frame and the reconstructed motion information is also downsampled into reconstructed motion information of multiple scales After that, multiple motion alignments are performed at the corresponding scales to obtain context predictions at multiple scales Such multi-scale motion alignment enhances the utilization of reference frame information and can, to a certain extent, alleviate the problem of inaccurate reconstructed motion information caused by large motion information at small scales. However, it is difficult to handle occlusion and newly emerging objects only using the prediction context generated from the temporal reference frame.

[0014] In another implementation, motion alignment is performed according to multiple motion information offset maps encoded from the motion information This method extracts the reference frame as features and uses multiple motion information offset maps to perform multiple alignment operations on different channels of the features to obtain the prediction context This prediction context subtracts the feature expression F t of the current frame to obtain the residual, and the residual is encoded and decoded to reduce the coding rate of the current frame. The compression efficiency of residual coding has been proven to be inferior to that of conditional coding methods; although different motion information offset maps are used for motion alignment to make more full use of the reference frame content, its reference content is still limited and it is difficult to handle occlusion and newly emerging objects.

[0015] The articles related to the second prior art are as follows:

[0016] Article 3: Sheng X, Li J, Li B, et al. Temporal context mining for learned video compression[J]. IEEE Transactions on Multimedia, 2022, 25: 7311-7322.

[0017] Article 4: Hu Z, Xu D, Lu G, et al. Fvc: An end-to-end framework towards deep video compression in feature space[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 45(4): 4569-4585. Summary of the Invention

[0018] The object of the present invention is to provide an end-to-end video encoding and decoding method, system, device and storage medium, which can effectively model large motion information and can also effectively describe newly emerging objects in the current frame.

[0019] The object of the present invention is achieved by the following technical solutions:

[0020] An end-to-end video encoding and decoding method, comprising:

[0021] Downsampling process: performing spatial downsampling on the current frame to obtain a video frame called the low-resolution current frame;

[0022] Low-resolution encoding and decoding: encoding and decoding the low-resolution current frame to obtain a low-resolution reconstructed frame and reconstructed motion information;

[0023] Iterative enhanced context prediction: combining the full-resolution temporal reference frame, the low-resolution reconstructed frame and the reconstructed motion information, iteratively enhancing the reconstructed motion information, and thereby enhancing the context information. The finally output context information is used for the full-resolution encoding and decoding of the current frame, and, upsampling is performed during the gap of iterative enhancement so that the resolution of the finally output context information is the same as the resolution of the current frame; wherein, the full-resolution temporal reference frame is the reconstructed frame obtained by full-resolution encoding and decoding when encoding and decoding the previous frame;

[0024] Full-resolution encoding and decoding: combining the context information obtained by iterative enhanced context prediction, encoding and decoding the current frame to obtain a reconstructed frame.

[0025] An end-to-end video encoding and decoding system, comprising:

[0026] A downsampling module, configured to perform spatial downsampling on the current frame to obtain a video frame called the low-resolution current frame;

[0027] A low-resolution encoding and decoding model, configured to encode and decode the low-resolution current frame to obtain a low-resolution reconstructed frame and reconstructed motion information;

[0028] An iterative enhanced context prediction model is used to combine a full-resolution time-domain reference frame, a low-resolution reconstructed frame, and reconstructed motion information, iteratively enhance the reconstructed motion information, and thereby enhance the context information. The finally output context information is used for full-resolution encoding and decoding of the current frame. Moreover, upsampling is performed during the gap of iterative enhancement so that the resolution of the finally output context information is the same as that of the current frame. Among them, the full-resolution time-domain reference frame is a reconstructed frame obtained through full-resolution encoding and decoding when encoding and decoding the previous frame.

[0029] A full-resolution encoding and decoding model is used to combine the context information obtained by iterative enhanced context prediction, perform encoding and decoding on the current frame, and obtain a reconstructed frame.

[0030] A processing device includes: one or more processors; a memory for storing one or more programs;

[0031] Among them, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the foregoing method.

[0032] A readable storage medium stores a computer program, and when the computer program is executed by a processor, the foregoing method is implemented.

[0033] It can be seen from the technical solutions provided by the present invention above that additional spatial domain references are introduced for end-to-end video coding, and a new motion information encoding and decoding and motion alignment method (corresponding to the iterative enhanced context prediction scheme provided by the present invention) is designed to solve the problems encountered by end-to-end video in representing large motion information and describing newly emerging objects. Specifically, a low-resolution video encoder is introduced, and the decoded low-resolution reconstructed motion information and low-resolution reconstructed frame are used for iterative enhancement at the decoding end to obtain higher-quality motion information. At the same time, the low-resolution reconstructed frame will also be continuously enhanced during iterative enhancement to provide higher-quality additional spatial domain reference information. Based on higher-quality motion information and richer reference frames, the quality of the predicted context can be enhanced, and the encoding and decoding performance can be improved. Description of the Drawings

[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0035] Figure 1 It is a flowchart of an end-to-end video encoding and decoding method provided by an embodiment of the present invention;

[0036] Figure 2Schematic diagram of the overall framework of an end-to-end video encoding and decoding method provided by an embodiment of the present invention;

[0037] Figure 3 Schematic diagram of low-resolution video encoding and decoding provided by an embodiment of the present invention;

[0038] Figure 4 Schematic diagram of iterative enhanced context prediction provided by an embodiment of the present invention;

[0039] Figure 5 Schematic diagram of the quality comparison between the reconstructed motion information after iterative enhancement of the present invention and the existing model DCVC-DC provided by an embodiment of the present invention;

[0040] Figure 6 Schematic diagram of the quality comparison between the context information after iterative enhancement of the present invention and the existing model DCVC-DC provided by an embodiment of the present invention;

[0041] Figure 7 Schematic diagram of an end-to-end video encoding and decoding system provided by an embodiment of the present invention;

[0042] Figure 8 Schematic diagram of a processing device provided by an embodiment of the present invention. Detailed implementation manners

[0043] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0044] First, the following explanations are given for the terms that may be used in this article:

[0045] Descriptions with semantic meanings such as "including", "comprising", "containing", "having" or other similar ones should be interpreted as non-exclusive inclusion. For example, including a certain technical feature element (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, processes, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or articles, etc.) should be interpreted as not only including the clearly listed certain technical feature element, but also including other well-known technical feature elements in the art that are not clearly listed.

[0046] The term "consisting of" means excluding any technical feature elements not explicitly listed. If this term is used in a claim, it will make the claim a closed claim, excluding technical feature elements other than those explicitly listed, except for conventional impurities related thereto. If this term only appears in a certain clause of a claim, it only limits the elements explicitly listed in that clause, and the elements recorded in other clauses are not excluded from the overall claim.

[0047] The following provides a detailed description of an end-to-end video encoding and decoding method, system, device, and storage medium provided by the present invention. The content not described in detail in the embodiments of the present invention belongs to the prior art well known to those skilled in the art. For the conditions not specified in the embodiments of the present invention, they are carried out according to the conventional conditions in the art or the conditions recommended by the manufacturer. The reagents or instruments used in the embodiments of the present invention without indicating the manufacturer are all conventional products that can be obtained through commercial purchase.

[0048] Embodiment 1

[0049] In the embodiment of the present invention, an end-to-end video encoding and decoding method is adopted, as Figure 1 shown, mainly including the following steps:

[0050] Step 1: Downsampling processing.

[0051] In this step, spatial downsampling is performed on the current frame, and the obtained video frame is called the low-resolution current frame.

[0052] Step 2: Low-resolution encoding and decoding.

[0053] In this step, the low-resolution current frame is encoded and decoded to obtain a low-resolution reconstructed frame and reconstructed motion information.

[0054] The preferred implementation of this step is as follows: Combining the low-resolution temporal reference frame, the low-resolution current frame is motion-encoded to obtain motion information, and through decoding, reconstructed motion information is obtained wherein the low-resolution temporal reference frame is the low-resolution reconstructed frame obtained when encoding and decoding the previous frame (that is, when performing the encoding and decoding method of the present invention on the previous frame, obtained through Step 2), t is the index of the current frame; after context prediction by combining the motion information and the low-resolution temporal reference frame, the low-resolution current frame is encoded, and the low-resolution reconstructed frame is obtained through decoding

[0055] Step 3: Iterative enhanced context prediction.

[0056] In this step, by combining the full-resolution temporal reference frame, the low-resolution reconstructed frame, and the reconstructed motion information, the reconstructed motion information is iteratively enhanced, and thereby the context information is enhanced. The finally output context information is used for the full-resolution encoding and decoding of the current frame. Moreover, upsampling is performed during the gap between iterative enhancements, such that the resolution of the finally output context information is the same as that of the current frame; wherein, the full-resolution temporal reference frame is the reconstructed frame obtained through full-resolution encoding and decoding when encoding and decoding the previous frame (i.e., obtained through step 4 when performing the encoding and decoding method of the present invention on the previous frame).

[0057] The preferred implementation manner of this step is as follows:

[0058] In the i-th enhancement stage, combine the full-resolution temporal reference frame the enhanced context information obtained in the previous enhancement stage and the enhanced motion information obtained in the previous enhancement stage to perform motion enhancement to obtain the enhanced motion information in the i-th enhancement stage wherein, i is the index of the enhancement stage. When i = 1, the enhanced context information obtained in the previous enhancement stage is the low-resolution reconstructed frame the enhanced motion information obtained in the previous enhancement stage is the reconstructed motion information t is the index of the current frame.

[0059] Utilize the enhanced motion information to perform motion information compensation on the full-resolution temporal reference frame to obtain a temporally aligned reference frame. Use the temporally aligned reference frame to perform context prediction enhancement on the enhanced context information obtained in the previous enhancement stage to obtain the enhanced context information in the i-th enhancement stage

[0060] Repeat N times continuously. The enhanced context information obtained in the N-th enhancement stage is the finally output context information; and, during the gap between iterative enhancements, i.e., between adjacent enhancement stages, perform upsampling on the enhanced context information, and the total upsampling ratio is the same as the spatial downsampling ratio.

[0061] Step 4, Full-resolution encoding and decoding.

[0062] In this step, by combining the context information obtained through iterative enhanced context prediction, perform encoding and decoding on the current frame to obtain a reconstructed frame.

[0063] In the embodiments of the present invention, if it is the first frame, full-resolution encoding and decoding are performed through an image encoder in combination with existing solutions. Starting from the second frame, the solution provided by the present invention is used for encoding and decoding. That is, the range of the current frame described in the present invention is from the second frame to the last frame. The above is only introduced by taking the encoding and decoding process of the current frame as an example, and all frames in the video are processed in the above manner, and finally the encoding and decoding of the entire video are completed.

[0064] Preferably: The full-resolution encoding and decoding are implemented through a full-resolution encoding and decoding model, the low-resolution encoding and decoding are implemented through a low-resolution encoding and decoding model, and the iterative enhanced context prediction is implemented through an iterative enhanced context prediction model. The full-resolution encoding and decoding model, the low-resolution encoding and decoding model, and the iterative enhanced context prediction model are trained in the following manner:

[0065] Construct the following loss function:

[0066]

[0067] where t is the index of the current frame, t = 2, …, T, T is the total number of frames; L1, L2, and L3 are three loss functions, D(·, ·) is the distortion loss, R(.) is the bitrate requirement; x t is the current frame, is the reconstructed frame; is the low-resolution current frame, is the low-resolution reconstructed frame; λ is a parameter for controlling the bitrate, and w is a parameter for regulating the low-resolution reconstruction quality (i.e., the distortion loss of the low-resolution encoding and decoding model) and the full-resolution reconstruction quality (the distortion loss of the full-resolution encoding and decoding model);

[0068] Use the loss function L1 to optimize the low-resolution video encoding and decoding model;

[0069] Use the loss function L2 to optimize the full-resolution encoding and decoding model and the iterative enhanced context prediction model;

[0070] After that, use the loss function L3 to optimize the low-resolution video encoding and decoding model, the full-resolution encoding and decoding model, and the iterative enhanced context prediction model.

[0071] In the above solution provided by the embodiments of the present invention, the downsampling method can be arbitrarily selected, the downsampling ratio of the low resolution can be arbitrarily specified, and the number of iterations N of the iterative enhancement is not limited.

[0072] In order to more clearly show the technical solution provided by the present invention and the technical effects produced, the following uses specific embodiments to describe in detail the method provided by the embodiments of the present invention.

[0073] I. Overall introduction of the solution.

[0074] Considering the existing motion information encoding and decoding and motion alignment methods, which encode a single motion alignment with limited temporal reference frames, on the one hand, it is difficult to effectively model large motion information, and on the other hand, the limited temporal reference frames are difficult to effectively describe newly emerging objects in the current frame.

[0075] The present invention provides an end-to-end video encoding and decoding method to solve the above problems. The present invention introduces additional spatial references (i.e., when encoding full-resolution frames x t a low-resolution reconstructed frame is introduced ), and designs a new motion information encoding and decoding and motion alignment method (i.e., the iterative enhanced context prediction scheme provided above) to solve the problems encountered in the end-to-end video in representing large motion information and describing newly emerging objects. Specifically, the full-resolution motion information encoding and decoding is improved to a low-resolution video codec, and the decoded low-resolution reconstructed motion information and low-resolution reconstructed frames are used for iterative enhancement at the decoding end to obtain higher-quality motion information. At the same time, the low-resolution reconstructed frames will also be continuously enhanced during the iterative enhancement to provide more additional spatial reference frames. Based on the higher-quality motion information and richer reference frames, the quality of the prediction context is enhanced.

[0076] The above method provided by the present invention only improves some links, so it can be applied to any end-to-end video coding system.

[0077] Figure 2 Shows the overall framework of the end-to-end video encoding and decoding method provided by the present invention. The current frame encoding and decoding process can be implemented with reference to conventional techniques, so it will not be elaborated here; the following mainly introduces the two parts of low-resolution encoding and decoding (low-resolution video codec) and iterative enhanced context prediction.

[0078] 1. Low-resolution encoding and decoding.

[0079] As Figure 3 shown, it is a schematic diagram of low-resolution encoding and decoding. This process is similar to a common video encoding / decoding system and relies on temporal reference frames for motion information encoding, context prediction, and current frame encoding and decoding. The difference is that all encodings are performed on the low-resolution spatial domain, that is, the low-resolution current frame to be encoded is the low-resolution frame obtained by spatially downsampling the current frame x t , and the low-resolution temporal reference frame is the low-resolution reconstructed frame obtained by encoding the previous low-resolution frame x t-1

[0080] The main process is as follows: Combining the low-resolution temporal reference frame, perform motion encoding on the low-resolution current frame to obtain motion information, and through decoding, obtain the reconstructed motion information Encode the low-resolution current frame after context prediction by combining motion information with a low-resolution temporal reference frame, and obtain a low-resolution reconstructed frame through decoding

[0081] 2. Iteratively enhance context prediction

[0082] Through the aforementioned low-resolution encoding and decoding, low-resolution reconstructed motion information can be obtained and a low-resolution reconstructed frame along with the temporal reference frame enter the iterative enhanced context prediction process together. In the iterative enhanced context prediction, multiple enhancement stages will be carried out to continuously enhance the quality of motion information and context prediction. Among them, the low-resolution reconstructed motion information and the low-resolution reconstructed frame provide motion information and a spatial domain reference for the encoding of the current frame

[0083] As Figure 4 shown, it is a schematic diagram of iterative enhanced context prediction. The initial input and are equivalent to and the final output is equivalent to the final output context information C t . In each enhancement stage, the temporal reference frame the enhanced context information obtained in the previous enhancement stage and the enhanced motion information obtained in the previous enhancement stage will perform a motion enhancement together to obtain the enhanced motion information of the i-th enhancement stage (i represents the index of the enhancement stage). Obtaining the enhanced motion information of the i-th enhancement stage will perform motion information compensation on the temporal reference frame to obtain a temporally aligned reference frame. This temporally aligned reference frame will perform a context prediction enhancement together with the enhanced context information obtained in the previous enhancement stage to obtain the enhanced context information of the i-th enhancement stage

[0084] During the gap between multiple enhancement stages, the enhanced context and the enhanced motion information need to be continuously upsampled. The finally obtained enhanced context is the final output context information C tThere is no additional restriction on the gap of upsampling here. The user can set it according to the actual situation or experience, as long as it is ensured that the upsampling ratio is the same as the aforementioned spatial downsampling ratio, so that the resolution of the finally output context information is the same as that of the current frame image.

[0085] In addition, each model involved in the method provided by the embodiments of the present invention needs to be trained.

[0086] Denote the distortion loss as D(·,·), and the reconstructed frame is obtained by encoding The estimated bitrate requirement is Define the following three loss functions:

[0087]

[0088] Among them, λ is a parameter for controlling the bitrate, used to train models with different compression ratios; w is a parameter for regulating the low-resolution reconstruction quality and the full-resolution reconstruction quality. The smaller w means less bitrate allocation for low-resolution reconstruction, and the larger w means more bitrate allocation for low-resolution reconstruction.

[0089] The specific training process is as follows:

[0090] Training of the low-resolution video codec model: Use low-resolution video frame data and optimize it in combination with the loss function L1;

[0091] Training of the full-resolution codec model and the iterative enhanced context prediction model: Use low-resolution video frame data and full-resolution video frame data and optimize it in combination with the loss function L2;

[0092] Training of the overall model: Use low-resolution video frame data and full-resolution video frame data and optimize it in combination with the loss function L3.

[0093] In the above three optimization processes, the first two optimization processes can be executed in parallel or in any order. After the first two optimization processes are completed, enter the third optimization process (i.e., overall model training).

[0094] II. Example introduction.

[0095] Example 1: A method based on 4-fold low-resolution video codec and 6 enhancement stages.

[0096] Set the downsampling ratio to 4 times, use bicubic (bilinear interpolation algorithm) for downsampling, the number of enhancement stages N = 6, and perform upsampling every 2 enhancements (a total of 2 upsamplings, and no upsampling in the last 2 times).

[0097] Example 2: A method based on 4-fold low-resolution video codec and 9 enhancement stages.

[0098] Set the downsampling ratio to 4 times, use bilinear downsampling, the number of enhancement stages N = 9, and perform upsampling every 3 enhancements (a total of 2 upsamplings, and no upsampling in the last 2 times).

[0099] Example 3: A method based on 8-fold low-resolution video coding and decoding with 8 enhancement stages.

[0100] Set the downsampling ratio to 8 times, use bicubic downsampling, the number of enhancement stages N = 8, and perform upsampling every 2 enhancements (a total of 3 upsamplings, and no upsampling in the last 2 times).

[0101] The sampling ratio, sampling method, and number of enhancement stages involved in the above process are all examples. In actual applications, they can be set by users according to actual situations or experience, and the present invention does not make any restrictions.

[0102] III. Effect description.

[0103] Here, the effect of the present invention is mainly illustrated through experiments. In the experiments, the end-to-end video coding model DCVC-DC is used as the low-resolution encoder.

[0104] 1. Effect of large motion information modeling.

[0105] On large motion sequences, it can be observed that the quality of the reconstructed motion information after iterative enhancement has been significantly improved, as Figure 5 shown.

[0106] Taking the large motion sequence as an example, Figure 5 the motion information of RAFT GT (pseudo motion information) in [[ ]] is used as the pseudo label for measuring the quality of motion information; AEPE (Average End Point Error) is used as an index for measuring the quality of motion information, and the lower the AEPE, the more accurate the motion information representation; SSIM (Structural Similarity Index) is used as another index for measuring the quality of motion information, which reflects the effect of alignment using motion information, and the higher the SSIM, the better the alignment effect. The present invention has better motion information representation and motion alignment effects compared with DCVC-DC, thus indicating better reconstructed motion information.

[0107] 2. Effect of describing newly emerging objects.

[0108] In the front and rear frames with obvious newly emerging objects, it can be observed that the context prediction after iterative enhancement accurately describes the newly emerging objects, as Figure 6 shown.

[0109] In [[ ]] Figure 6In part (c), the object (hammer) that appears in the current frame box does not appear in the temporal reference frame shown in part (a) of Figure 6 Therefore, the context prediction generated only using the temporal reference in the existing method, that is Figure 6 in part (b) of Figure 6 has difficulty in describing the object and the prediction is inaccurate, thus leading to a decrease in coding efficiency. In the present invention, due to the introduction of a low-resolution spatial reference and continuously enhancing its description in the context prediction during the iterative enhanced context prediction process, it can be seen that the predicted context of the present invention, that is Figure 6 in part (d) of

[0110] 3. Positive effect on coding performance.

[0111] For the above two aspects of effects, the present invention has achieved excellent effects on general test data sets such as HEVC, UVG, MCL-JCV, and USTC-TD. As shown in Table 1, the BD-rate (RGB-BDBR) is used to measure the coding performance gain. A negative value represents the percentage of bitrate savings, and a positive value represents the percentage of bitrate increase, with the existing technology as the baseline for comparison. Applying low-resolution video coding and decoding and iterative enhanced context prediction can significantly improve the compression ratio.

[0112] Table 1: Performance comparison between the present invention and existing end-to-end video coding models

[0113]

[0114] The first column in Table 1 is the name of each type of solution. The first 6 rows are the performance performances of the models corresponding to the existing solutions on different data. The last row SEVC (ours) is the performance performance of the solution of the present invention on different data. In the experiment, the solution described in Example 1 above is adopted. The smaller the value in Table 1, the better the coding performance (the higher the bitrate savings). The comparison anchor is VTM (the encoder of the H.266 standard), and all the values of VTM are 0.0. Therefore, if the value is negative, it represents the bitrate savings relative to VTM.

[0115] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software or by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.

[0116] Embodiment 2

[0117] The present invention also provides an end-to-end video encoding and decoding system, which is mainly used to implement the method provided in the foregoing embodiment, such as Figure 7 As shown, the system mainly includes:

[0118] A downsampling module, which is used to perform spatial domain downsampling on the current frame to obtain a video frame called the low-resolution current frame;

[0119] A low-resolution encoding and decoding model, which is used to encode and decode the low-resolution current frame to obtain a low-resolution reconstructed frame and reconstructed motion information;

[0120] An iterative enhanced context prediction model, which is used to combine the full-resolution temporal reference frame, the low-resolution reconstructed frame and the reconstructed motion information, iteratively enhance the reconstructed motion information, and thereby enhance the context information. The finally output context information is used for the full-resolution encoding and decoding of the current frame, and, upsampling is performed during the gap of iterative enhancement, so that the resolution of the finally output context information is the same as the resolution of the current frame; wherein, the full-resolution temporal reference frame is the reconstructed frame obtained through full-resolution encoding and decoding when encoding and decoding the previous frame;

[0121] A full-resolution encoding and decoding model, which is used to encode and decode the current frame by combining the context information obtained by iterative enhanced context prediction to obtain a reconstructed frame.

[0122] Further, the encoding and decoding of the low-resolution current frame to obtain a low-resolution reconstructed frame and reconstructed motion information includes:

[0123] Combining the low-resolution temporal reference frame, the low-resolution current frame Performing motion encoding to obtain motion information, and obtaining reconstructed motion information through decoding Wherein, the low-resolution temporal reference frame is the low-resolution reconstructed frame obtained when encoding and decoding the previous frame t is the index of the current frame;

[0124] After performing context prediction by combining the motion information and the low-resolution temporal reference frame, encoding the low-resolution current frame, and obtaining the low-resolution reconstructed frame through decoding

[0125] Further, the combination of the full-resolution temporal reference frame, the low-resolution reconstructed frame and the reconstructed motion information, iteratively enhancing the reconstructed motion information, and thereby enhancing the context information, and the finally output context information being used for the encoding and decoding of the next frame includes:

[0126] In the i-th enhancement stage, combining the full-resolution temporal reference frame The enhanced context information obtained in the previous enhancement stage and the enhanced motion information obtained in the previous enhancement stage perform motion enhancement to obtain the enhanced motion information of the i-th enhancement stage where i is the index of the enhancement stage. When i = 1, the enhanced context information obtained in the previous enhancement stage is the low-resolution reconstructed frame the enhanced motion information obtained in the previous enhancement stage is the reconstructed motion information t is the index of the current frame;

[0127] Utilize the enhanced motion information to compensate the motion information of the full-resolution temporal reference frame to obtain a temporally aligned reference frame, and utilize the temporally aligned reference frame to perform context prediction enhancement on the enhanced context information obtained in the previous enhancement stage to obtain the enhanced context information of the i-th enhancement stage

[0128] Repeat N times continuously to obtain the enhanced context information of the N-th enhancement stage which is the finally output context information; and, between the gaps of iterative enhancement, i.e., between adjacent enhancement stages, upsample the enhanced context information, and the total upsampling magnification is the same as the spatial downsampling magnification.

[0129] Furthermore, the full-resolution codec model, the low-resolution codec model, and the iterative enhancement context prediction model are trained in the following manner:

[0130] Construct the following loss function:

[0131]

[0132] where t is the index of the current frame; L1, L2, and L3 are three loss functions, D(·,·) is the distortion loss, R(.) is the rate requirement; x t is the current frame, is the reconstructed frame; is the low-resolution current frame, is the low-resolution reconstructed frame; λ is the parameter for controlling the rate, and w is the parameter for regulating the low-resolution reconstruction quality and the full-resolution reconstruction quality;

[0133] Use the loss function L1 to optimize the low-resolution video codec model;

[0134] Use the loss function L2 to optimize the full-resolution codec model and the iterative enhancement context prediction model;

[0135] After that, the loss function L3 is used to optimize the low-resolution video coding and decoding model, the full-resolution coding and decoding model, and the iterative enhanced context prediction model.

[0136] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above-mentioned division of each functional module is used as an example. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above.

[0137] Embodiment III

[0138] The present invention also provides a processing device, as Figure 8 shown, which mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the foregoing embodiment.

[0139] Further, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, the memory, the input device, and the output device are connected through a bus.

[0140] In the embodiments of the present invention, the specific types of the memory, the input device, and the output device are not limited; for example:

[0141] The input device can be a touch screen, an image acquisition device, a physical button, or a mouse, etc.;

[0142] The output device can be a display terminal;

[0143] The memory can be a random access memory (RAM), or a non-volatile memory, such as a disk memory.

[0144] Embodiment IV

[0145] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the foregoing embodiment when the computer program is executed by a processor.

[0146] In the embodiments of the present invention, the readable storage medium as a computer-readable storage medium can be disposed in the foregoing processing device, for example, as the memory in the processing device. In addition, the readable storage medium can also be various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc.

[0147] As described above, it is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims. The information disclosed in the background art part of this article is only intended to deepen the understanding of the overall background art of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes the prior art known to those skilled in the art.

Claims

1. An end-to-end video encoding and decoding method, characterized in that: include: Downsampling processing: Downsample the current frame in the spatial domain, and the obtained video frame is called the low-resolution current frame; Low-resolution encoding and decoding: encoding and decoding the low-resolution current frame to obtain a low-resolution reconstructed frame and reconstructed motion information; Iteratively enhanced context prediction: combining the full-resolution temporal reference frame, the low-resolution reconstructed frame and the reconstructed motion information, iteratively enhancing the reconstructed motion information, and thereby enhancing the context information, the context information finally output is used for full-resolution encoding and decoding of the current frame, and upsampling is performed in the gaps of iterative enhancement so that the resolution of the context information finally output is the same as the resolution of the current frame; wherein the full-resolution temporal reference frame is a reconstructed frame obtained by full-resolution encoding and decoding when encoding and decoding the previous frame; Full-resolution encoding and decoding: Combine the context information obtained by iterative enhanced context prediction to encode and decode the current frame to obtain a reconstructed frame.

2. The end-to-end video encoding and decoding method according to claim 1, characterized in that: The encoding and decoding of the low-resolution current frame to obtain a low-resolution reconstructed frame and reconstructed motion information includes: Combined with the low-resolution temporal reference frame, the low-resolution current frame Perform motion coding to obtain motion information, and then decode to obtain reconstructed motion information Among them, the low-resolution temporal reference frame is the low-resolution reconstructed frame obtained when encoding and decoding the previous frame t is the index of the current frame; Combining motion information with low-resolution temporal reference frames for context prediction, the low-resolution current frame is encoded and decoded to obtain a low-resolution reconstructed frame.

3. The end-to-end video encoding and decoding method according to claim 1, characterized in that: The combining of the full-resolution temporal reference frame, the low-resolution reconstructed frame and the reconstructed motion information, iteratively enhancing the reconstructed motion information, and thereby enhancing the context information, and finally outputting the context information for encoding and decoding of the next frame includes: In the i-th enhancement stage, combined with the full-resolution temporal reference frame Enhanced context information obtained in the previous enhancement stage and the enhanced motion information obtained in the previous enhancement stage Perform motion enhancement to obtain the enhanced motion information of the i-th enhancement stage Where i is the index of the enhancement stage. When i=1, the enhanced context information obtained in the previous enhancement stage is Reconstruct frames for low resolution Enhanced motion information obtained in the previous enhancement stage To reconstruct motion information t is the index of the current frame; Using enhanced motion information Full resolution temporal reference frame Perform motion information compensation to obtain a time-domain aligned reference frame, and use the time-domain aligned reference frame to enhance the context information obtained in the previous enhancement stage Perform context prediction enhancement to obtain enhanced context information for the i-th enhancement stage Repeat N times to obtain the enhanced context information of the Nth enhancement stage That is the context information of the final output; and, in the gap of iterative enhancement, that is, between adjacent enhancement stages, the enhanced context information is upsampled, and the total upsampling ratio is the same as the spatial domain downsampling ratio.

4. The end-to-end video encoding and decoding method according to claim 1, characterized in that: Full-resolution encoding and decoding is implemented by a full-resolution encoding and decoding model, low-resolution encoding and decoding is implemented by a low-resolution encoding and decoding model, and iterative enhanced context prediction is implemented by an iterative enhanced context prediction model; the full-resolution encoding and decoding model, the low-resolution encoding and decoding model and the iterative enhanced context prediction model are trained in the following manner: Construct the following loss function: Where t is the index of the current frame; L1, L2 and L3 are three loss functions, D(·,·) is the distortion loss, R(.) is the bit rate requirement; x t is the current frame, To reconstruct the frame; is the low-resolution current frame, is the low-resolution reconstructed frame; λ is the parameter for controlling the bit rate, and w is the parameter for regulating the low-resolution reconstruction quality and the full-resolution reconstruction quality; Use the loss function L1 to optimize the low-resolution video encoding and decoding model; Use the loss function L2 to optimize the full-resolution encoding and decoding model and the iterative enhanced context prediction model; Afterwards, the loss function L3 is used to optimize the low-resolution video codec model, the full-resolution codec model and the iterative enhanced context prediction model.

5. An end-to-end video encoding and decoding system, characterized in that: include: The downsampling module is used to perform spatial downsampling on the current frame, and the obtained video frame is called a low-resolution current frame; A low-resolution encoding and decoding model, used for encoding and decoding the low-resolution current frame to obtain a low-resolution reconstructed frame and reconstructed motion information; An iterative enhancement context prediction model is used to combine a full-resolution temporal reference frame, a low-resolution reconstructed frame and reconstructed motion information, iteratively enhance the reconstructed motion information, and thereby enhance the context information, and the context information finally output is used for full-resolution encoding and decoding of the current frame, and upsampling is performed in the gaps of iterative enhancement so that the resolution of the context information finally output is the same as the resolution of the current frame; wherein the full-resolution temporal reference frame is a reconstructed frame obtained by full-resolution encoding and decoding when encoding and decoding a previous frame; The full-resolution encoding and decoding model is used to encode and decode the current frame in combination with the context information obtained by iterative enhanced context prediction to obtain a reconstructed frame.

6. The end-to-end video encoding and decoding system according to claim 5, characterized in that: The encoding and decoding of the low-resolution current frame to obtain a low-resolution reconstructed frame and reconstructed motion information includes: Combined with the low-resolution temporal reference frame, the low-resolution current frame Perform motion coding to obtain motion information, and then decode to obtain reconstructed motion information Among them, the low-resolution temporal reference frame is the low-resolution reconstructed frame obtained when encoding and decoding the previous frame t is the index of the current frame; Combining motion information with low-resolution temporal reference frames for context prediction, the low-resolution current frame is encoded and decoded to obtain a low-resolution reconstructed frame.

7. The end-to-end video encoding and decoding system according to claim 5, characterized in that: The combining of the full-resolution temporal reference frame, the low-resolution reconstructed frame and the reconstructed motion information, iteratively enhancing the reconstructed motion information, and thereby enhancing the context information, and finally outputting the context information for encoding and decoding of the next frame includes: In the i-th enhancement stage, combined with the full-resolution temporal reference frame Enhanced context information obtained in the previous enhancement stage and the enhanced motion information obtained in the previous enhancement stage Perform motion enhancement to obtain the enhanced motion information of the i-th enhancement stage Where i is the index of the enhancement stage. When i=1, the enhanced context information obtained in the previous enhancement stage is Reconstruct frames for low resolution Enhanced motion information obtained in the previous enhancement stage To reconstruct motion information t is the index of the current frame; Using enhanced motion information Full resolution temporal reference frame Perform motion information compensation to obtain a time-domain aligned reference frame, and use the time-domain aligned reference frame to enhance the context information obtained in the previous enhancement stage Perform context prediction enhancement to obtain enhanced context information for the i-th enhancement stage Repeat N times to obtain the enhanced context information of the Nth enhancement stage That is the context information of the final output; and, in the gap of iterative enhancement, that is, between adjacent enhancement stages, the enhanced context information is upsampled, and the total upsampling ratio is the same as the spatial domain downsampling ratio.

8. The end-to-end video encoding and decoding system according to claim 5, characterized in that: The full-resolution codec model, the low-resolution codec model and the iterative enhanced context prediction model are trained in the following manner: Construct the following loss function: Where t is the index of the current frame; L1, L2 and L3 are three loss functions, D(·,·) is the distortion loss, R(.) is the bit rate requirement; x t is the current frame, To reconstruct the frame; is the low-resolution current frame, is the low-resolution reconstructed frame; λ is the parameter for controlling the bit rate, and w is the parameter for regulating the low-resolution reconstruction quality and the full-resolution reconstruction quality; Use the loss function L1 to optimize the low-resolution video encoding and decoding model; Use the loss function L2 to optimize the full-resolution encoding and decoding model and the iterative enhanced context prediction model; Afterwards, the loss function L3 is used to optimize the low-resolution video codec model, the full-resolution codec model and the iterative enhanced context prediction model.

9. A processing device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 4.

10. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • End-to-end intelligent video coding method and device

    CN115278262A

  • Remote sensing image super-resolution method based on context sensing edge enhancement

    CN117217997A

  • Video coding method and system

    CN117939146A

  • Neural video coding and decoding

    CN118381944A

  • Method, apparatus, and recording medium for image encoding / decoding

    CN119325707A