A method and device for optimizing time domain context

By generating and fusing directional temporal context in video coding, the problems of insufficient motion estimation accuracy and temporal context quality in motion compensation methods are solved, and the efficiency and accuracy of video coding are improved.

CN120128725BActive Publication Date: 2025-09-16UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510449828.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-09-16
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

Existing motion compensation methods in video coding suffer from poor motion estimation accuracy and poor temporal context quality, resulting in inaccurate inter-frame correlation modeling and high data redundancy.

Method used

By obtaining the reference frame and reconstructing the motion vector, the predicted frame is generated, and the motion estimation network is used to extract the inter-frame correlation, determine the directional motion vector for time domain alignment and feature extraction, generate the directional time domain context, combine the propagation time domain context for feature fusion calculation, and optimize the time domain context quality.

Benefits of technology

It significantly improves the accuracy of motion compensation, optimizes the video encoding process, reduces data transmission and storage requirements, and achieves more efficient video encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128725B_ABST
    Figure CN120128725B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for optimizing time domain context, which obtains reference frames and reconstructs motion vectors and generates predicted frames; extracts the inter-frame correlation between the predicted frame and the reference frame through a motion estimation network, and determines the directional motion vector based on the inter-frame correlation; uses the directional motion vector and the reference frame to perform time domain alignment and feature extraction to generate a directional time domain context; performs feature fusion calculation on the directional time domain context and the propagation time domain context to obtain a compensated time domain context. The present invention is based on an end-to-end video coding framework of conditional coding, and uses reference frames, reference features, and reconstructed motion vectors to perform iterative calculations at the decoding end to fine-tune the time domain context, thereby significantly improving the quality and accuracy of the time domain context. This optimization effectively improves the accuracy of motion compensation and optimizes the entire video coding process, reducing unnecessary data transmission and storage requirements, and ultimately achieving more efficient video coding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video coding and standard technology, and in particular to a method and device for optimizing time domain context. Background Art

[0002] In recent years, end-to-end video coding frameworks have gained increasing attention. In particular, within conditional coding frameworks, motion compensation, as an important technical tool, plays a crucial role in reducing video data redundancy and improving coding efficiency. Currently, motion compensation methods in video coding fall into two main categories: feature-domain motion compensation and reference-frame refresh-based motion compensation. However, these existing methods still face challenges in practical applications, primarily due to poor motion estimation accuracy and poor temporal context quality.

[0003] Among them, feature-domain motion compensation methods can extract high-quality inter-frame features through deep learning models. However, in the prediction chain, the updated reconstructed features may contain information irrelevant to the encoding, resulting in error accumulation, which affects the accuracy of inter-frame correlation modeling. On the other hand, methods based on reference frame refresh improve motion estimation accuracy by continuously updating the reference frame. However, manually switching reference features cannot fully utilize the adaptive learning capabilities of neural networks and, in some cases, cannot optimize the efficiency of the reference mechanism.

[0004] Therefore, how to improve the quality of temporal context and thus enhance the effect of inter-frame redundancy removal has become a key issue that needs to be urgently addressed in the current end-to-end video coding framework. Summary of the Invention

[0005] In view of this, an embodiment of the present invention provides a method and apparatus for optimizing a time domain context to solve the problem of poor quality of the time domain context.

[0006] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:

[0007] A first aspect of the present invention discloses a method for optimizing a time domain context, the method comprising:

[0008] Get the reference frame and reconstruct the motion vector and generate the predicted frame;

[0009] Extracting inter-frame correlation between the predicted frame and the reference frame through a motion estimation network, and determining a directional motion vector according to the inter-frame correlation;

[0010] Performing temporal alignment and feature extraction using the directional motion vector and a reference frame to generate a directional temporal context;

[0011] The directional time domain context and the propagation time domain context are subjected to feature fusion calculation to obtain a compensation time domain context; the propagation time domain context is generated according to the reference features in the prediction chain.

[0012] Preferably, the obtaining of the reference frame, reconstructing the motion vector and generating the predicted frame includes:

[0013] Get reference frame and reconstruct motion vectors;

[0014] Performing time domain alignment on the reference frame and the reconstructed motion vector;

[0015] Motion compensation is performed using the time-domain aligned reference frame and the reconstructed motion vector to generate a predicted frame.

[0016] Preferably, the performing time domain alignment and feature extraction using the directional motion vector and the reference frame to generate a directional time domain context includes:

[0017] Adjusting the reference frame using the directional motion vector to align the reference frame with the predicted frame in time domain;

[0018] Through the deep neural network model, features are extracted from the time-aligned reference frames according to the preset convolution kernel to generate directional time domain context.

[0019] Preferably, the adjusting the temporal alignment between the reference frame and the predicted frame by using the directional motion vector includes:

[0020] Pixels in the reference frame are adjusted using the directional motion vector so that the reference frame and the predicted frame are temporally aligned in the pixel domain.

[0021] Preferably, after extracting the inter-frame correlation between the predicted frame and the reference frame through a motion estimation network and determining the directional motion vector according to the inter-frame correlation, the method further includes:

[0022] Extract deep reference features of reference frames using deep neural networks;

[0023] Accordingly, performing temporal alignment and feature extraction using the directional motion vector and the reference frame to generate a directional temporal context includes:

[0024] Temporally aligning the directional motion vector and the deep reference feature in a feature domain;

[0025] Through the deep neural network model, features are extracted from the time-aligned reference frames according to the preset convolution kernel to generate directional time domain context.

[0026] A second aspect of the present invention discloses a device for optimizing a time domain context, the device comprising:

[0027] An acquisition unit, configured to acquire a reference frame, reconstruct a motion vector, and generate a predicted frame;

[0028] a determining unit, configured to extract an inter-frame correlation between the predicted frame and the reference frame through a motion estimation network, and determine a directional motion vector according to the inter-frame correlation;

[0029] A generating unit, configured to perform temporal alignment and feature extraction using the directional motion vector and a reference frame to generate a directional temporal context;

[0030] A feature fusion calculation unit is used to perform feature fusion calculation on the directional time domain context and the propagation time domain context to obtain a compensated time domain context; the propagation time domain context is generated according to the reference features in the prediction chain.

[0031] Preferably, the acquisition unit includes:

[0032] An acquisition module, used for acquiring a reference frame and reconstructing a motion vector;

[0033] A time domain alignment module, configured to perform time domain alignment on the reference frame and the reconstructed motion vector;

[0034] The generating module is used for performing motion compensation by using the time-domain aligned reference frame and the reconstructed motion vector to generate a predicted frame.

[0035] Preferably, the generating unit includes:

[0036] A first time domain alignment module, configured to adjust the reference frame using the directional motion vector to align the reference frame with the predicted frame in time domain;

[0037] The first extraction module is used to extract features from the time-domain aligned reference frame through a deep neural network model according to a preset convolution kernel to generate a directional time domain context.

[0038] Preferably, the first time-domain alignment module is specifically configured to: adjust pixels in the reference frame using the directional motion vector, so that the reference frame and the predicted frame are time-domain aligned in the pixel domain.

[0039] Preferably, the device further comprises:

[0040] An extraction unit, configured to extract deep reference features of a reference frame using a deep neural network;

[0041] Accordingly, the generating unit includes:

[0042] A second time domain alignment module, configured to perform time domain alignment on the directional motion vector and the deep reference feature in a feature domain;

[0043] The second extraction module is used to extract features from the time-domain aligned reference frames through a deep neural network model according to a preset convolution kernel to generate a directional time domain context.

[0044] Based on the above-mentioned embodiment of the present invention, a method and device for optimizing the time domain context are provided, which obtains a reference frame and a reconstructed motion vector and generates a predicted frame; extracts the inter-frame correlation between the predicted frame and the reference frame through a motion estimation network, and determines the directional motion vector based on the inter-frame correlation; uses the directional motion vector and the reference frame to perform time domain alignment and feature extraction to generate a directional time domain context; performs feature fusion calculation on the directional time domain context and the propagation time domain context to obtain a compensated time domain context. The present invention is based on an end-to-end video coding framework of conditional coding, and uses reference frames, reference features, and reconstructed motion vectors to perform iterative calculations at the decoding end to fine-tune the time domain context, thereby significantly improving the quality and accuracy of the time domain context. This optimization effectively improves the accuracy of motion compensation, optimizes the entire video coding process, reduces unnecessary data transmission and storage requirements, and ultimately achieves more efficient video coding. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0046] Figure 1 A schematic diagram of an end-to-end video coding system based on decoding-end motion estimation provided by an embodiment of the present invention;

[0047] Figure 2 A flowchart of a time domain context optimization method provided by an embodiment of the present invention;

[0048] Figure 3 A schematic diagram of a flow chart of a method for optimizing a time domain context according to an embodiment of the present invention;

[0049] Figure 4 A graph showing the relationship between reconstruction quality and bit rate as the number of encoded frames is provided in an embodiment of the present invention;

[0050] Figure 5 Another flow chart of a time domain context optimization method provided by an embodiment of the present invention;

[0051] Figure 6 Another flow chart of a time domain context optimization method provided by an embodiment of the present invention;

[0052] Figure 7Another schematic diagram of a flow chart of a method for optimizing a time domain context according to an embodiment of the present invention;

[0053] Figure 8 A structural block diagram of a time domain context optimization device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0055] In this application, the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0056] As we know from the background, in feature-domain motion compensation methods, the updated reconstructed features in the prediction chain may contain information irrelevant to the encoding, leading to error accumulation and affecting the accuracy of inter-frame correlation modeling. On the other hand, the manual switching of reference features based on reference frame refresh methods cannot fully utilize the adaptive learning capabilities of neural networks and, in some cases, cannot optimize the efficiency of the reference mechanism.

[0057] Therefore, an embodiment of the present invention provides a method and device for optimizing the time domain context, which obtains a reference frame and a reconstructed motion vector and generates a predicted frame; extracts the inter-frame correlation between the predicted frame and the reference frame through a motion estimation network, and determines the directional motion vector based on the inter-frame correlation; uses the directional motion vector and the reference frame to perform time domain alignment and feature extraction to generate a directional time domain context; performs feature fusion calculation on the directional time domain context and the propagation time domain context to obtain a compensated time domain context. The present invention is based on an end-to-end video coding framework of conditional coding, and uses reference frames, reference features, and reconstructed motion vectors to perform iterative calculations at the decoding end to fine-tune the time domain context, thereby significantly improving the quality and accuracy of the time domain context. This optimization effectively improves the accuracy of motion compensation, optimizes the entire video coding process, reduces unnecessary data transmission and storage requirements, and ultimately achieves more efficient video coding.

[0058] It is understandable that, since this method retains the temporal context branch of the original framework, it can be widely applied to any end-to-end video coding framework based on conditional coding, and has strong versatility and application potential. Figure 1 The diagram shows an end-to-end video coding system based on decoding-end motion estimation.

[0059] like Figure 1 As shown, when the video is decoded, the input frame x t Input motion estimation network (i.e. Figure 1 "Motion Estimation"), output motion vector v t . Motion vector v t Input to the motion coding module and output the reconstructed motion vector .

[0060] In the process of temporal context mining, according to the reconstructed motion vector and propagation reference features , output the propagation time domain context (such as the propagation time domain context of different scales ).

[0061] In the temporal context modulation process, according to the reconstructed motion vector and propagation time context at different scales , and the reference frame , output compensated temporal context (e.g. ).

[0062] Then, the entropy model, context encoder and context decoder are used to calculate the input frame x. t , compensate for the temporal context and output the reconstructed features In addition, the frame generator is used to finally generate the decoded frame, that is, the decoded frame , complete the video reconstruction.

[0063] Figure 1 The entire video encoding and decoding process from motion estimation to frame generation is demonstrated, involving the mining and modulation of temporal context, as well as entropy encoding and decoding. Figure 2 , shows a flowchart of a time domain context optimization method provided by an embodiment of the present invention, the method comprising:

[0064] It is understandable that this method can be used to optimize the propagation time domain context of any scale at the decoding end to obtain the compensated time domain context. The following takes the propagation time domain context of a single scale as an example for detailed description.

[0065] Step S201: Acquire a reference frame and reconstruct a motion vector and generate a predicted frame.

[0066] In the specific implementation of step S201, Figure 3 The content shown, get the reference frame and reconstructed motion vectors ; Set the reference frame and reconstructed motion vectors Perform time domain alignment; use the time domain aligned reference frame and reconstructed motion vector to perform motion compensation and generate the predicted frame .

[0067] Specifically, the reference frame and the reconstructed motion vector are aligned in time domain, and based on the time domain alignment, the prediction frame is generated by using the image data and motion vector of the reference frame through a motion compensation process.

[0068] It is understandable that the predicted frame Based on the reference frame and reconstructed motion vectors The calculated,prediction frame can be considered as an approximation or estimation of the,current frame.

[0069] In practical applications, the process of obtaining a reference frame, reconstructing motion vectors, and generating a predicted frame is as follows: first, image data is obtained from a reference frame in video coding, and then processed based on the reconstructed motion vectors obtained through the motion estimation process. The motion vectors describe the movement of pixel blocks between the reference frame and the target frame.

[0070] Secondly, the purpose of temporal alignment is to align the reference frame and the reconstructed motion vectors with the target frame (current frame) on the temporal axis. This step can be understood as ensuring that the image motion between the reference frame and the current frame is temporally aligned, especially when there are frames with different timestamps.

[0071] After time domain alignment, the time domain aligned reference frame and motion vector are used to perform motion compensation on the current frame to generate a predicted frame.

[0072] Step S202: extracting the inter-frame correlation between the predicted frame and the reference frame through a motion estimation network, and determining a directional motion vector according to the inter-frame correlation.

[0073] In the specific implementation of step S202, combined with Figure 3 The content shown, the predicted frame is extracted through the motion estimation network With reference frame The inter-frame correlation between the two frames gives the directional motion vector .

[0074] For example, the optical flow between the predicted frame and the reference frame is extracted through the optical flow estimation network, that is, the directional optical flow is extracted, and the directional motion vector is determined according to the directional optical flow. .

[0075] As you can understand, in video encoding or processing, a reference frame is used to predict the image information of the current frame (the predicted frame). The inter-frame correlation between the two frames describes the motion pattern between them—that is, how the pixels in one frame move to the other. By calculating this inter-frame correlation, the encoder can use fewer bits to describe the video frame, thereby achieving compression.

[0076] It should be noted that the directional motion vector is usually a two-dimensional vector, which indicates the direction and size of the pixel block moving from the reference frame to the predicted frame.

[0077] Step S203: performing temporal alignment and feature extraction using the directional motion vector and the reference frame to generate a directional temporal context.

[0078] In the specific implementation of step S203, first, the reference frame is adjusted using the directional motion vector to achieve temporal alignment between the reference frame and the predicted frame; second, features are extracted from the temporally aligned reference frame to generate a directional temporal context.

[0079] That is to say, the direction and size of the directional motion vector are used to adjust the reference frame so that it is more accurately aligned with the predicted frame on the time axis, thereby extracting features and generating a directional temporal context.

[0080] It should be noted that the deep neural network model can be used to extract features from the time-aligned reference frame according to the preset convolution kernel (for example, 3×3 convolution kernel) to generate the directional time domain context (such as Figure 3 Directed temporal context shown ).

[0081] Step S204: performing feature fusion calculation on the directional time domain context and the propagation time domain context to obtain the compensated time domain context.

[0082] It should be noted that the propagation time domain context is generated based on the reference features in the prediction chain.

[0083] In the specific implementation of step S204, Figure 3 The content shown obtains the propagation time domain context generated by the reference features of the propagation update in the prediction chain ; Directed temporal context and propagation time domain context Perform feature fusion to form compensated time domain context .

[0084] It can be understood that feature fusion refers to combining information from different contexts to obtain richer feature representations. Here, the directional temporal context and propagation time domain context Fusion can take advantage of their respective advantages and enhance the expressiveness of time domain features.

[0085] Fusion can be performed using several common methods, such as weighted averaging, concatenation, and element-by-element addition or multiplication.

[0086] Combine Figure 2 The optimization method based on the single-scale time domain context shown in FIG. 3 can be extended to three-scale propagation time domain context for time domain context optimization in practical applications to obtain the compensated time domain context.

[0087] Specifically, mining obtains high-scale propagation time domain context , medium-scale propagation time context (corresponding to the scale after downsampling twice the original resolution) and the low-scale propagation time domain context (corresponding to the scale after downsampling four times the original resolution).

[0088] First, for high-scale propagation time context , optimize the time domain context, and obtain the process of compensating the time domain context as follows Figure 2 Steps S201 to S204 are shown.

[0089] Secondly, for medium-scale propagation time context , the process of optimizing the time domain context and obtaining the compensated time domain context is as follows:

[0090] Reference frame and directional motion vectors Perform bilinear downsampling twice and convert the downsampled directional motion vector The value is reduced to 1 / 2 of the original value. Then, using the downsampled reference frame and directional motion vectors ,according to Figure 2 Steps S203 to S204 shown in the figure are used to obtain the medium-scale directional temporal context. , and then propagate the temporal context with the medium scale Perform fusion calculations to obtain compensation time domain context .

[0091] Finally, for low-scale propagation time context , the process of optimizing the time domain context and obtaining the compensated time domain context is as follows:

[0092] Reference frame and directional motion vectors Perform bilinear four-fold downsampling and convert the four-fold downsampled motion vector The value of is reduced to 1 / 4 of the original value. Then, the reference frame after four times downsampling is used. and directional motion vectors ,according to Figure 2 Steps S203 to S204 shown generate the lowest scale directional temporal context Finally, the lowest-scale directional temporal context Propagating temporal context with the lowest scale Perform fusion calculation to obtain the lowest scale compensated time domain context .

[0093] Combine Figure 4 The relationship between reconstruction quality and bit rate as a function of the number of coded frames is shown in the graph. Under a coding configuration with a total of 96 coded frames and Intra-period (IP) set to -1, this method embodiment can more effectively mitigate the error accumulation problem caused by the coding method in a long prediction chain compared to DCVC-FM.

[0094] In this embodiment of the present invention, an end-to-end video coding framework based on conditional coding uses reference frames, reference features, and reconstructed motion vectors for iterative calculations at the decoding end to fine-tune the temporal context, significantly improving its quality and accuracy. This optimization effectively improves the accuracy of motion compensation and the entire video coding process, reducing unnecessary data transmission and storage requirements, ultimately achieving more efficient video coding.

[0095] In some specific embodiments, see Figure 5 , shows another flow chart of a time domain context optimization method provided by an embodiment of the present invention, the method comprising:

[0096] Step S501: Obtain a reference frame and reconstruct a motion vector and generate a predicted frame.

[0097] Step S502: extracting the inter-frame correlation between the predicted frame and the reference frame through a motion estimation network, and determining a directional motion vector according to the inter-frame correlation.

[0098] It should be noted that the specific implementation of step S501 and step S502 is detailed in the above Figure 2 The relevant contents of step S201 and step S202 are not repeated here.

[0099] Step S503: Based on the directional motion vector and the reference frame, perform pixel-domain temporal alignment and feature extraction to generate a directional temporal context.

[0100] Specifically, using directional motion vector Adjust reference frame pixels in the reference frame With predicted frame Temporal alignment is performed in the pixel domain. Then, the temporally aligned reference frame With predicted frame Input into the feature extraction network for feature extraction to generate directional time domain context .

[0101] Step S504: performing feature fusion calculation on the directional time domain context and the propagation time domain context to obtain the compensated time domain context.

[0102] It should be noted that the specific implementation of step S504 is detailed in the above Figure 2 The relevant content of step S204 is not repeated here.

[0103] In some specific embodiments, see Figure 6 , shows another flow chart of a time domain context optimization method provided by an embodiment of the present invention, the method comprising:

[0104] Step S601: Obtain a reference frame and reconstruct a motion vector and generate a predicted frame.

[0105] Step S602: extracting the inter-frame correlation between the predicted frame and the reference frame through a motion estimation network, and determining a directional motion vector according to the inter-frame correlation.

[0106] It should be noted that the specific implementation of step S601 and step S602 is detailed in the above Figure 2 The relevant contents of step S201 and step S202 are not repeated here.

[0107] Step S603: Utilize a deep neural network to extract deep reference features of the reference frame.

[0108] Specific, combined Figure 7 The content shown uses a deep neural network to extract reference frames Deep reference features.

[0109] Step S604: Based on the directional motion vector and the deep reference features of the reference frame, perform temporal alignment and feature extraction of the feature domain to generate a directional temporal context.

[0110] In the specific implementation of step S604, the directional motion vector The deep reference features of the reference frame are aligned in the feature domain in the time domain; the deep neural network model is used to extract features from the time-aligned reference frame according to the preset convolution kernel to generate a directional time domain context.

[0111] For example, the reference frame with time domain alignment of feature domain is input into the deep neural network model, and features are extracted based on the preset convolution kernel (3×3), and directional time domain context is further generated. .

[0112] Step S605: performing feature fusion calculation on the directional time domain context and the propagation time domain context to obtain the compensated time domain context.

[0113] Combined with the actual application results, it can be seen that the time domain context optimization method proposed in the embodiment of the present invention has achieved excellent results on common test datasets such as the High Efficiency Video Coding (HEVC) standard, the University of Vigo (UVG) universal video image dataset, and the Large Scale Video Coding (MCL-JCV) dataset.

[0114] As shown in Table 1, BD-rate (RGB-BDBR) is used to measure the encoding performance gain. Negative values ​​indicate the percentage of bit rate savings, and positive values ​​indicate the percentage of bit rate increases. As shown in Table 1, compared with the existing traditional coding standard H.266 / VVC standard reference software VTM-13.2 and the best end-to-end video coding method - Deep Conditional Video Compression with Flow Modeling (DCVC-FM), the present invention Figure 5 and Figure 6 The illustrated method embodiments can all achieve higher compression performance.

[0115] Table 1

[0116] HEVC B HEVC C HEVC D HEVC E MCL-JCV UVG VTM-13.2 0.0 0.0 0.0 0.0 0.0 0.0 DCVC-FM -11.7 -8.2 -28.5 -26.6 -12.5 -24.3 Figure 5 Example of the method shown -16.8 -14.8 -34.6 -36.5 -22.8 -35.6 Figure 6 Example of the method shown -14.8 -11.1 -29.5 -31.8 -21.3 -33.5

[0117] As shown in Table 1, it can be preliminarily determined that Figure 5 The illustrated method embodiment can achieve higher compression performance.

[0118] Corresponding to the above embodiment of the present invention, a method for optimizing a time domain context is provided. Figure 8 The structure block diagram of a time domain context optimization device is shown.

[0119] The device includes: an acquisition unit 801, a determination unit 802, a generation unit 803 and a feature fusion calculation unit 804.

[0120] The acquisition unit 801 is configured to acquire a reference frame and reconstruct a motion vector and generate a predicted frame.

[0121] The determination unit 802 is configured to extract the inter-frame correlation between the prediction frame and the reference frame through a motion estimation network, and determine a directional motion vector according to the inter-frame correlation.

[0122] The generating unit 803 is configured to perform time domain alignment and feature extraction using the directional motion vector and the reference frame to generate a directional time domain context.

[0123] The feature fusion calculation unit 804 is used to perform feature fusion calculation on the directional time domain context and the propagation time domain context to obtain the compensation time domain context; the propagation time domain context is generated according to the reference features in the prediction chain.

[0124] In this embodiment of the present invention, an end-to-end video coding framework based on conditional coding uses reference frames, reference features, and reconstructed motion vectors for iterative calculations at the decoding end to fine-tune the temporal context, significantly improving its quality and accuracy. This optimization effectively improves the accuracy of motion compensation and the entire video coding process, reducing unnecessary data transmission and storage requirements, ultimately achieving more efficient video coding.

[0125] Combine Figure 8 The content shown, the acquisition unit 801, includes: an acquisition module, a time domain alignment module and a generation module.

[0126] The acquisition module is used to obtain the reference frame and reconstruct the motion vector.

[0127] The time domain alignment module is used to perform time domain alignment on the reference frame and the reconstructed motion vector.

[0128] The generation module is used to perform motion compensation using the time-domain aligned reference frame and the reconstructed motion vector to generate a predicted frame.

[0129] Combine Figure 8 The content shown, the generating unit 803, includes: a first time domain alignment module and a first extraction module.

[0130] The first time domain alignment module is used to adjust the reference frame using the directional motion vector to align the reference frame with the predicted frame in the time domain.

[0131] The first time domain alignment module is specifically used to adjust the pixels in the reference frame using the directional motion vector so as to perform time domain alignment between the reference frame and the prediction frame in the pixel domain.

[0132] The first extraction module is used to extract features from the time-domain aligned reference frame through a deep neural network model according to a preset convolution kernel to generate a directional time domain context.

[0133] Combine Figure 8 As shown in the content, the device also includes an extraction unit for extracting deep reference features of the reference frame using a deep neural network.

[0134] Correspondingly, the generating unit 803 includes: a second time domain alignment module and a second extraction module.

[0135] The second time domain alignment module is used to perform time domain alignment on the directional motion vector and the deep reference feature in the feature domain.

[0136] The second extraction module is used to extract features from the time-domain aligned reference frames through a deep neural network model according to a preset convolution kernel to generate a directional time domain context.

[0137] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0138] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0139] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A time domain context optimization method, characterized in that: The method comprises: Get the reference frame and reconstruct the motion vector and generate the predicted frame; Extracting inter-frame correlation between the predicted frame and the reference frame through a motion estimation network, and determining a directional motion vector according to the inter-frame correlation; Performing temporal alignment and feature extraction using the directional motion vector and a reference frame to generate a directional temporal context; The directional time domain context and the propagation time domain context are subjected to feature fusion calculation to obtain a compensation time domain context; the propagation time domain context is generated according to the reference features in the prediction chain.

2. The method according to claim 1, characterized in that The obtaining of the reference frame, reconstructing the motion vector and generating the predicted frame includes: Get reference frame and reconstruct motion vectors; Performing time domain alignment on the reference frame and the reconstructed motion vector; Motion compensation is performed using the time-domain aligned reference frame and the reconstructed motion vector to generate a predicted frame.

3. The method according to claim 1, characterized in that The step of performing temporal alignment and feature extraction using the directional motion vector and the reference frame to generate a directional temporal context includes: Adjusting the reference frame using the directional motion vector to align the reference frame with the predicted frame in time domain; Through the deep neural network model, features are extracted from the time-aligned reference frames according to the preset convolution kernel to generate directional time domain context.

4. The method according to claim 3, characterized in that The adjusting the temporal alignment between the reference frame and the predicted frame by using the directional motion vector includes: Pixels in the reference frame are adjusted using the directional motion vector so that the reference frame and the predicted frame are temporally aligned in the pixel domain.

5. The method according to claim 1, characterized in that After extracting the inter-frame correlation between the predicted frame and the reference frame through a motion estimation network and determining a directional motion vector according to the inter-frame correlation, the method further includes: Extract deep reference features of reference frames using deep neural networks; Accordingly, performing temporal alignment and feature extraction using the directional motion vector and the reference frame to generate a directional temporal context includes: Temporally aligning the directional motion vector and the deep reference feature in a feature domain; Through the deep neural network model, features are extracted from the time-aligned reference frames according to the preset convolution kernel to generate directional time domain context.

6. A device for optimizing time domain context, characterized in that: The device comprises: An acquisition unit, configured to acquire a reference frame, reconstruct a motion vector, and generate a predicted frame; a determining unit, configured to extract an inter-frame correlation between the predicted frame and the reference frame through a motion estimation network, and determine a directional motion vector according to the inter-frame correlation; A generating unit, configured to perform temporal alignment and feature extraction using the directional motion vector and a reference frame to generate a directional temporal context; A feature fusion calculation unit is used to perform feature fusion calculation on the directional time domain context and the propagation time domain context to obtain a compensated time domain context; the propagation time domain context is generated according to the reference features in the prediction chain.

7. The device according to claim 6, characterized in that The acquisition unit includes: An acquisition module, used for acquiring a reference frame and reconstructing a motion vector; A time domain alignment module, configured to perform time domain alignment on the reference frame and the reconstructed motion vector; The generating module is used for performing motion compensation by using the time-domain aligned reference frame and the reconstructed motion vector to generate a predicted frame.

8. The device according to claim 6, characterized in that The generating unit comprises: A first time domain alignment module, configured to adjust the reference frame using the directional motion vector to align the reference frame with the predicted frame in time domain; The first extraction module is used to extract features from the time-domain aligned reference frame through a deep neural network model according to a preset convolution kernel to generate a directional time domain context.

9. The device according to claim 8, characterized in that The first time-domain alignment module is specifically configured to adjust pixels in the reference frame using the directional motion vector, so that the reference frame and the predicted frame are time-domain aligned in the pixel domain.

10. The device according to claim 8, characterized in that The device further comprises: An extraction unit, configured to extract deep reference features of a reference frame using a deep neural network; Accordingly, the generating unit includes: A second time domain alignment module, configured to perform time domain alignment on the directional motion vector and the deep reference feature in a feature domain; The second extraction module is used to extract features from the time-domain aligned reference frames through a deep neural network model according to a preset convolution kernel to generate a directional time domain context.

Citation Information

Patent Citations

  • Video quality enhancement method and device, video quality compression method and device, terminal and medium

    CN117714714A

  • Learable video coding method, system and device and storage medium

    CN117750034A