Method, apparatus, and medium for visual data processing
A joint local and global motion compensation scheme using optical flow and deformable neural networks addresses inefficiencies in existing video compression methods, enhancing coding quality and efficiency by capturing both local and global redundancies.
Patent Information
- Application Number
- PCT/CN2025/075574
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-03
- Filing Date
- 2025-01-27
- Publication Date
- 2025-08-07
AI Technical Summary
Current flow-based motion estimation and compensation methods in learning-based video compression struggle with large motions and limited receptive fields, leading to inefficiencies in capturing global redundancy and local redundancy.
Implement a joint local and global motion compensation scheme using optical flow networks, deformable neural networks, and cross attention networks to enhance coding quality and efficiency by combining local and global motion estimation and compensation techniques.
The proposed method improves coding quality and efficiency by effectively handling large motions and capturing both local and global redundancies, optimizing video compression processes.
Smart Images

Figure CN2025075574_07082025_PF_FP_ABST
Abstract
Description
METHOD, APPARATUS, AND MEDIUM FOR VISUAL DATA PROCESSINGFIELDS
[0001] Embodiments of the present disclosure relates generally to video processing techniques, and more particularly, to joint local and global motion compensation for learning-based video compression.BACKGROUND
[0002] The past decade has witnessed the rapid development of deep learning in a variety of areas, especially in computer vision and image processing. Neural network was invented originally with the interdisciplinary research of neuroscience and mathematics. It has shown strong capabilities in the context of non-linear transform and classification. Neural network-based image / video compression technology has gained significant progress during the past half decade. It is reported that the latest neural network-based image compression algorithm achieves comparable rate-distortion (R-D) performance with Versatile Video Coding (VVC) . With the performance of neural image compression continually being improved, neural network-based video compression has become an actively developing research area. However, coding efficiency and / or coding quality of neural network-based visual data coding is generally expected to be further improved.SUMMARY
[0003] Embodiments of the present disclosure provide a solution for visual data processing.
[0004] In a first aspect, a method for visual data processing is proposed. The method comprises: constructing, for a conversion between visual data and a bitstream of the visual data, at least one of a motion compensation scheme or a motion estimation scheme based on at least one of: an optical flow network, a deformable neural network, or a cross attention network; processing the visual data by applying the at least one of the motion compensation scheme or the motion estimation scheme in a learning-based video compression; and performing the conversion based on the processed visual data.
[0005] According to the method in accordance with the first aspect of the present disclosure, the motion compensation scheme and / or the motion estimation scheme is constructed. In addition, the visual data is processed by applying the motion compensation scheme and / or the motion estimation scheme. In this way, the proposed method can advantageously improve the coding quality and coding efficiency.
[0006] In a second aspect, an apparatus for visual data processing is proposed. The apparatus comprises a processor and a non-transitory memory with instructions thereon. The instructions upon execution by the processor, cause the processor to perform a method in accordance with the first aspect of the present disclosure.
[0007] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to perform a method in accordance with the first aspect of the present disclosure.
[0008] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream of visual data which is generated by a method performed by an apparatus for visual data processing. The method comprises: constructing at least one of a motion compensation scheme or a motion estimation scheme based on at least one of: an optical flow network, a deformable neural network, or a cross attention network; processing the visual data by applying the at least one of the motion compensation scheme or the motion estimation scheme in a learning-based video compression; and generating the bitstream based on the processed visual data.
[0009] In a fifth aspect, a method for storing a bitstream of visual data is proposed. The method comprises: constructing at least one of a motion compensation scheme or a motion estimation scheme based on at least one of: an optical flow network, a deformable neural network, or a cross attention network; processing the visual data by applying the at least one of the motion compensation scheme or the motion estimation scheme in a learning-based video compression; generating the bitstream based on the processed visual data; storing the bitstream in a non-transitory computer-readable recording medium.
[0010] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Through the following detailed description with reference to the accompanying drawings, the above and other objectives, features, and advantages of example embodiments of the present disclosure will become more apparent. In the example embodiments of the present disclosure, the same reference numerals usually refer to the same components.
[0012] Fig. 1 illustrates a block diagram that illustrates an example visual data coding system, in accordance with some embodiments of the present disclosure;
[0013] Fig. 2 illustrates an overall framework of the proposed method;
[0014] Fig. 3 illustrates an illustration of the joint local and global motion compensation module (LGMC) at encoder side;
[0015] Fig. 4 illustrates an illustration of the joint local and global motion compensation module (LGMC) at decoder side;
[0016] Fig. 5 illustrates an overall framework of the proposed method;
[0017] Fig. 6 illustrates an illustration of the joint local and global motion compensation module (LGMC) at encoder side;
[0018] Fig. 7 illustrates an illustration of the joint local and global motion compensation module (LGMC) at decoder side;
[0019] Fig. 8 illustrates a flowchart of a method for visual data processing in accordance with embodiments of the present disclosure;
[0020] Fig. 9 illustrates a block diagram of a computing device in which various embodiments of the present disclosure can be implemented.
[0021] Throughout the drawings, the same or similar reference numerals usually refer to the same or similar elements.DETAILED DESCRIPTION
[0022] Principle of the present disclosure will now be described with reference to some embodiments. It is to be understood that these embodiments are described only for the purpose of illustration and help those skilled in the art to understand and implement the present disclosure, without suggesting any limitation as to the scope of the disclosure. The disclosure described herein can be implemented in various manners other than the ones described below.
[0023] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.
[0024] References in the present disclosure to “one embodiment, ” “an embodiment, ” “an example embodiment, ” and the like indicate that the embodiment described may include a particular feature, structure, or characteristic, but it is not necessary that every embodiment includes the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is submitted that it is within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.
[0025] It shall be understood that although the terms “first” and “second” etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and similarly, a second element could be termed a first element, without departing from the scope of example embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0026] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments. As used herein, the singular forms “a” , “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” , “comprising” , “has” , “having” , “includes” and / or “including” , when used herein, specify the presence of stated features, elements, and / or components etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof. Example Environment
[0027] Fig. 1 is a block diagram that illustrates an example visual data coding system 100 that may utilize the techniques of this disclosure. As shown, the visual data coding system 100 may include a source device 110 and a destination device 120. The source device 110 can be also referred to as a visual data encoding device, and the destination device 120 can be also referred to as a visual data decoding device. In operation, the source device 110 can be configured to generate encoded visual data and the destination device 120 can be configured to decode the encoded visual data generated by the source device 110. The source device 110 may include a visual data source 112, a visual data encoder 114, and an input / output (I / O) interface 116.
[0028] The visual data source 112 may include a source such as a visual data capture device. Examples of the visual data capture device include, but are not limited to, an interface to receive visual data from a visual data provider, a computer graphics system for generating visual data, and / or a combination thereof.
[0029] The visual data may comprise one or more pictures of a video or one or more images. The visual data encoder 114 encodes the visual data from the visual data source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the visual data. The bitstream may include coded pictures and associated visual data. The coded picture is a coded representation of a picture. The associated visual data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The encoded visual data may be transmitted directly to destination device 120 via the I / O interface 116 through the network 130A. The encoded visual data may also be stored onto a storage medium / server 130B for access by destination device 120.
[0030] The destination device 120 may include an I / O interface 126, a visual data decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded visual data from the source device 110 or the storage medium / server 130B. The visual data decoder 124 may decode the encoded visual data. The display device 122 may display the decoded visual data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120 which is configured to interface with an external display device.
[0031] The visual data encoder 114 and the visual data decoder 124 may operate according to a visual data coding standard, such as video coding standard or still picture coding standard and other current and / or further standards.
[0032] Some example embodiments of the present disclosure will be described in detailed hereinafter. It should be understood that section headings are used in the present document to facilitate ease of understanding and do not limit the embodiments disclosed in a section to only that section. Furthermore, while certain embodiments are described with reference to Versatile Video Coding or other specific visual data codecs, the disclosed techniques are applicable to other coding technologies also. Furthermore, while some embodiments describe coding steps in detail, it will be understood that corresponding steps decoding that undo the coding will be implemented by a decoder. Furthermore, the term visual data processing encompasses visual data coding or compression, visual data decoding or decompression and visual data transcoding in which visual data are represented from one compressed format into another compressed format or at a different compressed bitrate. 1. Brief Summary
[0033] The present disclosure is related to learning-based end-to-end optimized video coding technologies. Specifically, it is related to flow-based motion estimation and motion compensation. It may be applied to the existing learning-based video coding models like DCVC, DCVC-TCM . It may also be applicable to future video coding standards or video codec powered by AI. 2. Introduction
[0034] Due to the rise and emergence of social media and video applications, it is vital to restore them efficiently, which makes image and video compression become an active area. In recent years, learned video compression attracts lots of attention, most of the learned video compression models are based on predictive coding paradigm. They use an optical flow net or Deformable Convolutional Network (DCN) to predict the motion information between the decoded frame and current frame, use a motion codec to compress the motion information, and use a residual codec or a contextual codec to compress the residuals or contexts. These models are optimized in an end-to-end manner, which makes it flexible to meet a variety of demands, such as perceptual quality. 2.1. Deep Contextual Video Compression
[0035] The learning-based video compression model consists of an independent intra codec, a contextual or encoder EC, a contextual decoder DC, a motion encoder, a motion decoder, and an optical flow net.
[0036] The intra codec is used to compress the key frame (I frame) in each GOP (group of images) . When compress the P frames xt, optical flow net is used to estimate the flow between previous decoded frame and current P frame xt. Flow map is employed as the motion vector.
[0037] At encoder, a motion codec is used to compress the flow map and the reconstructed flow map is employed for consistency between encoding and decoding. The decoded frame is warped to context based on the decoded optical flow When compressing P frame xt, xt is concatenated with context and the result is fed to the contextual codec. The concatenation between xt and is to let the network learn conditional coding.
[0038] At decoder, the key frame or I frame is first decoded using the independent intra codec decoder. The optical flow is decoded using the motion codec decoder. The decoded frame is warped to context based on the decoded optical flow. The inter codec is used to decompress the learned conditional coding information and it is concatenated with the context to obtain the reconstructed frame
[0039] The overall process is as follows: 2.2. Multi-Scale Deep Contextual Video Compression
[0040] The compression of frame xt is taken as an example. Multi-scale features are first extracted from propagated feature The propagated feature is the feature extracted from previous frames.
[0041] Motion vector is adopted to warp the multi-scale features to multi-scale local contexts The is concatenated to the current frame xt, is concatenated the to the mid-feature and the is concatenated to mid-feature The concatenation lets the network learn how to conduct conditional coding by itself. When decoding, multi-scale contexts are also concatenated to recover the frame. The overall process can be formulated as: where EC and ED are the contextual encoder and the contextual decoder. 3. Problems
[0042] The current Flow-based motion estimation and motion compensation has the following problems: 1. Flow-based motion estimation can only handle small motions. The large motions are ignored in the current design of end-to-end video compression. 2. There can exist global redundancy even in the case of small motions. Limited receptive field makes them can only capture local redundancy. 4. Detailed solutions
[0043] The detailed solutions below should be considered as examples to explain general concepts. These solutions should not be interpreted in a narrow way. Furthermore, these solutions can be combined in any manner.
[0044] It is noted that the Predictive (P) frame at time t is xt∈R3×H×W, the decoded frame at time t-1 is where H is the height of each frame and W is the width of each frame. The offset or motion vector between frame xt and frame is denoted as vt∈R2×H×W. The decompressed offset or motion vector is denoted as At the encoder, the vt is compressed and decompressed into by a predefined motion codec. At the decoder side, the offsets / motion vector is required to be decompressed firstly. The compressor of the contextual encoder will down-sample the input frame xt for four times. The features are denoted as The decompressor of the contextual decoder will up-sample the quantized input latent representation four times. The features are denoted as
[0045] Firstly, the key symbol / notation or process in pixel-space end-to-end optimized video coding is defined. 1. The frame and decoded offset are warped to The is concatenated with xt in channel di- mension to get the result The is compressed by a predefined contextual codec. Herein, the is denoted as local context 2. The is employed to warp decoded frame to The contextual decoder from contextual codec decompresses the bit-stream to obtain decoded The is concatenated with in channel dimension to recover the reconstructed frame Herein, the is denoted as local context
[0046] Secondly, the key symbol / notation or process in feature-space end-to-end optimized video coding is defined. 1. At the encoder, the feature is extracted from decoded frame or propagated feature The is employed to warp feature to local context 2. At the decoder, the feature is extracted from decoded frame or propagated feature The is employed to warp feature to local context
[0047] Thirdly, the key symbol / notation or process in multiple-feature-space end-to-end optimized video coding is defined. 1. At the encoder, the multi-scale features are extracted from decoded frame or propagated feature by a predefined neural network. The is employed to warp features to local contexts 2. At the decoder, the multi-scale features are extracted from decoded frame or propagated feature by a predefined neural network. The is employed to warp features to local contexts Propose GMC 1. It is proposed to use a global motion compensation / estimation (GMC) scheme in learning-based video compression. a. In one example, the GMC scheme may be used to generate global motion compensation or extract motion content / information. b. In one example, the GMC scheme may be built / constructed by learning-based method, such as neural network. c. In one example, the GMC scheme may be employed in pixel space or feature space. d. In one example, the GMC scheme may be employed for single-scale features or multi-scale features in feature space. Propose LMC 2. It is proposed to use a local motion compensation / estimation (LMC) scheme in learning-based video compression. a. In one example, the LMC scheme may be used to generate local motion compensation or extract motion content / information. b. In one example, the LMC scheme may be built / constructed by learning-based method, such as neural network. c. In one example, the LMC scheme may be employed in pixel space or feature space. d. In one example, the LMC scheme may be employed for single-scale features or multi-scale features in feature space. Propose LGMC 3. It is proposed to use a joint the local and global motion compensation / estimation scheme (LGMC) in learning-based video compression. a. In one example, the LGMC scheme may be used to generate joint local and global motion compensation or extract motion content / information. b. In one example, the LGMC scheme may be built / constructed by learning-based method, such as neural network. c. In one example, the LGMC scheme may be combined by an LMC scheme and an GMC scheme. d. In one example, the LGMC scheme may be employed in pixel space or feature space. e. In one example, the LGMC scheme may be employed for single-scale feature or multi-scale features in feature space. How to design the LMC. 4. It is proposed to use Optical Flow network (denoted as FlowNet) to construct LMC. a. In one example, the FlowNet is employed to estimate the offsets or motion vector vt between frame xt and frame i. In one example, the LMC is employed in pixel space. ii. In one example, the LMC is employed in feature space. iii. In one example, the LMC is employed for multi-scale features. 5. It is proposed to use Deformable Neural Network (DCN) to construct LMC. a. In one example, the DCN is employed to estimate the offsets or motion vector vt between frame xt and frame i. In one example, the LMC is employed in pixel space. ii. In one example, the LMC is employed in feature space. iii. In one example, the LMC is employed for multi-scale features. How to design the GMC. 6. It is proposed to use cross attention network to construct GMC. a. In one example, the GMC is employed in pixel space. i. In one example, at encoder side, the cross attention is computed via a predefined cross attention method between and xt when conducting global motion compensation. The process is formulated as: 1. or 2. or 3. ii. In one example, at encoder side, the global motion compensation process is formu- lated as follows: 1. 2. 3. 4. iii. In one example, at decoder side, the cross attention is computed via a predefined cross attention method between and The process is formulated as: 1. or 2. or 3. iv. In one example, at decoder side, the global motion compensation process is formu- lated as follows: 1. 2. 3. 4. 5. b. In one example, the GMC is employed in feature space. i. In one example, at encoder side, the cross attention is computed via a predefined cross attention method between feature and xt. is extracted from or propagated feature The process is formulated as: 1. or 2. or 3. ii. In one example, at encoder side, the global motion compensation process is formu- lated as follows: 1. 2. 3. 4. iii. In one example, at decoder side, the cross attention is computed via a predefined cross attention method between feature and is extracted from or propagated feature The process is formulated as: 1. or 2. or 3. iv. In one example, at decoder side, the global motion compensation process is formu- lated as follows: 1. 2. 3. 4. 5. c. In one example, GMC is employed for multi-scale features. i. In one example, at encoder side, the cross attention is computed via a predefined cross attention method between and xt, and and The process is formulated as follows: 1. or or 2. or or 3. or or ii. In one example, at encoder side, the global motion compensation process is formu- lated as follows: 1. 2. 3. 4. iii. In one example, at decoder side, the cross attention is computed via a predefined cross attention method between and and and and The process is formulated as follows: 1. or or 2. or or 3. or or iv. In one example, at decoder side, the motion compensation process is formulated as follows: 1. 2. 3. 4. 5. d. In one example, the cross attention G is computed via a predefined cross attention method between a transform FA (*) of input signal A and a transform FB (*) of input signal B. i. In one example, the process is formulated as follows: 1. G=softmax (FA (A) (FB (B) ) T) FA (B) or 2. G=softmax (FA (A) ) softmax ( (FA (B) ) T) FB (B) 3. G=softmax (FA (A) ) * [softmax ( (FA (B) ) T) *FB (B) ] . 4. G=softmax (FA (A) ) * [softmax (FA (B) ) T*FB (B) ] . ii. In one example, FA (*) and FB (*) may be removed: 1. G=softmax (A* BT) *B or 2. G=softmax (A) *softmax (BT) *B 3. G=softmax (A) * [softmax (BT) *B] 4. G=softmax (A) * [softmax (B) T*B] iii. In one example, FA (*) and FB (*) may be dependent on the input signal. iv. In one example, FA (*) may be same as FB (*) . v. In one example, the transform is conducted by a network such as Depth Residual Bottleneck (DepthRB) . 1. In one example, the DepthRB is designed as follows: a. The DepthRB consists of a 1×1 convolutional layer, a 3×3 convo- lutional layer and a 1×1 convolutional layer. The output of DepthRB is the summation of the input and the output for the sec-ond 1×1 convolutional layer. vi. In one example, a network is applied to cross attention to get the final cross attention. 1. In one example, the network consists of a convolution layer and a DepthRB block. 2. In one example, G = DepthRB (conv5x5 (G) ) . vii. In one example, the calculation method of cross attention is suitable for any bullets mentioned above. 1. In one example, A may be xt or and B may be 2. In one example, A may be xt or and B may be 3. In one example, A may be or and B may be 4. In one example, A may be or and B may be How to design the LGMC 7. It is proposed to combine LMC and GMC to get LGMC. a. In one example, combination is concatenation in channel dimension. b. In one example, LGMC is employed in pixel or feature space. i. In one example, the is obtained via GMC and is obtained via LMC. ii. In one example, the are concatenated in the channel dimension. 1. In one example, at encoder side, the are concatenated with xt in the channel dimension. The process is formulated as follows: a. b. c. d. 2. In one example, at decoder side, the are concatenated with in the channel dimension. The process is formulated as follows: a. b. c. d. e. iii. In one example, the are fused by a network. 1. In one example, at encoder side, the are fused with xt. 2. In one example, at decoder side, the are fused with c. In one example, LGMC is employed for multi-scale features space. i. In one example, the are obtained via GMC and are obtained via LMC. ii. In one example, the are concatenated and / or are concatenated and / or are concatenated in the channel dimension. 1. In one example, at encoder side, the and are concatenated with xt in the channel dimension. and are concatenated the to the mid-feature in the channel dimension and the and are concatenated to mid-feature in the channel dimension. The process is formulated as follows: a. b. c. d. 2. In one example, at decoder side, the and are concatenated with in the channel dimension. and are concatenated the to the mid-feature in the channel dimension and the and are concatenated to mid-feature in the channel dimension. The process is formulated as follows: a. b. c. d. e. iii. In one example, the are fused and / or are fused and / or are fused by a network. 1. In one example, at encoder side, the are fused with xt, and / or the are fused with and / or the are fused with 2. In one example, at decoder side, the are fused with and / or the are fused with and / or the are fused with 5. Embodiment 5.1. Embodiment 1
[0048] The paradigm of the proposed LVC-LGMC is illustrated in Fig. 2. We adopt temporal propagated multi-scale features for local compensation. Fig. 2 shows an overall framework of the proposed method. EC is the contextual encoder, and DC is the contextual decoder. EM is the MV encoder, and DMis the MV decoder. LGMC is the proposed joint local and global motion compensation module. Fig. 3 is an illustration of the joint local and global motion compensation module (LGMC) at encoder side. Fig. 4 is an illustration of the joint local and global motion compensation module (LGMC) at decoder side. 5.1.1. Flow-based Local Compensation
[0049] When compressing the t-th frame xt, first, we conduct the local compensation. In particular, multi-scale features are extracted from propagated feature Motion vector is employed to warp the multi-scale features to multi-scale local contexts The are concatenated with the current frame xt, middle-feature and middle-feature respectively. As such, the network could well understand how to conduct conditional coding. During decoding, the multi-scale contexts are also concatenated to recover the frame. The overall process can be formulated as: 5.1.2. Attention-based Global Compensation
[0050] The compression of frame xt is taken as an example. Given multi-scales features and current frame xt, mid features and are taken as an example, where C is the channel number, and L = HW, H is the height and W is the width. and are first fed into an embedding layer. The vanilla approach adopts cross attention. The process is formulated as:
[0051] Because it can be treated as the similarity metric. It computes the simi-larity between a symbol and all other symbols, which makes it capture global dependency. The is concatenated with to let the network learn to conduct conditional coding on the basis of global de-pendency. The overall process is as follows:
[0052] The computational complexity of Vanilla Approach is O (L2) , which makes the vanilla approach cannot be employed for high-resolution video coding. The quadratic is caused by the softmax operation, which specifies the order of matrix multiplication. To solve the quadratic complexity, the softmax operation on in row and the softmax operation on in column is adopted.
[0053] Since it can be treated as a similarity metric. Larger values mean more similarity. is computed first in practice, which makes the computational complexity be O (C2L) . The overall process is as follows: 5.1.3. Mixed Flow-Attention for Joint Local and Global Motion Compensation
[0054] Flow is adopted to obtain local contexts and attention is adopted to obtain global contexts Local context, global context with current frames or mid feature are concatenated to let the network learn to use local and global contexts for conditional coding. The overall process is 5.2. Embodiment 2
[0055] The paradigm of the proposed LVC-LGMC is illustrated in Fig. 5. We adopt temporal propagated multi-scale features for local compensation. Fig. 5 shows overall framework of the proposed method. EC is the contextual encoder, and DC is the contextual decoder. EM is the MV encoder, and DMis the MV decoder. LGMC is the proposed joint local and global motion compensation module. Fig. 6 is an illustration of the joint local and global motion compensation module (LGMC) at encoder side. Fig. 7 is an illustration of the joint local and global motion compensation module (LGMC) at decoder side. 5.2.1. Flow-based Local Compensation
[0056] When compressing the t-th frame xt, first, we conduct the local compensation. In particular, multi-scale features are extracted from propagated feature Motion vector is employed to warp the multi-scale features to multi-scale local contexts The are concatenated with the current frame xt, middle-feature and middle-feature respectively. As such, the network could well understand how to conduct conditional coding. During decoding, the multi-scale contexts are also concatenated to recover the frame. The overall process can be formulated as: 5.2.2. Attention-based Global Compensation
[0057] The compression of frame xt is taken as an example. Given multi-scales features and current frame xt, mid features and are taken as an example, where C is the channel number, and L = HW, H is the height and W is the width. and are first fed into an embedding layer. The vanilla approach adopts cross attention. The process is formulated as:
[0058] Because it can be treated as the similarity metric. It computes the simi-larity between a symbol and all other symbols, which makes it capture global dependency. The is concatenated with to let the network learn to conduct conditional coding on the basis of global de-pendency. The overall process is as follows:
[0059] The computational complexity of Vanilla Approach is O (L2) , which makes the vanilla approach cannot be employed for high-resolution video coding. The quadratic is caused by the softmax operation, which specifies the order of matrix multiplication. To solve the quadratic complexity, the softmax operation on in row and the softmax operation on in column is adopted.
[0060] Since it can be treated as a similarity metric. Larger values mean more similarity. is computed first in practice, which makes the computational complexity be O (C2L) . The overall process is as follows: 5.2.3. Mixed Flow-Attention for Joint Local and Global Motion Compensation
[0061] Flow is adopted to obtain local contexts and attention is adopted to obtain global contexts Local context, global context with current frames or mid feature are concatenated to let the network learn to use local and global contexts for conditional coding. The overall process is
[0062] Fig. 8 illustrates a flowchart of a method 800 for visual data processing in accordance with embodiments of the present disclosure. The method 800 is implemented during a conversion between visual data and a bitstream of the visual data.
[0063] At block 810, for a conversion between visual data and a bitstream of the visual data, at least one of a motion compensation scheme or a motion estimation scheme is constructed based on at least one of: an optical flow network, a deformable neural network, or a cross attention network. In some embodiments, the at least one of the motion compensation scheme or the motion estimation scheme may include at least one of: a global motion compensation (GMC) scheme, a global motion estimation scheme, a local motion compensation (LMC) scheme, a local motion estimation scheme, a joint local and global motion compensation (LGMC) scheme, or a joint local and global motion estimation scheme.
[0064] At block 820, the visual data is processed by applying the at least one of the motion compensation scheme or the motion estimation scheme in a learning-based video compression.
[0065] At block 830, the conversion is performed based on the processed visual data. In some embodiments, the conversion may include encoding the visual data into the bitstream. In some other embodiments, the conversion may include decoding the visual data from the bitstream.
[0066] The method 800 enables the motion compensation scheme and / or the motion estimation scheme to be constructed. In addition, the visual data is processed by applying the motion compensation scheme and / or the motion estimation scheme. In this way, the method 800 can advantageously improve the coding quality and coding efficiency.
[0067] In some embodiments, applying the at least one of the motion compensation scheme or the motion estimation scheme in the learning-based video compression may include: applying a global motion compensation (GMC) scheme and / or a global motion estimation scheme in the learning-based video compression. For example, the global motion compensation scheme and / or the global motion estimation scheme may be used to generate global motion compensation. Alternatively, the global motion compensation scheme and / or the global motion estimation scheme may be used to derive at least one of motion content or motion information.
[0068] In some embodiments, applying the at least one of the motion compensation scheme or the motion estimation scheme in the learning-based video compression may include: applying a local motion compensation (LMC) scheme and / or a local motion estimation scheme in the learning-based video compression. For example, the local motion compensation scheme and / or the local motion estimation scheme may be used to generate a local motion compensation. Alternatively, the local motion compensation scheme and / or the local motion estimation scheme may be used to derive at least one of motion content or motion information.
[0069] In some embodiments, applying the at least one of the motion compensation scheme or the motion estimation scheme in the learning-based video compression may include: applying a joint local and global motion compensation (LGMC) scheme and / or a joint local and global motion estimation scheme in the learning-based video compression. For example, the joint local and global motion compensation scheme and / or the joint local and global motion estimation scheme may be used to generate joint local and global motion compensation. Alternatively, the joint local and global motion compensation scheme and / or the joint local and global motion estimation scheme may be used to derive at least one of motion content or motion information. In some embodiments, the joint local and global motion compensation scheme may be combined by a global motion compensation (GMC) scheme and a local motion compensation (LMC) scheme. Alternatively, the joint local and global motion estimation scheme may be combined by a global motion estimation scheme and a local motion estimation scheme.
[0070] In some embodiments, the at least one of the motion compensation scheme or the motion estimation scheme may be constructed based on a learning-based approach. For example, the learning-based approach may include a neural network. In some embodiments, the at least one of the motion compensation scheme or the motion estimation scheme may be applied to at least one of: a pixel space or a feature space. In some other embodiments, the at least one of the motion compensation scheme or the motion estimation scheme may be applied to at least one of the following in a feature space: a single-scale feature or a multi-scale feature.
[0071] In some embodiments, constructing the at least one of the motion compensation scheme or the motion estimation scheme based on the at least one of: the optical flow network, the deformable neural network, or the cross attention network may include: constructing a local motion compensation (LMC) scheme and / or a local motion estimation scheme by using the optical flow network. In some embodiments, the optical flow network may be applied to determine at least one of an offset or motion vector between a first frame and a second frame. In this case, the first frame may include a predictive frame at a time point, and the second frame may include a decoded frame at the time point minus a time interval. In some embodiments, the local motion compensation scheme and / or the local motion estimation scheme may be applied in at least one of: a pixel space or a feature space. In some other embodiments, the local motion compensation scheme and / or the local motion estimation scheme may be applied in a multi-scale feature.
[0072] In some embodiments, constructing the at least one of the motion compensation scheme or the motion estimation scheme based on the at least one of: the optical flow network, the deformable neural network, or the cross attention network may include: constructing a local motion compensation (LMC) scheme and / or a local motion estimation scheme by using the deformable neural network. In some embodiments, the deformable neural network may be applied to determine at least one of an offset or motion vector between a first frame and a second frame. In this case, the first frame may include a predictive frame at a time point, and the second frame may include a decoded frame at the time point minus a time interval. In some embodiments, the local motion compensation scheme and / or the local motion estimation scheme may be applied in at least one of: a pixel space or a feature space. In some other embodiments, the local motion compensation scheme and / or the local motion estimation scheme may be applied in a multi-scale feature.
[0073] In some embodiments, constructing the at least one of the motion compensation scheme or the motion estimation scheme based on the at least one of: the optical flow network, the deformable neural network, or the cross attention network may include: constructing a global motion compensation (GMC) scheme and / or a global motion estimation scheme by using the cross attention network. In some embodiments, the global motion compensation (GMC) scheme and / or the global motion estimation scheme may be applied in a pixel space.
[0074] In some embodiments, if the global motion compensation is performed, may be derived at encoder side by a predetermined cross attention approach between and xt. In this case, represents a cross attention, xt represents a predictive frame at a time point, and represents a decoded frame at the time point minus 1. In some embodiments, the cross attention may be derived by the following: In some other embodiments, the cross attention may be derived by the following: Alternatively, the cross attention may be derived by the following:
[0075] In some embodiments, a global motion compensation process at encoder side may be performed as the following: In this case, represents cross attention, xt represents a predictive frame at a time point, represents a first down-sample feature, represents a second down-sample feature, represents a third down-sample feature, and represents a fourth down-sample feature.
[0076] In some embodiments, may be derived at decoder side by another predetermined cross attention approach between and In this case, represents another cross attention, represents a decoded frame at a time point minus 1, and represents a reconstructed frame. In some embodiments, the other cross attention may be derived by the following: In some other embodiments, the other cross attention may be derived by the following: Alternatively, the other cross attention may be derived by the following:
[0077] In some embodiments, a global motion compensation process at decoder side may be performed as the following: In this case, represents a first up-sample feature, represents a second up-sample feature, represents a third up-sample feature, represents a fourth up-sample feature, represents a reconstructed frame, and represents another cross attention.
[0078] In some embodiments, the global motion compensation (GMC) scheme and / or the global motion estimation scheme may be applied to a feature space. In some embodiments, may be derived at encoder side by a predetermined cross attention approach between and xt. In this case, represents a cross attention, xt represents a predictive frame at a time point, and represents a feature derived from or a propagated feature and represents a decoded frame at the time point minus 1. In some embodiments, the cross attention may be derived by the following: In some other embodiments, the cross attention may be derived by the following: Alternatively, the cross attention may be derived by the following:
[0079] In some embodiments, a global motion compensation process at encoder side may be performed as the following: In this case, represents a cross attention, xt represents a predictive frame at a time point, represents a first down-sample feature, represents a second down-sample feature, represents a third down-sample feature, and represents a fourth down-sample feature.
[0080] In some embodiments, may be derived at decoder side by another predetermined cross attention approach between and In this case, represents another cross attention, represents a reconstructed frame, and represents a feature derived from or a propagated feature and represents a decoded frame at a time point minus 1. In some embodiments, the other cross attention may be derived by the following: In some other embodiments, the other cross attention may be derived by the following: Alternatively, the other cross attention may be derived by the following:
[0081] In some embodiments, a global motion compensation process at decoder side may be performed as the following: In this case, represents a first up-sample feature, represents a second up-sample feature, represents a third up-sample feature, represents a fourth up-sample feature, represents a reconstructed frame, and represents another cross attention.
[0082] In some embodiments, the global motion compensation (GMC) scheme and / or the global motion estimation scheme may be applied to multi-scale features. In some embodiments, cross attention may be derived at encoder side by a predetermined cross attention approach between and xt, and and In this case, xt represents a predictive frame at a time point, represents a first down-sample feature, represents a second down-sample feature, and and represent multi-scale features derived from or a propagated feature and represents a decoded frame at the time point minus 1.
[0083] In some embodiments, the cross attention may be derived by the following: or or In this case, represents a first cross attention. In some other embodiments, the cross attention may be derived by the following: or or In this case, represents a second cross attention. Alternatively, the cross attention may be derived by the following: or or In this case, represents a third cross attention.
[0084] In some embodiments, a global motion compensation process at encoder side may be performed as the following: In this case, represents a first cross attention, represents a second cross attention, represents a third cross attention, xt represents a predictive frame at a time point, represents a first down-sample feature, represents a second down-sample feature, represents a third down-sample feature, and represents a fourth down-sample feature.
[0085] In some embodiments, another cross attention may be derived at decoder side by another predetermined cross attention approach between and and and In this case, represents a reconstructed frame, represents a fourth up-sample feature, represents a third up-sample feature, and and represent multi-scale features derived from or a propagated feature and represents a decoded frame at a time point minus 1.
[0086] In some embodiments, the other cross attention may be derived by one of the following: or or In this case, represents a fourth cross attention. In some other embodiments, the other cross attention may be derived by one of the following: or or In this case, represents a fifth cross attention. Alternatively, the other cross attention may be derived by one of the following: or or In this case, represents a sixth cross attention.
[0087] In some embodiments, a motion compensation process at decoder side may be performed as the following: In this case, represents a first up-sample feature, represents a second up-sample feature, represents a third up-sample feature, represents a fourth up-sample feature, represents a reconstructed frame, represents a fourth cross attention, represents a sixth cross attention, and represents a sixth cross attention.
[0088] In some embodiments, G may be derived by a predetermined cross attention approach between FA (A) and FB (B) . In this case, G represents a cross attention, FA (A) represents a first transform of A, FB (B) represents a second transform of B, and A represents a first input signal, B represents a second input signal. In some embodiments, G may be derived by one of the following: G=softmax (FA (A) (FB (B) ) T) FA (B) , or G=softmax (FA (A) ) softmax ( (FA (B) ) T) FB (B) , or G=softmax (FA (A) ) * [softmax ( (FA (B) ) T) *FB (B) ] , or G=softmax (FA (A) ) * [softmax (FA (B) ) T*FB (B) ] . In some other embodiments, G may be derived by one of the following: G=softmax (A* BT) *B, or G=softmax (A) *softmax (BT) *B, or G=softmax (A) * [softmax (BT) *B] , or G=softmax (A) * [softmax (B) T*B] . In some embodiments, the first transform and / or the second transform may be based on an input signal. In some embodiments, the first transform may be same as the second transform.
[0089] In some embodiments, the first transform and / or the second transform may be performed by a network, such as a depth residual bottleneck (DepthRB) . For example, the DepthRB may include a first 1×1 convolutional layer, a 3×3 convolutional layer, and a second 1×1 convolutional layer, and an output of the DepthRB is a sum of an input and an output of the second 1×1 convolutional layer.
[0090] In some embodiments, a network may be applied to the cross attention to determine a final cross attention. In some embodiments, the network may include a convolution layer and a DepthRB block. In some other embodiments, G may equal to DepthRB (conv5x5 (G) ) . In this case, DepthRB represents a depth residual bottleneck and conv5x5 represents a 5×5 convolutional layer.
[0091] In some embodiments, the first input signal may include xt or and the second input signal may include In this case, xt represents a predictive frame at a time point, represents a reconstructed frame, and represents a decoded frame at the time point minus 1. In some other embodiments, the first input signal may include xt or and the second input signal may include In this case, xt represents a predictive frame at a time point, represents a reconstructed frame, and represents a first multi-scale features derived from or a propagated feature and represents a decoded frame at the time point minus 1. In some embodiments, the first input signal may include or and the second input signal may include In this case, represents a first down-sample feature, represents a fourth up-sample feature, and represent a second multi-scale feature derived from or a propagated feature and represents a decoded frame at a time point minus 1. Alternatively, the first input signal may include or and the second input signal may include In this case, represents a second down-sample feature, represents a third up-sample feature, and represent a third multi-scale feature derived from or a propagated feature and represents a decoded frame at a time point minus 1.
[0092] In some embodiments, a joint local and global motion compensation (LGMC) scheme may be obtained by combining a global motion compensation (GMC) scheme and a local motion compensation (LMC) scheme. Alternatively, a joint local and global motion estimation scheme may be obtained by combining a global motion estimation scheme and a local motion estimation scheme. For example, the combination may include a concatenation in a channel dimension.
[0093] In some embodiments, the joint local and global motion compensation (LGMC) scheme and / or the joint local and global motion estimation scheme may be applied to at least one of a pixel or a feature space. In some embodiments, may be obtained by the global motion compensation (GMC) scheme and / or the global motion estimation scheme, and may be obtained by the local motion compensation (LMC) scheme and / or the local motion estimation scheme. In this case, represents a first cross attention, and represents a first local context.
[0094] In some embodiments, and may be concatenated in a channel dimension. In this case, represents a first cross attention, and represents a first local context. In some embodiments, at encoder side, and may be concatenated with xt in the channel dimension, which is performed as the following: In this case, xt represents a predictive frame at a time point, represents a first down-sample feature, represents a second down-sample feature, represents a third down-sample feature, and represents a fourth down-sample feature.
[0095] In some embodiments, at decoder side, and may be concatenated with in the channel dimension, which is performed as the following: In this case, represents a first up-sample feature, represents a second up-sample feature, represents a third up-sample feature, represents a fourth up-sample feature, represents a reconstructed frame.
[0096] In some embodiments, and may be fused by a network. In this case, represents a first cross attention, and represents a first local context. In some embodiments, at encoder side, and may be fused with xt. In this case, xt represents a predictive frame at a time point. Alternatively, at decoder side, and may be fused with In this case, represents a reconstructed frame.
[0097] In some embodiments, the joint local and global motion compensation (LGMC) scheme and / or the joint local and global motion estimation scheme may be applied in a multi-scale features space. In some embodiments, and may be obtained by the global motion compensation (GMC) scheme and / or the global motion estimation scheme, and and are obtained by the local motion compensation (LMC) scheme and / or the local motion estimation scheme. In this case, represents a first cross attention, represents a second cross attention, represents a third cross attention, represents a first local context, represents a second local context, and represents a third local context.
[0098] In some embodiments, and may be concatenated in a channel dimension, and / or and may be concatenated in the channel dimension, and / or and may be concatenated in the channel dimension. In this case, represents a first cross attention, represents a second cross attention, represents a third cross attention, represents a first local context, represents a second local context, and represents a third local context.
[0099] In some embodiments, at encoder side, and may be concatenated with xt in the channel dimension, and may be concatenated with a mid-feature in the channel dimension, and and may be concatenated with another mid-feature in the channel dimension, which is performed as the following: In this case, xt represents a predictive frame at a time point, represents a first down-sample feature, represents a second down-sample feature, represents a third down-sample feature, and represents a fourth down-sample feature.
[0100] In some embodiments, at decoder side, and may be concatenated with in the channel dimension, and may be concatenated with a mid-feature in the channel dimension, and and may be concatenated with another mid-feature in the channel dimension, which is performed as the following: In this case, represents a first up-sample feature, represents a second up-sample feature, represents a third up-sample feature, represents a fourth up-sample feature, represents a reconstructed frame, represents a fourth cross attention, represents a fifth cross attention, and represents a sixth cross attention.
[0101] In some embodiments, and may be fused by a network, and / or and may be fused by the network, and / or and may be fused by the network. In this case, represents a first cross attention, represents a second cross attention, represents a third cross attention, represents a first local context, represents a second local context, and represents a third local context. In some embodiments, at encoder side, and may be fused with xt, and / or and may be fused with and / or and may be fused with In this case, represents a first down-sample feature, represents a second down-sample feature, and xt represents a predictive frame at a time point. In some other embodiments, at decoder side, and may be fused with and / or and may be fused with and / or and may be fused with In this case, represents a third up-sample feature, represents a fourth up-sample feature, and represents a reconstructed frame.
[0102] According to further embodiments of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of visual data which is generated by a method performed by an apparatus for visual data processing. The method comprises: constructing at least one of a motion compensation scheme or a motion estimation scheme based on at least one of: an optical flow network, a deformable neural network, or a cross attention network; processing the visual data by applying the at least one of the motion compensation scheme or the motion estimation scheme in a learning-based video compression; and generating the bitstream based on the processed visual data.
[0103] According to still further embodiments of the present disclosure, a method for storing bitstream of visual data is provided. The method comprises: constructing at least one of a motion compensation scheme or a motion estimation scheme based on at least one of: an optical flow network, a deformable neural network, or a cross attention network; processing the visual data by applying the at least one of the motion compensation scheme or the motion estimation scheme in a learning-based video compression; generating the bitstream based on the processed visual data; storing the bitstream in a non-transitory computer-readable recording medium.
[0104] Implementations of the present disclosure can be described in view of the following clauses, the features of which can be combined in any reasonable manner.
[0105] Clause 1. A method of visual data processing, comprising: constructing, for a conversion between visual data and a bitstream of the visual data, at least one of a motion compensation scheme or a motion estimation scheme based on at least one of: an optical flow network, a deformable neural network, or a cross attention network; processing the visual data by applying the at least one of the motion compensation scheme or the motion estimation scheme in a learning-based video compression; and performing the conversion based on the processed visual data.
[0106] Clause 2. The method of clause 1, wherein applying the at least one of the motion compensation scheme or the motion estimation scheme in the learning-based video compression comprises: applying a global motion compensation (GMC) scheme and / or a global motion estimation scheme in the learning-based video compression.
[0107] Clause 3. The method of clause 3, wherein the global motion compensation scheme and / or the global motion estimation scheme is used to generate global motion compensation, or wherein the global motion compensation scheme and / or the global motion estimation scheme is used to derive at least one of motion content or motion information.
[0108] Clause 4. The method of clause 1, wherein applying the at least one of the motion compensation scheme or the motion estimation scheme in the learning-based video compression comprises: applying a local motion compensation (LMC) scheme and / or a local motion estimation scheme in the learning-based video compression.
[0109] Clause 5. The method of clause 4, wherein the local motion compensation scheme and / or the local motion estimation scheme is used to generate a local motion compensation, or wherein the local motion compensation scheme and / or the local motion estimation scheme is used to derive at least one of motion content or motion information.
[0110] Clause 6. The method of clause 1, wherein applying the at least one of the motion compensation scheme or the motion estimation scheme in the learning-based video compression comprises: applying a joint local and global motion compensation (LGMC) scheme and / or a joint local and global motion estimation scheme in the learning-based video compression.
[0111] Clause 7. The method of clause 6, wherein the joint local and global motion compensation scheme and / or the joint local and global motion estimation scheme is used to generate joint local and global motion compensation, or wherein the joint local and global motion compensation scheme and / or the joint local and global motion estimation scheme is used to derive at least one of motion content or motion information.
[0112] Clause 8. The method of clause 6, wherein the joint local and global motion compensation scheme is combined by a global motion compensation (GMC) scheme and a local motion compensation (LMC) scheme, or wherein the joint local and global motion estimation scheme is combined by a global motion estimation scheme and a local motion estimation scheme.
[0113] Clause 9. The method of clause 1, wherein the at least one of the motion compensation scheme or the motion estimation scheme is constructed based on a learning-based approach.
[0114] Clause 10. The method of clause 9, wherein the learning-based approach comprises a neural network.
[0115] Clause 11. The method of clause 1, wherein the at least one of the motion compensation scheme or the motion estimation scheme is applied to at least one of: a pixel space or a feature space.
[0116] Clause 12. The method of clause 1, wherein the at least one of the motion compensation scheme or the motion estimation scheme is applied to at least one of the following in a feature space: a single-scale feature or a multi-scale feature.
[0117] Clause 13. The method of any of clauses 1 to 12, wherein the at least one of the motion compensation scheme or the motion estimation scheme comprises at least one of: a global motion compensation (GMC) scheme, a global motion estimation scheme, a local motion compensation (LMC) scheme, a local motion estimation scheme, a joint local and global motion compensation (LGMC) scheme, or a joint local and global motion estimation scheme.
[0118] Clause 14. The method of clause 1, wherein constructing the at least one of the motion compensation scheme or the motion estimation scheme based on the at least one of: the optical flow network, the deformable neural network, or the cross attention network comprises: constructing a local motion compensation (LMC) scheme and / or a local motion estimation scheme by using the optical flow network.
[0119] Clause 15. The method of clause 14, wherein the optical flow network is applied to determine at least one of an offset or motion vector between a first frame and a second frame, wherein the first frame comprises a predictive frame at a time point, and the second frame comprises a decoded frame at the time point minus a time interval.
[0120] Clause 16. The method of clause 15, wherein the local motion compensation scheme and / or the local motion estimation scheme is applied in at least one of: a pixel space or a feature space.
[0121] Clause 17. The method of clause 15, wherein the local motion compensation scheme and / or the local motion estimation scheme is applied in a multi-scale feature.
[0122] Clause 18. The method of clause 1, wherein constructing the at least one of the motion compensation scheme or the motion estimation scheme based on the at least one of: the optical flow network, the deformable neural network, or the cross attention network comprises: constructing a local motion compensation (LMC) scheme and / or a local motion estimation scheme by using the deformable neural network.
[0123] Clause 19. The method of clause 18, wherein the deformable neural network is applied to determine at least one of an offset or motion vector between a first frame and a second frame, wherein the first frame comprises a predictive frame at a time point, and the second frame comprises a decoded frame at the time point minus a time interval.
[0124] Clause 20. The method of clause 19, wherein the local motion compensation scheme and / or the local motion estimation scheme is applied in at least one of: a pixel space or a feature space.
[0125] Clause 21. The method of clause 19, wherein the local motion compensation scheme and / or the local motion estimation scheme is applied in a multi-scale feature.
[0126] Clause 22. The method of clause 1, wherein constructing the at least one of the motion compensation scheme or the motion estimation scheme based on the at least one of: the optical flow network, the deformable neural network, or the cross attention network comprises: constructing a global motion compensation (GMC) scheme and / or a global motion estimation scheme by using the cross attention network.
[0127] Clause 23. The method of clause 22, wherein the global motion compensation (GMC) scheme and / or the global motion estimation scheme is applied in a pixel space.
[0128] Clause 24. The method of clause 23, wherein if the global motion compensation is performed, is derived at encoder side by a predetermined cross attention approach between and xt, wherein represents a cross attention, xt represents a predictive frame at a time point, and represents a decoded frame at the time point minus 1.
[0129] Clause 25. The method of clause 24, wherein the cross attention is derived by the following:
[0130] Clause 26. The method of clause 24, wherein the cross attention is derived by the following:
[0131] Clause 27. The method of clause 24, wherein the cross attention is derived by the following:
[0132] Clause 28. The method of clause 23, wherein a global motion compensation process at encoder side is performed as the following: wherein represents cross attention, xt represents a predictive frame at a time point, represents a first down-sample feature, represents a second down-sample feature, represents a third down-sample feature, and represents a fourth down-sample feature.
[0133] Clause 29. The method of clause 23, wherein is derived at decoder side by another predetermined cross attention approach between and wherein represents another cross attention, represents a decoded frame at a time point minus 1, and represents a reconstructed frame.
[0134] Clause 30. The method of clause 29, wherein the other cross attention is derived by the following:
[0135] Clause 31. The method of clause 29, wherein the other cross attention is derived by the following:
[0136] Clause 32. The method of clause 29, wherein the other cross attention is derived by the following:
[0137] Clause 33. The method of clause 23, wherein a global motion compensation process at decoder side is performed as the following: wherein represents a first up-sample feature, represents a second up-sample feature, represents a third up-sample feature, represents a fourth up-sample feature, represents a reconstructed frame, and represents another cross attention.
[0138] Clause 34. The method of clause 22, wherein the global motion compensation (GMC) scheme and / or the global motion estimation scheme is applied to a feature space.
[0139] Clause 35. The method of clause 34, wherein is derived at encoder side by a predetermined cross attention approach between and xt, wherein represents a cross attention, xt represents a predictive frame at a time point, and represents a feature derived from or a propagated feature wherein represents a decoded frame at the time point minus 1.
[0140] Clause 36. The method of clause 35, wherein the cross attention is derived by the following:
[0141] Clause 37. The method of clause 35, wherein the cross attention is derived by the following:
[0142] Clause 38. The method of clause 35, wherein the cross attention is derived by the following:
[0143] Clause 39. The method of clause 34, wherein a global motion compensation process at encoder side is performed as the following: wherein represents a cross attention, xt represents a predictive frame at a time point, represents a first down-sample feature, represents a second down-sample feature, represents a third down-sample feature, and represents a fourth down-sample feature.
[0144] Clause 40. The method of clause 34, wherein is derived at decoder side by another predetermined cross attention approach between and wherein represents another cross attention, represents a reconstructed frame, and represents a feature derived from or a propagated feature wherein represents a decoded frame at a time point minus 1.
[0145] Clause 41. The method of clause 40, wherein the other cross attention is derived by the following:
[0146] Clause 42. The method of clause 40, wherein the other cross attention is derived by the following:
[0147] Clause 43. The method of clause 40, wherein the other cross attention is derived by the following:
[0148] Clause 44. The method of clause 34, wherein a global motion compensation process at decoder side is performed as the following: wherein represents a first up-sample feature, represents a second up-sample feature, represents a third up-sample feature, represents a fourth up-sample feature, represents a reconstructed frame, and represents another cross attention.
[0149] Clause 45. The method of clause 22, wherein the global motion compensation (GMC) scheme and / or the global motion estimation scheme is applied to multi-scale features.
[0150] Clause 46. The method of clause 45, wherein cross attention is derived at encoder side by a predetermined cross attention approach between and xt, and and wherein xt represents a predictive frame at a time point, represents a first down-sample feature, represents a second down-sample feature, and and represent multi-scale features derived from or a propagated feature wherein represents a decoded frame at the time point minus 1.
[0151] Clause 47. The method of clause 46, wherein the cross attention is derived by the following: or or wherein represents a first cross attention.
[0152] Clause 48. The method of clause 46, wherein the cross attention is derived by the following: or or wherein represents a second cross attention.
[0153] Clause 49. The method of clause 46, wherein the cross attention is derived by the following: or or wherein represents a third cross attention.
[0154] Clause 50. The method of clause 45, wherein a global motion compensation process at encoder side is performed as the following: wherein represents a first cross attention, represents a second cross attention, represents a third cross attention, xt represents a predictive frame at a time point, represents a first down-sample feature, represents a second down-sample feature, represents a third down-sample feature, and represents a fourth down-sample feature.
[0155] Clause 51. The method of clause 45, wherein another cross attention is derived at decoder side by another predetermined cross attention approach between and and and wherein represents a reconstructed frame, represents a fourth up-sample feature, represents a third up-sample feature, and and represent multi-scale features derived from or a propagated feature wherein represents a decoded frame at a time point minus 1.
[0156] Clause 52. The method of clause 51, wherein the other cross attention is derived by one of the following: or or wherein represents a fourth cross attention.
[0157] Clause 53. The method of clause 51, wherein the other cross attention is derived by one of the following: or or wherein represents a fifth cross attention.
[0158] Clause 54. The method of clause 51, wherein the other cross attention is derived by one of the following: or or wherein represents a sixth cross attention.
[0159] Clause 55. The method of clause 45, wherein a motion compensation process at decoder side is performed as the following: wherein represents a first up-sample feature, represents a second up-sample feature, represents a third up-sample feature, represents a fourth up-sample feature, represents a reconstructed frame, represents a fourth cross attention, represents a sixth cross attention, and represents a sixth cross attention.
[0160] Clause 56. The method of clause 22, wherein G is derived by a predetermined cross attention approach between FA (A) and FB (B) , wherein G represents a cross attention, FA (A) represents a first transform of A, FB (B) represents a second transform of B, wherein A represents a first input signal, B represents a second input signal.
[0161] Clause 57. The method of clause 56, wherein G is derived by one of the following: G=softmax (FA (A) (FB (B) ) T) FA (B) , or G=softmax (FA (A) ) softmax ( (FA (B) ) T) FB (B) , or G=softmax (FA (A) ) * [softmax ( (FA (B) ) T) *FB (B) ] , or G=softmax (FA (A) ) * [softmax (FA (B) ) T*FB (B) ] .
[0162] Clause 58. The method of clause 56, wherein G is derived by one of the following: G=softmax (A* BT) *B, or G=softmax (A) *softmax (BT) *B, or G=softmax (A) * [softmax (BT) *B] , or G=softmax (A) * [softmax (B) T*B] .
[0163] Clause 59. The method of clause 56, wherein the first transform and / or the second transform is based on an input signal.
[0164] Clause 60. The method of clause 56, wherein the first transform is same as the second transform.
[0165] Clause 61. The method of clause 56, wherein the first transform and / or the second transform is performed by a network, wherein the network comprises a depth residual bottleneck (DepthRB) .
[0166] Clause 62. The method of clause 61, wherein the DepthRB comprises a first 1×1 convolutional layer, a 3×3 convolutional layer, and a second 1×1 convolutional layer, and an output of the DepthRB is a sum of an input and an output of the second 1×1 convolutional layer.
[0167] Clause 63. The method of clause 56, wherein a network is applied to the cross attention to determine a final cross attention.
[0168] Clause 64. The method of clause 63, wherein the network comprises a convolution layer and a DepthRB block.
[0169] Clause 65. The method of clause 63, wherein G equals to DepthRB (conv5x5 (G) ) , wherein DepthRB represents a depth residual bottleneck and conv5x5 represents a 5×5 convolutional layer.
[0170] Clause 66. The method of clause 56, wherein the first input signal comprises xt or and the second input signal comprises wherein xt represents a predictive frame at a time point, represents a reconstructed frame, and represents a decoded frame at the time point minus 1.
[0171] Clause 67. The method of clause 56, wherein the first input signal comprises xt or and the second input signal comprises wherein xt represents a predictive frame at a time point, represents a reconstructed frame, and represents a first multi-scale features derived from or a propagated feature wherein represents a decoded frame at the time point minus 1.
[0172] Clause 68. The method of clause 56, wherein the first input signal comprises or and the second input signal comprises wherein represents a first down-sample feature, represents a fourth up-sample feature, and represent a second multi-scale feature derived from or a propagated feature wherein represents a decoded frame at a time point minus 1.
[0173] Clause 69. The method of clause 56, wherein the first input signal comprises or and the second input signal comprises wherein represents a second down-sample feature, represents a third up-sample feature, and represent a third multi-scale feature derived from or a propagated feature wherein represents a decoded frame at a time point minus 1.
[0174] Clause 70. The method of clause 1, wherein a joint local and global motion compensation (LGMC) scheme is obtained by combining a global motion compensation (GMC) scheme and a local motion compensation (LMC) scheme, or wherein a joint local and global motion estimation scheme is obtained by combining a global motion estimation scheme and a local motion estimation scheme.
[0175] Clause 71. The method of clause 70, wherein the combination comprises a concatenation in a channel dimension.
[0176] Clause 72. The method of clause 70, wherein the joint local and global motion compensation (LGMC) scheme and / or the joint local and global motion estimation scheme is applied to at least one of a pixel or a feature space.
[0177] Clause 73. The method of clause 72, wherein is obtained by the global motion compensation (GMC) scheme and / or the global motion estimation scheme, and is obtained by the local motion compensation (LMC) scheme and / or the local motion estimation scheme, wherein represents a first cross attention, and represents a first local context.
[0178] Clause 74. The method of clause 72, wherein and are concatenated in a channel dimension, wherein represents a first cross attention, and represents a first local context.
[0179] Clause 75. The method of clause 74, wherein at encoder side, and are concatenated with xt in the channel dimension, which is performed as the following: wherein xt represents a predictive frame at a time point, represents a first down-sample feature, represents a second down-sample feature, represents a third down-sample feature, and represents a fourth down-sample feature.
[0180] Clause 76. The method of clause 74, wherein at decoder side, and are concatenated with in the channel dimension, which is performed as the following: wherein represents a first up-sample feature, represents a second up-sample feature, represents a third up-sample feature, represents a fourth up-sample feature, represents a reconstructed frame.
[0181] Clause 77. The method of clause 72, wherein and are fused by a network, wherein represents a first cross attention, and represents a first local context.
[0182] Clause 78. The method of clause 77, wherein at encoder side, and are fused with xt, wherein xt represents a predictive frame at a time point.
[0183] Clause 79. The method of clause 77, wherein at decoder side, and are fused with wherein represents a reconstructed frame.
[0184] Clause 80. The method of clause 70, wherein the joint local and global motion compensation (LGMC) scheme and / or the joint local and global motion estimation scheme is applied in a multi-scale features space.
[0185] Clause 81. The method of clause 80, wherein and are obtained by the global motion compensation (GMC) scheme and / or the global motion estimation scheme, and and are obtained by the local motion compensation (LMC) scheme and / or the local motion estimation scheme, wherein represents a first cross attention, represents a second cross attention, represents a third cross attention, represents a first local context, represents a second local context, and represents a third local context.
[0186] Clause 82. The method of clause 80, wherein and are concatenated in a channel dimension, and / or and are concatenated in the channel dimension, and / or and are concatenated in the channel dimension, wherein represents a first cross attention, represents a second cross attention, represents a third cross attention, represents a first local context, represents a second local context, and represents a third local context.
[0187] Clause 83. The method of clause 82, wherein at encoder side, and are concatenated with xt in the channel dimension, and are concatenated with a mid-feature in the channel dimension, and and are concatenated with another mid-feature in the channel dimension, which is performed as the following: wherein xt represents a predictive frame at a time point, represents a first down-sample feature, represents a second down-sample feature, represents a third down-sample feature, and represents a fourth down-sample feature.
[0188] Clause 84. The method of clause 82, wherein at decoder side, and are concatenated with in the channel dimension, and are concatenated with a mid-feature in the channel dimension, and and are concatenated with another mid-feature in the channel dimension, which is performed as the following: wherein represents a first up-sample feature, represents a second up-sample feature, represents a third up-sample feature, represents a fourth up-sample feature, represents a reconstructed frame, represents a fourth cross attention, represents a fifth cross attention, and represents a sixth cross attention.
[0189] Clause 85. The method of clause 80, wherein and are fused by a network, and / or and are fused by the network, and / or and are fused by the network, wherein represents a first cross attention, represents a second cross attention, represents a third cross attention, represents a first local context, represents a second local context, and represents a third local context.
[0190] Clause 86. The method of clause 85, wherein at encoder side, and are fused with xt, and / or and are fused with and / or and are fused with wherein represents a first down-sample feature, represents a second down-sample feature, and xt represents a predictive frame at a time point.
[0191] Clause 87. The method of clause 85, wherein at decoder side, and are fused with and / or and are fused with and / or and are fused with wherein represents a third up-sample feature, represents a fourth up-sample feature, and represents a reconstructed frame.
[0192] Clause 88. The method of any of clauses 1-87, wherein the conversion includes encoding the visual data into the bitstream.
[0193] Clause 89. The method of any of clauses 1-87, wherein the conversion includes decoding the visual data from the bitstream.
[0194] Clause 90. An apparatus for visual data processing comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to perform a method in accordance with any of clauses 1-89.
[0195] Clause 91. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform a method in accordance with any of clauses 1-89.
[0196] Clause 92. A non-transitory computer-readable recording medium storing a bitstream of visual data which is generated by a method performed by an apparatus for visual data processing, wherein the method comprises: constructing at least one of a motion compensation scheme or a motion estimation scheme based on at least one of: an optical flow network, a deformable neural network, or a cross attention network; processing the visual data by applying the at least one of the motion compensation scheme or the motion estimation scheme in a learning-based video compression; and generating the bitstream based on the processed visual data.
[0197] Clause 93. A method for storing a bitstream of visual data, comprising: constructing at least one of a motion compensation scheme or a motion estimation scheme based on at least one of: an optical flow network, a deformable neural network, or a cross attention network; processing the visual data by applying the at least one of the motion compensation scheme or the motion estimation scheme in a learning-based video compression; generating the bitstream based on the processed visual data; storing the bitstream in a non-transitory computer-readable recording medium. Example Device
[0198] Fig. 9 illustrates a block diagram of a computing device 900 in which various embodiments of the present disclosure can be implemented. The computing device 900 may be implemented as or included in the source device 110 (or the visual data encoder 114) or the destination device 120 (or the visual data decoder 124) .
[0199] It would be appreciated that the computing device 900 shown in Fig. 9 is merely for purpose of illustration, without suggesting any limitation to the functions and scopes of the embodiments of the present disclosure in any manner.
[0200] As shown in Fig. 9, the computing device 900 includes a general-purpose computing device 900. The computing device 900 may at least comprise one or more processors or processing units 910, a memory 920, a storage unit 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960.
[0201] In some embodiments, the computing device 900 may be implemented as any user terminal or server terminal having the computing capability. The server terminal may be a server, a large-scale computing device or the like that is provided by a service provider. The user terminal may for example be any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, station, unit, device, multimedia computer, multimedia tablet, Internet node, communicator, desktop computer, laptop computer, notebook computer, netbook computer, tablet computer, personal communication system (PCS) device, personal navigation device, personal digital assistant (PDA) , audio / video player, digital camera / video camera, positioning device, television receiver, radio broadcast receiver, E-book device, gaming device, or any combination thereof, including the accessories and peripherals of these devices, or any combination thereof. It would be contemplated that the computing device 900 can support any type of interface to a user (such as “wearable” circuitry and the like) .
[0202] The processing unit 910 may be a physical or virtual processor and can implement various processes based on programs stored in the memory 920. In a multi-processor system, multiple processing units execute computer executable instructions in parallel so as to improve the parallel processing capability of the computing device 900. The processing unit 910 may also be referred to as a central processing unit (CPU) , a microprocessor, a controller or a microcontroller.
[0203] The computing device 900 typically includes various computer storage medium. Such medium can be any medium accessible by the computing device 900, including, but not limited to, volatile and non-volatile medium, or detachable and non-detachable medium. The memory 920 can be a volatile memory (for example, a register, cache, Random Access Memory (RAM) ) , a non-volatile memory (such as a Read-Only Memory (ROM) , Electrically Erasable Programmable Read-Only Memory (EEPROM) , or a flash memory) , or any combination thereof. The storage unit 930 may be any detachable or non-detachable medium and may include a machine-readable medium such as a memory, flash memory drive, magnetic disk or another other media, which can be used for storing information and / or data and can be accessed in the computing device 900.
[0204] The computing device 900 may further include additional detachable / non-detachable, volatile / non-volatile memory medium. Although not shown in Fig. 9, it is possible to provide a magnetic disk drive for reading from and / or writing into a detachable and non-volatile magnetic disk and an optical disk drive for reading from and / or writing into a detachable non-volatile optical disk. In such cases, each drive may be connected to a bus (not shown) via one or more data medium interfaces.
[0205] The communication unit 940 communicates with a further computing device via the communication medium. In addition, the functions of the components in the computing device 900 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, the computing device 900 can operate in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs) or further general network nodes.
[0206] The input device 950 may be one or more of a variety of input devices, such as a mouse, keyboard, tracking ball, voice-input device, and the like. The output device 960 may be one or more of a variety of output devices, such as a display, loudspeaker, printer, and the like. By means of the communication unit 940, the computing device 900 can further communicate with one or more external devices (not shown) such as the storage devices and display device, with one or more devices enabling the user to interact with the computing device 900, or any devices (such as a network card, a modem and the like) enabling the computing device 900 to communicate with one or more other computing devices, if required. Such communication can be performed via input / output (I / O) interfaces (not shown) .
[0207] In some embodiments, instead of being integrated in a single device, some or all components of the computing device 900 may also be arranged in cloud computing architecture. In the cloud computing architecture, the components may be provided remotely and work together to implement the functionalities described in the present disclosure. In some embodiments, cloud computing provides computing, software, data access and storage service, which will not require end users to be aware of the physical locations or configurations of the systems or hardware providing these services. In various embodiments, the cloud computing provides the services via a wide area network (such as Internet) using suitable protocols. For example, a cloud computing provider provides applications over the wide area network, which can be accessed through a web browser or any other computing components. The software or components of the cloud computing architecture and corresponding data may be stored on a server at a remote position. The computing resources in the cloud computing environment may be merged or distributed at locations in a remote data center. Cloud computing infrastructures may provide the services through a shared data center, though they behave as a single access point for the users. Therefore, the cloud computing architectures may be used to provide the components and functionalities described herein from a service provider at a remote location. Alternatively, they may be provided from a conventional server or installed directly or otherwise on a client device.
[0208] The computing device 900 may be used to implement visual data encoding / decoding in embodiments of the present disclosure. The memory 920 may include one or more visual data coding modules 925 having one or more program instructions. These modules are accessible and executable by the processing unit 910 to perform the functionalities of the various embodiments described herein.
[0209] In the example embodiments of performing visual data encoding, the input device 950 may receive visual data as an input 970 to be encoded. The visual data may be processed, for example, by the visual data coding module 925, to generate an encoded bitstream. The encoded bitstream may be provided via the output device 960 as an output 980.
[0210] In the example embodiments of performing visual data decoding, the input device 950 may receive an encoded bitstream as the input 970. The encoded bitstream may be processed, for example, by the visual data coding module 925, to generate decoded visual data. The decoded visual data may be provided via the output device 960 as the output 980.
[0211] While this disclosure has been particularly shown and described with references to preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the present application as defined by the appended claims. Such variations are intended to be covered by the scope of this present application. As such, the foregoing description of embodiments of the present application is not intended to be limiting.
Claims
1.A method of visual data processing, comprising:constructing, for a conversion between visual data and a bitstream of the visual data, at least one of a motion compensation scheme or a motion estimation scheme based on at least one of: an optical flow network, a deformable neural network, or a cross attention network;processing the visual data by applying the at least one of the motion compensation scheme or the motion estimation scheme in a learning-based video compression; andperforming the conversion based on the processed visual data.2.The method of claim 1, wherein applying the at least one of the motion compensation scheme or the motion estimation scheme in the learning-based video compression comprises:applying a global motion compensation (GMC) scheme and / or a global motion estimation scheme in the learning-based video compression.3.The method of claim 3, wherein the global motion compensation scheme and / or the global motion estimation scheme is used to generate global motion compensation, orwherein the global motion compensation scheme and / or the global motion estimation scheme is used to derive at least one of motion content or motion information.4.The method of claim 1, wherein applying the at least one of the motion compensation scheme or the motion estimation scheme in the learning-based video compression comprises:applying a local motion compensation (LMC) scheme and / or a local motion estimation scheme in the learning-based video compression.5.The method of claim 4, wherein the local motion compensation scheme and / or the local motion esti-mation scheme is used to generate a local motion compensation, orwherein the local motion compensation scheme and / or the local motion estimation scheme is used to derive at least one of motion content or motion information.6.The method of claim 1, wherein applying the at least one of the motion compensation scheme or the motion estimation scheme in the learning-based video compression comprises:applying a joint local and global motion compensation (LGMC) scheme and / or a joint local and global motion estimation scheme in the learning-based video compression.7.The method of claim 6, wherein the joint local and global motion compensation scheme and / or the joint local and global motion estimation scheme is used to generate joint local and global motion compensation, orwherein the joint local and global motion compensation scheme and / or the joint local and global motion estimation scheme is used to derive at least one of motion content or motion information.8.The method of claim 6, wherein the joint local and global motion compensation scheme is combined by a global motion compensation (GMC) scheme and a local motion compensation (LMC) scheme, orwherein the joint local and global motion estimation scheme is combined by a global motion estimation scheme and a local motion estimation scheme.9.The method of claim 1, wherein the at least one of the motion compensation scheme or the motion estimation scheme is constructed based on a learning-based approach.10.The method of claim 9, wherein the learning-based approach comprises a neural network.11.The method of claim 1, wherein the at least one of the motion compensation scheme or the motion estimation scheme is applied to at least one of: a pixel space or a feature space.12.The method of claim 1, wherein the at least one of the motion compensation scheme or the motion estimation scheme is applied to at least one of the following in a feature space: a single-scale feature or a multi-scale feature.13.The method of any of claims 1 to 12, wherein the at least one of the motion compensation scheme or the motion estimation scheme comprises at least one of: a global motion compensation (GMC) scheme, a global motion estimation scheme, a local motion compensation (LMC) scheme, a local motion estimation scheme, a joint local and global motion compensation (LGMC) scheme, or a joint local and global motion estimation scheme.14.The method of claim 1, wherein constructing the at least one of the motion compensation scheme or the motion estimation scheme based on the at least one of: the optical flow network, the deformable neural network, or the cross attention network comprises:constructing a local motion compensation (LMC) scheme and / or a local motion estimation scheme by using the optical flow network.15.The method of claim 14, wherein the optical flow network is applied to determine at least one of an offset or motion vector between a first frame and a second frame, wherein the first frame comprises a predictive frame at a time point, and the second frame comprises a decoded frame at the time point minus a time interval.16.The method of claim 15, wherein the local motion compensation scheme and / or the local motion estimation scheme is applied in at least one of: a pixel space or a feature space.17.The method of claim 15, wherein the local motion compensation scheme and / or the local motion estimation scheme is applied in a multi-scale feature.18.The method of claim 1, wherein constructing the at least one of the motion compensation scheme or the motion estimation scheme based on the at least one of: the optical flow network, the deformable neural network, or the cross attention network comprises:constructing a local motion compensation (LMC) scheme and / or a local motion estimation scheme by using the deformable neural network.19.The method of claim 18, wherein the deformable neural network is applied to determine at least one of an offset or motion vector between a first frame and a second frame, wherein the first frame comprises a predictive frame at a time point, and the second frame comprises a decoded frame at the time point minus a time interval.20.The method of claim 19, wherein the local motion compensation scheme and / or the local motion estimation scheme is applied in at least one of: a pixel space or a feature space.21.The method of claim 19, wherein the local motion compensation scheme and / or the local motion estimation scheme is applied in a multi-scale feature.22.The method of claim 1, wherein constructing the at least one of the motion compensation scheme or the motion estimation scheme based on the at least one of: the optical flow network, the deformable neural network, or the cross attention network comprises:constructing a global motion compensation (GMC) scheme and / or a global motion estimation scheme by using the cross attention network.23.The method of claim 22, wherein the global motion compensation (GMC) scheme and / or the global motion estimation scheme is applied in a pixel space.24.The method of claim 23, wherein if the global motion compensation is performed, is derived at encoder side by a predetermined cross attention approach between and xt, wherein represents a cross attention, xt represents a predictive frame at a time point, and represents a decoded frame at the time point minus 1.25.The method of claim 24, wherein the cross attention is derived by the following: 26.The method of claim 24, wherein the cross attention is derived by the following: 27.The method of claim 24, wherein the cross attention is derived by the following: 28.The method of claim 23, wherein a global motion compensation process at encoder side is performed as the following: whereinrepresents cross attention, xt represents a predictive frame at a time point, represents a first down-sample feature, represents a second down-sample feature, represents a third down-sample fea-ture, andrepresents a fourth down-sample feature.29.The method of claim 23, wherein is derived at decoder side by another predetermined cross atten-tion approach between and wherein represents another cross attention, represents a decoded frame at a time point minus 1, and represents a reconstructed frame.30.The method of claim 29, wherein the other cross attention is derived by the following: 31.The method of claim 29, wherein the other cross attention is derived by the following: 32.The method of claim 29, wherein the other cross attention is derived by the following: 33.The method of claim 23, wherein a global motion compensation process at decoder side is performed as the following: whereinrepresents a first up-sample feature, represents a second up-sample feature, represents a third up-sample feature, represents a fourth up-sample feature, represents a reconstructed frame, andrepresents another cross attention.34.The method of claim 22, wherein the global motion compensation (GMC) scheme and / or the global motion estimation scheme is applied to a feature space.35.The method of claim 34, wherein is derived at encoder side by a predetermined cross attention approach between and xt, wherein represents a cross attention, xt represents a predictive frame at a time point, and represents a feature derived from or a propagated feature wherein represents a decoded frame at the time point minus 1.36.The method of claim 35, wherein the cross attention is derived by the following: 37.The method of claim 35, wherein the cross attention is derived by the following: 38.The method of claim 35, wherein the cross attention is derived by the following: 39.The method of claim 34, wherein a global motion compensation process at encoder side is performed as the following: whereinrepresents a cross attention, xt represents a predictive frame at a time point, represents a first down-sample feature, represents a second down-sample feature, represents a third down-sample fea-ture, andrepresents a fourth down-sample feature.40.The method of claim 34, wherein is derived at decoder side by another predetermined cross atten-tion approach between and wherein represents another cross attention, represents a reconstructed frame, and represents a feature derived from or a propagated feature wherein represents a decoded frame at a time point minus 1.41.The method of claim 40, wherein the other cross attention is derived by the following: 42.The method of claim 40, wherein the other cross attention is derived by the following: 43.The method of claim 40, wherein the other cross attention is derived by the following: 44.The method of claim 34, wherein a global motion compensation process at decoder side is performed as the following: whereinrepresents a first up-sample feature, represents a second up-sample feature, represents a third up-sample feature, represents a fourth up-sample feature, represents a reconstructed frame, andrepresents another cross attention.45.The method of claim 22, wherein the global motion compensation (GMC) scheme and / or the global motion estimation scheme is applied to multi-scale features.46.The method of claim 45, wherein cross attention is derived at encoder side by a predetermined cross attention approach between and xt, and and wherein xt represents a predictive frame at a time point, represents a first down-sample feature, represents a second down-sample feature, and and represent multi-scale features derived from or a propagated feature wherein represents a de-coded frame at the time point minus 1.47.The method of claim 46, wherein the cross attention is derived by the following: or or whereinrepresents a first cross attention.48.The method of claim 46, wherein the cross attention is derived by the following: or or whereinrepresents a second cross attention.49.The method of claim 46, wherein the cross attention is derived by the following: or or whereinrepresents a third cross attention.50.The method of claim 45, wherein a global motion compensation process at encoder side is performed as the following: whereinrepresents a first cross attention, represents a second cross attention, represents a third cross attention, xt represents a predictive frame at a time point, represents a first down-sample feature, represents a second down-sample feature, represents a third down-sample feature, andrepresents a fourth down-sample feature.51.The method of claim 45, wherein another cross attention is derived at decoder side by another prede-termined cross attention approach between and and and wherein represents a recon-structed frame, represents a fourth up-sample feature, represents a third up-sample feature, and and represent multi-scale features derived from or a propagated feature wherein represents a decoded frame at a time point minus 1.52.The method of claim 51, wherein the other cross attention is derived by one of the following: or or whereinrepresents a fourth cross attention.53.The method of claim 51, wherein the other cross attention is derived by one of the following: or or whereinrepresents a fifth cross attention.54.The method of claim 51, wherein the other cross attention is derived by one of the following: or or whereinrepresents a sixth cross attention.55.The method of claim 45, wherein a motion compensation process at decoder side is performed as the following: whereinrepresents a first up-sample feature, represents a second up-sample feature, represents a third up-sample feature, represents a fourth up-sample feature, represents a reconstructed frame, rep-resents a fourth cross attention, represents a sixth cross attention, andrepresents a sixth cross attention.56.The method of claim 22, wherein G is derived by a predetermined cross attention approach between FA (A) and FB (B) , wherein G represents a cross attention, FA (A) represents a first transform of A, FB (B) repre-sents a second transform of B, wherein A represents a first input signal, B represents a second input signal.57.The method of claim 56, wherein G is derived by one of the following: G=softmax (FA (A) (FB (B) ) T) FA (B) , or G=softmax (FA (A) ) softmax ( (FA (B) ) T) FB (B) , or G=softmax (FA (A) ) * [softmax ( (FA (B) ) T) *FB (B) ] , or G=softmax (FA (A) ) * [softmax (FA (B) ) T*FB (B) ]58.The method of claim 56, wherein G is derived by one of the following: G=softmax (A*BT) *B, or G=softmax (A) *softmax (BT) *B, or G=softmax (A) * [softmax (BT) *B] , or G=softmax (A) * [softmax (B) T*B]59.The method of claim 56, wherein the first transform and / or the second transform is based on an input signal.60.The method of claim 56, wherein the first transform is same as the second transform.61.The method of claim 56, wherein the first transform and / or the second transform is performed by a network, wherein the network comprises a depth residual bottleneck (DepthRB) .62.The method of claim 61, wherein the DepthRB comprises a first 1×1 convolutional layer, a 3×3 convolutional layer, and a second 1×1 convolutional layer, and an output of the DepthRB is a sum of an input and an output of the second 1×1 convolutional layer.63.The method of claim 56, wherein a network is applied to the cross attention to determine a final cross attention.64.The method of claim 63, wherein the network comprises a convolution layer and a DepthRB block.65.The method of claim 63, wherein G equals to DepthRB (conv5x5 (G) ) , wherein DepthRB represents a depth residual bottleneck and conv5x5 represents a 5×5 convolutional layer.66.The method of claim 56, wherein the first input signal comprises xt or and the second input signal comprises wherein xt represents a predictive frame at a time point, represents a reconstructed frame, and represents a decoded frame at the time point minus 1.67.The method of claim 56, wherein the first input signal comprises xt or and the second input signal comprises wherein xt represents a predictive frame at a time point, represents a reconstructed frame, and represents a first multi-scale features derived from or a propagated feature wherein repre-sents a decoded frame at the time point minus 1.68.The method of claim 56, wherein the first input signal comprises or and the second input signal comprises wherein represents a first down-sample feature, represents a fourth up-sample feature, and represent a second multi-scale feature derived from or a propagated feature wherein repre-sents a decoded frame at a time point minus 1.69.The method of claim 56, wherein the first input signal comprises or and the second input signal comprises wherein represents a second down-sample feature, represents a third up-sample feature, and represent a third multi-scale feature derived from or a propagated feature wherein represents a decoded frame at a time point minus 1.70.The method of claim 1, wherein a joint local and global motion compensation (LGMC) scheme is obtained by combining a global motion compensation (GMC) scheme and a local motion compensation (LMC) scheme, orwherein a joint local and global motion estimation scheme is obtained by combining a global motion estimation scheme and a local motion estimation scheme.71.The method of claim 70, wherein the combination comprises a concatenation in a channel dimension.72.The method of claim 70, wherein the joint local and global motion compensation (LGMC) scheme and / or the joint local and global motion estimation scheme is applied to at least one of a pixel or a feature space.73.The method of claim 72, wherein is obtained by the global motion compensation (GMC) scheme and / or the global motion estimation scheme, and is obtained by the local motion compensation (LMC) scheme and / or the local motion estimation scheme, wherein represents a first cross attention, and represents a first local context.74.The method of claim 72, wherein and are concatenated in a channel dimension, wherein represents a first cross attention, and represents a first local context.75.The method of claim 74, wherein at encoder side, and are concatenated with xt in the channel dimension, which is performed as the following: wherein xt represents a predictive frame at a time point, represents a first down-sample feature, represents a second down-sample feature, represents a third down-sample feature, andrepresents a fourth down-sample feature.76.The method of claim 74, wherein at decoder side, and are concatenated with in the channel dimension, which is performed as the following: whereinrepresents a first up-sample feature, represents a second up-sample feature, represents a third up-sample feature, represents a fourth up-sample feature, represents a reconstructed frame.77.The method of claim 72, wherein and are fused by a network, wherein represents a first cross attention, and represents a first local context.78.The method of claim 77, wherein at encoder side, and are fused with xt, wherein xt represents a predictive frame at a time point.79.The method of claim 77, wherein at decoder side, and are fused with wherein represents a reconstructed frame.80.The method of claim 70, wherein the joint local and global motion compensation (LGMC) scheme and / or the joint local and global motion estimation scheme is applied in a multi-scale features space.81.The method of claim 80, wherein and are obtained by the global motion compensation (GMC) scheme and / or the global motion estimation scheme, and and are obtained by the local motion compensation (LMC) scheme and / or the local motion estimation scheme, wherein represents a first cross attention, represents a second cross attention, represents a third cross attention, represents a first local context, represents a second local context, and represents a third local context.82.The method of claim 80, wherein and are concatenated in a channel dimension, and / or and are concatenated in the channel dimension, and / or and are concatenated in the channel dimension, wherein represents a first cross attention, represents a second cross attention, represents a third cross attention, represents a first local context, represents a second local context, and represents a third local context.83.The method of claim 82, wherein at encoder side, and are concatenated with xt in the channel dimension, and are concatenated with a mid-feature in the channel dimension, and and are con-catenated with another mid-feature in the channel dimension, which is performed as the following: wherein xt represents a predictive frame at a time point, represents a first down-sample feature, represents a second down-sample feature, represents a third down-sample feature, andrepresents a fourth down-sample feature.84.The method of claim 82, wherein at decoder side, and are concatenated with in the channel dimension, and are concatenated with a mid-feature in the channel dimension, and and are con-catenated with another mid-feature in the channel dimension, which is performed as the following: whereinrepresents a first up-sample feature, represents a second up-sample feature, represents a third up-sample feature, represents a fourth up-sample feature, represents a reconstructed frame, rep-resents a fourth cross attention, represents a fifth cross attention, andrepresents a sixth cross attention.85.The method of claim 80, wherein and are fused by a network, and / or and are fused by the network, and / or and are fused by the network, wherein represents a first cross attention, repre-sents a second cross attention, represents a third cross attention, represents a first local context, repre-sents a second local context, and represents a third local context.86.The method of claim 85, wherein at encoder side, and are fused with xt, and / or and are fused with and / or and are fused with wherein represents a first down-sample feature, repre-sents a second down-sample feature, and xt represents a predictive frame at a time point.87.The method of claim 85, wherein at decoder side, and are fused with and / or and are fused with and / or and are fused with wherein represents a third up-sample feature, repre-sents a fourth up-sample feature, and represents a reconstructed frame.88.The method of any of claims 1-87, wherein the conversion includes encoding the visual data into the bitstream.89.The method of any of claims 1-87, wherein the conversion includes decoding the visual data from the bitstream.90.An apparatus for visual data processing comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to perform a method in accordance with any of claims 1-89.91.A non-transitory computer-readable storage medium storing instructions that cause a processor to perform a method in accordance with any of claims 1-89.92.A non-transitory computer-readable recording medium storing a bitstream of visual data which is generated by a method performed by an apparatus for visual data processing, wherein the method comprises:constructing at least one of a motion compensation scheme or a motion estimation scheme based on at least one of: an optical flow network, a deformable neural network, or a cross attention network;processing the visual data by applying the at least one of the motion compensation scheme or the motion estimation scheme in a learning-based video compression; andgenerating the bitstream based on the processed visual data.93.A method for storing a bitstream of visual data, comprising:constructing at least one of a motion compensation scheme or a motion estimation scheme based on at least one of: an optical flow network, a deformable neural network, or a cross attention network;processing the visual data by applying the at least one of the motion compensation scheme or the motion estimation scheme in a learning-based video compression;generating the bitstream based on the processed visual data; andstoring the bitstream in a non-transitory computer-readable recording medium.
Citation Information
Patent Citations
Video compression method based on deep learning
CN111294604A
Video enhancement method, device and equipment and computer medium
CN116996692A
Image encoding, decoding method and device, coder-decoder
US20230171435A1