Video coding method, computer device, equipment and computer readable medium
By generating virtual reference data through a deep neural network-based inter-frame prediction method, the problem of low efficiency in traditional video coding when handling complex motion is solved, achieving more efficient video compression and quality improvement.
Patent Information
- Application Number
- CN202511118250.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2021-09-16
- Filing Date
- 2021-09-24
- Publication Date
- 2025-11-18
AI Technical Summary
Existing video coding standards are inefficient when dealing with complex motion and non-translational motion. Traditional block-based motion vector estimation methods are prone to errors and have difficulty effectively reducing artifacts at motion boundaries.
A deep neural network-based inter-frame prediction method is adopted. This method generates virtual reference data for inter-frame prediction and uses multiple adjacent reference frames to generate feature maps and refine them to improve frame quality and reduce artifacts on motion boundaries.
It improves the compression efficiency of video encoding, reduces artifacts at motion boundaries such as noise and blur, and enhances video quality.
Smart Images

Figure CN120980245A_ABST
Abstract
Description
This application is a divisional application of the original application with the original application number 202180030747X and the original application filing date of 2021-09-24. Cross Reference to Related Applications
[0001] This application is based on and claims priority to U.S. Provisional Patent Application No. 63 / 131,625, filed on December 29, 2020, and U.S. Patent Application No. 17 / 476,928, filed on September 16, 2021, the disclosures of which are incorporated by reference in their entireties. TECHNICAL FIELD
[0002] The present disclosure relates to a method of video coding, a computer device, an apparatus, and a computer readable medium. BACKGROUND
[0003] Uncompressed digital video can include a series of pictures, each picture having a spatial dimension, for example 1920 x 1080 luma samples and associated chroma samples. The series of pictures can have a fixed or variable picture rate (also known as frame rate), for example 60 pictures per second or 60 Hz. Uncompressed video has significant bandwidth requirements. For example, 1080p60 4:2:0 video at 8 bit per sample (1920x1080 luma samples at 60 Hz) requires almost 1.5 Gbit / s bandwidth. An hour of such video requires more than 600 GBytes of storage.
[0004] Traditional video coding standards, such as H.264 / Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC), share a similar (recursive) block-based hybrid prediction / transform framework, where individual coding tools like intra / inter prediction, integer transform and context adaptive entropy coding are centrally hand-crafted to optimize overall efficiency. SUMMARY
[0005] According to embodiments, a method of video coding using neural network based inter prediction is provided, the method is performed by at least one processor, and the method includes: generating an intermediate stream based on an input frame; generating a reconstructed frame by performing backward warping of the input frame with the intermediate stream; generating a fusion map and a residual map based on the input frame, the intermediate stream, and the reconstructed frame; generating a feature map having a plurality of levels using a first neural network based on a current reference frame, a first reference frame, and a second reference frame; generating a predicted frame based on aligned features from the generated feature map by refining the current reference frame, the first reference frame, and the second reference frame; generating a final residual based on the predicted frame; and computing an enhanced frame as an output by adding the final residual to the current reference frame.
[0006] According to embodiments, there is provided a computer apparatus, characterized in that the computer apparatus comprises one or more computer-readable non-transitory storage media configured to store computer program code; and one or more computer processors configured to access said computer program code and to carry out the above-described method of video coding using neural network-based inter prediction as instructed by the computer program code.
[0007] According to embodiments, there is provided an apparatus for video coding using neural network-based inter prediction, characterized in that the apparatus comprises a first generating unit configured to generate an intermediate stream based on an input frame; a second generating unit configured to perform backward warping of the input frame with the intermediate stream to generate a reconstructed frame; a fusion unit configured to generate a fusion map and a residual map based on the input frame, the intermediate stream, and the reconstructed frame; a third generating unit configured to generate a feature map with multiple levels using a first neural network based on a current reference frame, a first reference frame, and a second reference frame; a prediction unit configured to predict a frame based on aligned features from the generated feature map by refining the current reference frame, the first reference frame, and the second reference frame; a residual unit configured to generate a final residual based on the predicted frame; and a fourth generating unit configured to generate an enhanced frame as an output by adding the final residual to the current reference frame.
[0008] According to embodiments, there is provided a non-transitory computer readable medium storing instructions that, when executed by at least one processor for video coding using neural network-based inter prediction, cause the at least one processor to perform the above-described method of video coding using neural network-based inter prediction.
[0009] The method, computer apparatus, apparatus, and computer readable medium of video coding of the present disclosure describe a deep neural network (DNN) based model related to video coding and decoding. More specifically, as a video frame interpolation (VFI) task and a detail enhancement module, the (DNN) based model generates virtual reference data using and from neighboring reference frames for inter prediction to further improve the quality of frames and reduce artifacts (e.g., noise, blurring, blocking effects, etc.) on motion boundaries. BRIEF DESCRIPTION OF DRAWINGS
[0010] Figure 1 is a diagram of an environment in which the methods, apparatus and systems described herein can be implemented in accordance with embodiments.
[0011] Figure 2 is a block diagram of example components of one or more devices of Figure 1 .
[0012] Figure 3 is a schematic diagram of an example of virtual reference picture generation and insertion into a reference picture list.
[0013] Figure 4 is a block diagram of a test apparatus for a virtual reference generation process during a test phase according to an embodiment.
[0014] Figure 5 is a detailed block diagram of an optical flow estimation and intermediate frame synthesis module from the test apparatus of Figure 4 during a test phase according to an embodiment.
[0015] Figure 6 is a detailed block diagram of a detail enhancement module from the test apparatus of Figure 4 during a test phase according to an embodiment.
[0016] Figure 7 is a detailed block diagram of a PCD alignment module during a test phase according to an embodiment.
[0017] Figure 8 is a detailed block diagram of a TSA fusion module during a test phase according to an embodiment.
[0018] Figure 9 is a detailed block diagram of a TSA fusion module during a test phase according to another embodiment.
[0019] Figure 10 is a flow diagram of a method for video coding using neural network based inter prediction according to an embodiment.
[0020] Figure 11 is a block diagram of an apparatus for video coding using neural network based inter prediction according to an embodiment. DETAILED DESCRIPTION
[0021] The present disclosure describes deep neural network (DNN) based models related to video coding and decoding. More specifically, as a video frame interpolation (VFI) task and a detail enhancement module, the (DNN) based models generate virtual reference data using and from neighboring reference frames for inter prediction to further improve the quality of frames and reduce artifacts (e.g., noise, blurring, blockiness effects, etc.) on motion boundaries.
[0022] One objective of video encoding and decoding is to reduce redundancy in the input video signal through compression. Compression can help reduce bandwidth or storage requirements by two orders of magnitude or more in some cases. Lossless compression, lossy compression, and combinations thereof can be used. Lossless compression refers to techniques in which an exact copy of the original signal can be reconstructed from the compressed original signal. When using lossy compression, the reconstructed signal may differ from the original signal, but the distortion between the original and the reconstructed signal is small enough that the reconstructed signal is useful for the intended application. Lossy compression is widely used in the case of video. The amount of distortion tolerated depends on the application; for example, users of some consumer streaming applications may tolerate higher distortion than users of television streaming applications. The achievable compression ratio reflects this: higher permissible / acceptable distortion can result in a higher compression ratio.
[0023] Essentially, spatiotemporal pixel neighborhoods are used to construct the predicted signal to obtain the corresponding residuals for subsequent transformation, quantization, and entropy coding. On the other hand, deep neural networks (DNNs) are essentially used to extract different levels of spatiotemporal stimuli by analyzing the spatiotemporal information from the receptive fields of neighboring pixels. The ability to highly explore nonlinear and nonlocal spatiotemporal correlations offers promising opportunities to significantly improve compression quality.
[0024] One caveat of utilizing information from multiple adjacent video frames is the complex motion caused by moving cameras and dynamic scenes. Traditional block-based motion vectors do not perform well for non-translational motion. Learning-based optical flow methods can provide pixel-level accurate motion information, but unfortunately, they are error-prone, especially at the boundaries of moving objects. This disclosure proposes using a DNN-based model to implicitly handle arbitrarily complex motion in a data-driven manner without explicit motion estimation.
[0025] Figure 1 This is a diagram of an environment 100 in which the methods, apparatus and systems described herein can be implemented, according to an embodiment.
[0026] like Figure 1 As shown, environment 100 may include user device 110, platform 120, and network 130. Devices in environment 100 may be interconnected via wired connection, wireless connection, or a combination of wired and wireless connection.
[0027] User devices 110 include one or more devices capable of receiving, generating, storing, processing, and / or providing information associated with platform 120. For example, user devices 110 can include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smart phone, a wireless phone, etc.), a wearable device (e.g., a pair of smart glasses or a smart watch), or a similar device. In some implementations, user devices 110 can receive information from and / or send information to platform 120.
[0028] Platform 120 includes one or more devices as described elsewhere herein. In some implementations, platform 120 can include a cloud server or a group of cloud servers. In some implementations, platform 120 can be designed to be modular such that software components can be swapped in or out. Thus, platform 120 can be easily and / or quickly reconfigured for different uses.
[0029] In some implementations, as shown, platform 120 can be hosted in cloud computing environment 122. Notably, while the implementations described herein describe platform 120 as being hosted in cloud computing environment 122, in some implementations, platform 120 can not be cloud-based (i.e., can be implemented outside of a cloud computing environment) or can be partially cloud-based.
[0030] Cloud computing environment 122 includes an environment that hosts platform 120. Cloud computing environment 122 can provide computing, software, data access, storage, etc. services that do not require end users (e.g., user devices 110) to know the physical location and configuration of the system(s) and / or device(s) that host platform 120. As shown, cloud computing environment 122 can include a group of computing resources 124 (collectively referred to as “computing resources 124” and individually as “computing resource 124”).
[0031] Computing resources 124 include one or more personal computers, workstation computers, server devices, or other types of computation and / or communication devices. In some implementations, computing resources 124 can host platform 120. Cloud resources can include computing instances executing in computing resources 124, storage devices provided in computing resources 124, data transfer devices provided by computing resources 124, etc. In some implementations, computing resources 124 can communicate with other computing resources 124 via wired connections, wireless connections, or a combination of wired and wireless connections.
[0032] Further asFigure 1 As shown, computing resources 124 include a set of cloud resources such as one or more applications (“APPs”) 124-1, one or more virtual machines (“VMs”) 124-2, virtualized storage (“VSs”) 124-3, one or more hypervisors (“HYPs”) 124-4, and the like.
[0033] Applications 124-1 include one or more software applications that can be provided to and / or accessed by user devices 110 and / or platform 120. Applications 124-1 can eliminate the need to install and execute software applications on user devices 110. For example, applications 124-1 can include software associated with platform 120 and / or any other software capable of being provided via cloud computing environment 122. In some implementations, one application 124-1 can send / receive information to / from one or more other applications 124-1 via virtual machines 124-2.
[0034] Virtual machines 124-2 include a software implementation of a machine (e.g., a computer) that executes programs like a physical machine. Virtual machines 124-2 can be either system virtual machines or process virtual machines, depending on the degree and manner in which any real machine is leveraged by a virtual machine 124-2. System virtual machines can provide a complete system platform that supports execution of a complete operating system (“OS”). Process virtual machines can execute a single program and can support a single process. In some implementations, virtual machines 124-2 can execute on behalf of users (e.g., user devices 110) and can manage infrastructure of cloud computing environment 122, such as data management, synchronization, or long-duration data transfers.
[0035] Virtualized storage 124-3 includes one or more storage systems and / or one or more devices that use virtualization techniques within the storage systems or devices of computing resources 124. In some implementations, in the case of storage systems, types of virtualization can include block virtualization and file virtualization. Block virtualization can refer to decoupling (or detachment) of logical storage from physical storage such that a storage system can be accessed without regard to physical storage or heterogeneous structure. Detachment can allow an administrator of a storage system flexibility in how the administrator manages storage for end users. File virtualization can eliminate dependency on where data is physically stored for data accessed at a file level. This can enable optimization of storage usage, server consolidation, and / or performance of non-disruptive file migrations.
[0036] The hypervisor 124-4 can provide a hardware virtualization technique that enables multiple operating systems (e.g., “guest operating systems”) to execute concurrently on a host computer such as the computing resource 124. The hypervisor 124-4 can present a virtual operating platform to the guest operating systems and can manage the execution of the guest operating systems. Multiple instances of various operating systems can share virtualized hardware resources.
[0037] The network 130 includes one or more wired and / or wireless networks. For example, the network 130 can include a cellular network (e.g., a fifth generation (5G) network, a long-term evolution (LTE) network, a third generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a
[0038] The apparatuses and networks shown are examples. Figure 1 In practice, there can be additional apparatuses and / or networks, fewer apparatuses and / or networks, different apparatuses and / or networks, or differently arranged apparatuses and / or networks than those shown. Figure 1 than those shown, or a different arrangement of apparatuses and / or networks than those shown. In addition, or alternatively, Figure 1 two or more of the apparatuses shown can be implemented within a single apparatus, or Figure 1 a single apparatus shown can be implemented as multiple, distributed apparatuses. Additionally or alternatively, a set of apparatuses (e.g., one or more apparatuses) of the environment 100 can perform one or more functions of another set of apparatuses of the environment 100 that are described as performing the one or more functions.
[0039] Figure 2 is a block diagram of example components of one or more apparatuses of Figure 1
[0040] The apparatus 200 can correspond to the user apparatus 110 and / or the platform 120. As Figure 2 shown, the apparatus 200 can include a bus 210, a processor 220, a memory 230, a storage component 240, an input component 250, an output component 260, and a
[0041] Bus 210 includes a component that permits communication among the components of device 200. Processor 220 is implemented in hardware, firmware, or a combination of hardware and software. Processor 220 is a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or another type of processing component. In some implementations, processor 220 includes one or more processors capable of being programmed to perform a function. Memory 230 includes a random access memory (RAM), a read only memory (ROM), and / or another type of dynamic or static storage (e.g., a flash memory, a magnetic storage, and / or an optical storage) that stores information and / or instructions for use by processor 220.
[0042] Storage component 240 stores information and / or software related to the operation and use of device 200. For example, storage component 240 can include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optic disk, and / or a solid state disk), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cartridge, a magnetic tape, and / or another type of non-transitory computer-readable medium, along with a corresponding drive.
[0043] Input component 250 includes a component that permits device 200 to receive information, such as via user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, and / or a microphone). Additionally, or alternatively, input component 250 can include a sensor for sensing information (e.g., a global positioning system (GPS) component, an accelerometer, a gyroscope, and / or an actuator). Output component 260 includes a component that provides output information from device 200 (e.g., a display, a speaker, and / or one or more light-emitting diodes (LEDs)).
[0044] Communication interface 270 includes a transceiver-like component (e.g., a transceiver and / or a separate receiver and transmitter) that enables device 200 to communicate with other devices, such as via a wired connection, a wireless connection, or a combination of wired and wireless connections. Communication interface 270 can permit device 200 to receive information from another device and / or provide information to another device. For example, communication interface 270 can include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, and / or the like.
[0045] Apparatus 200 may perform one or more of the processes described herein. Apparatus 200 may perform these processes in response to processor 220 executing software instructions stored in non-transitory computer-readable media such as memory 230 and / or storage unit 240. Computer-readable media are defined herein as non-transitory memory devices. Memory devices include memory space within a single physical storage device or memory space distributed across multiple physical storage devices.
[0046] Software instructions can be read into memory 230 and / or storage unit 240 from another computer-readable medium or from another device via communication interface 270. When the software instructions stored in memory 230 and / or storage unit 240 are executed, the software instructions can cause processor 220 to perform one or more processes described herein. Additionally or alternatively, hardwired circuitry can be used in place of or in combination with the software instructions to perform one or more processes described herein. Therefore, the implementations described herein are not limited to any particular combination of hardware circuitry and software.
[0047] Provided Figure 2 The number and arrangement of components shown are for illustrative purposes only. In reality, besides... Figure 2 In addition to the components shown, device 200 may include additional components, fewer components, different components, or components arranged differently. Alternatively or additionally, a set of components of device 200 (e.g., one or more components) may perform one or more functions described as being performed by another set of components of device 200.
[0048] A typical video compression framework can be described as follows. In the first motion estimation step, given multiple image frames x1,…,x… T Given an input video x, divide these frames into spatial blocks. Each block can be iteratively divided into smaller blocks, and the current frame x can be computed for each block. t Compared with the previously reconstructed frame set The set of motion vectors m between t Note that the subscript t indicates the current t-th encoding cycle, which may not match the timestamp of the image frame. Furthermore, the previously reconstructed frame set... This includes frames from multiple previous coding cycles. Then, in the second motion compensation step, based on the motion vector m... t By copying the previously rebuilt frame set To obtain the predicted frame, use the corresponding pixels. And the original frame x can be obtained t With the predicted frame The residual r between t (Right now, In the third motion compensation step, the residual r is...t Quantization is performed. The quantization step gives quantized The motion vector m t and quantized are both encoded into a bitstream by entropy coding, which is sent to the decoder. Then, on the decoder side, quantized is dequantized (usually by an inverse transform similar to IDCT using dequantization coefficients) to obtain recovered residual The recovered residual is then added back to the predicted frame to obtain the reconstructed frame (i.e., ). Additional components are further used to improve the visual quality of the reconstructed frame . Typically, one or more of the following enhancement modules can be selected to process the reconstructed frame The enhancement modules include a deblocking filter (DF), a sample adaptive offset (SAO), an adaptive loop filter (ALF), and the like.
[0049] In HEVC, VVC, or other video coding frameworks or standards, the decoded picture can be included in a reference picture list (RPL) and can be used as a reference picture for motion-compensated prediction and other parameter prediction for encoding one or more pictures following in the encoding or decoding order, or can be used for intra prediction or block copy for encoding different regions or blocks of the current picture.
[0050] In embodiments, one or more virtual references can be generated and included in the RPL in both the encoder and the decoder or only in the decoder. The virtual reference picture can be generated by one or more processes including signal processing, spatial or temporal filtering, scaling, weighted averaging, up / down sampling, pooling, memory recursive processing, linear system processing, nonlinear system processing, neural network processing, deep learning based processing, artificial intelligence processing, pre-trained network processing, machine learning based processing, online trained network processing, or combinations thereof. For the process of generating one or more virtual references, zero or more forward reference pictures preceding the current picture in both the output / display order and the encoding / decoding order and zero or more backward reference pictures following the current picture in the output / display order but preceding the current picture in the encoding / decoding order are used as input data. The output of the process is the virtual / generated picture to be used as a new reference picture.
[0051] A DNN pre-trained network process for virtual reference picture generation is described according to embodiments. Figure 3An example of virtual reference picture generation and insertion into a reference picture list is shown, including a hierarchical GOP structure 300, a reference picture list 310, and a virtual reference generation process 320.
[0052] As Figure 3 shown, considering the hierarchical GOP structure 300, when the picture order count (POC) of the current picture is equal to 3, generally, decoded pictures with POC equal to 0, 2, 4, or 8 can be stored in a decoded picture buffer, and some of the decoded pictures are included in the reference picture list for decoding the current picture. As an example, the most recently decoded pictures with POC equal to 2 or 4 can be fed as input data into the virtual reference generation process 320. The virtual reference picture can be generated by one or more processes. The generated virtual reference picture can be stored in the decoded picture buffer and included in the reference picture list 310 of the current picture or one or more future pictures in decoding order. If the virtual reference picture is included in the reference picture list 310 of the current picture, the pixel data of the generated virtual reference picture can be used as reference data for motion-compensated prediction when the virtual reference picture is indicated to be used by a reference index.
[0053] In the same or another embodiment, the entire virtual reference generation process can include one of more signaling processing modules using one or more pre-trained neural network models or any predefined parameters. For example, as Figure 3 shown, the entire virtual reference generation process 320 can include a flow estimation module 330, a flow compensation module 340, and a detail enhancement module 350. The flow estimation module 330 can be an optical feature flow estimation process modeled using a DNN. The flow compensation module 340 can be an optical flow compensation and coarse intermediate frame synthesis process modeled using a DNN. The detail enhancement module 350 can be an enhancement process modeled using a DNN.
[0054] Methods and apparatuses for a DNN model for video frame interpolation (VFI) according to embodiments will now be described in detail.
[0055] Figure 4 is a block diagram of a test apparatus for a virtual reference generation process 400 during a test phase according to embodiments.
[0056] As Figure 4 shown, the virtual reference generation process 400 includes an optical flow estimation and intermediate frame synthesis module 410 and a detail enhancement module 420.
[0057] In the random access configuration of the VVC decoder, two reference frames are fed into the optical flow estimation and intermediate frame synthesis module 410 to generate a flow map and then a coarse intermediate frame is generated to be fed into the detail enhancement module 420 together with the forward / backward reference frames to further improve the quality of the frame. More detailed descriptions of the DNN modules for performing the reference frame generation process, the optical flow estimation and intermediate frame synthesis module 410, and the detail enhancement module 420 will be detailed later with reference to Figure 5 and Figure 6
[0058] Depending on the prediction structure and / or the coding configuration, two or more different models can be selectively trained and used. In the random access configuration where a hierarchical prediction structure can be used for motion-compensated prediction with both forward and backward prediction reference pictures, one or more forward reference pictures that precede the current picture in output (or display) order and one or more backward reference pictures that follow the current picture in output (or display) order are inputs to the network. In the low-delay configuration where only forward prediction reference pictures can be used for motion-compensated prediction, two or more forward reference pictures are inputs to the network. For each configuration, one or more different network models can be selected and used for inter prediction.
[0059] Once one or more reference picture lists (RPLs) are constructed before processing the network inference, a check is made in both the encoder and the decoder for the presence of a suitable reference picture that can be used as input to the network inference. If there is a backward reference picture that has the same POC distance from the current picture as the POC distance of a forward reference picture, the trained model for the random access configuration can be selected and used in the network inference to generate an intermediate (virtual) reference picture. If there is no suitable backward reference picture, two forward reference pictures can be selected and used in the network inference, one of which has a POC distance that is twice the POC distance of the other.
[0060] When the network model for inter prediction is explicitly used by the coding configuration, one or more syntax elements in the high-level syntax structure, such as parameter sets or headers, can indicate which models are used for the current sequence, picture, or slice. In an implementation, all available network models in the current coded video sequence are listed in the sequence parameter set (SPS), and the selected model for each coded picture or slice is indicated by one or more syntax elements in the picture parameter set (PPS) or the picture / slice header.
[0061] The network topology and parameters can be explicitly specified in the specification document of a video coding standard, e.g., VVC, HEVC or AV1. In this case, a pre-defined network model can be used for the entire coded video sequence. If two or more network models are defined in the specification, one or more of the network models can be selected for each coded video sequence, coded picture or slice.
[0062] In the case that a customized network is used for each coded video sequence, picture or slice, the network topology and parameters can be explicitly signaled in a high-level syntax structure, e.g., in a parameter set or SEI message in the base bitstream or in a parameter set or SEI message in the metadata track in file format. In an embodiment, one or more network models with network topology and parameters are specified in one or more SEI messages, where these SEI messages can be inserted into coded video sequences with different activation ranges. Furthermore, later SEI messages in the coded video bitstream can update the previously activated SEI messages.
[0063] Figure 5 is a detailed block diagram of the optical flow estimation and intermediate frame synthesis module 410 during the test phase according to an embodiment.
[0064] As shown in Figure 5 , the optical flow estimation and intermediate frame synthesis module 410 includes a flow estimation module 510, a backward warping module 520 and a fusion processing module 530.
[0065] Two reference frames {I0, I1} are used as inputs to the flow estimation module 510, which approximates the intermediate flow {F t , F 0->t} from the perspective of the frame I 1->t that is intended to be synthesized. The flow estimation module 510 employs a coarse-to-fine strategy with gradually increasing resolution: it iteratively updates the flow field. Conceptually, according to the iteratively updated flow field, the corresponding pixels are moved from the two input frames to the same positions in the potential intermediate frame.
[0066] With the estimated intermediate flow {F 0->t , F 1->t}, the backward warping module 520 generates the coarse reconstructed frames or warped frames {I 0->t , I 1->t} by performing backward warping on the input frames {I0, I1}. The backward warping of the frames can be performed by inversely mapping and sampling the pixels of the potential intermediate frame to the input frames {I0, I1} to generate the warped frames {I 0->t , I 1->t}. The fusion processing module 530 fuses the input frames {I0, I1}, the warped frames {I 0->t , I1->t} and the estimated intermediate flow {F 0->t ,F 1->t} as inputs. The fusion processing module 530 estimates a fusion map and another residual map. The fusion map and the residual map are estimated maps that map the changes in features in the input. Then, the deformed frames are linearly combined according to the fusion result and added to the residual map to be reconstructed to obtain the intermediate frame ( Figure 5 referred to as the reference frame I t ).
[0067] Figure 6 is a detailed block diagram of the detail enhancement module 420 during the testing phase according to an embodiment.
[0068] As shown in Figure 6 , the detail enhancement module 420 includes a PCD alignment module 610, a TSA fusion module 620, and a reconstruction module 630.
[0069] Assuming the reference frame I t is the reference that all other frames will be aligned to use and two reference frames {I t-1 ,I t+1} together as inputs, the PCD (pyramid, cascade, and deformable convolution) alignment module 610 refines the (I t-1 ,I t ,I t+1 ) features of the reference frames to aligned features F t-对准 . Using the aligned features F t-对准 , the TSA (temporal and spatial attention) fusion module 620 provides weights to the feature maps to concentrate attention on important features for subsequent restoration and output predicted frames I p . More detailed descriptions of the PCD alignment module 610 and the TSA fusion module 620 will be described later with respect to Figure 7 and Figure 8 , respectively.
[0070] Then, the reconstruction module formats the final residual of the reference frame I p based on the predicted frame I t . Finally, the addition module 640 adds the formatted final residual to the reference frame I t to obtain the enhanced frame I t-增强 as the final output of the detail enhancement module 420.
[0071] Figure 7 is a detailed block diagram of the PCD alignment module 610 within the detail enhancement module 420 during the testing phase according to an embodiment. Figure 7 The modules with the same name and number convention in may be the same or one of a plurality of modules performing the described functions (e.g., the deformable convolution module 730).
[0072] like Figure 7 As shown, the PCD alignment module 610 includes a feature extraction module 710, an offset generation module 720, and a deformable convolution module 730. The PCD alignment module 610 calculates the alignment features F between frames. t-对准 .
[0073] First, use {I0,I1…I t ,I t+1 Each reference frame in …} is used as input, and the feature extraction module 710 computes a feature map {F0, F1…F…} with three different levels using a feature extraction DNN through forward inference. t ,F t+1 …}. For feature compensation, different levels have different resolutions to capture different levels of spatial / temporal information. For example, with larger motion in the sequence, smaller feature maps will be able to have larger corresponding fields and can handle various object offsets.
[0074] For feature maps of different levels, the offset generation module 720 connects the feature maps {F... t ,F t+1 The connected feature maps are then used to generate an offset map ΔP via an offset-generating DNN. The deformable convolution module 730 then adds the offset map ΔP to the original position P using a temporally deformable convolution (TDC) operation. 原始 (that is, ΔP+P) 原始 To compute the DNN convolution kernel P 新 The new location. Note that reference frame I... t It can be {I0, I1…I t ,I t+1 Any frame within …}. Without loss of generality, frames can be arranged in emphasis order based on their timestamps. In one implementation, when the current goal is to enhance the current reconstructed frame I' t At that time, reference frame I t It is used as an intermediate frame.
[0075] Due to the new location P 新 The location may be irregular and may not be an integer, therefore a TDC operation can be performed using interpolation (e.g., bilinear interpolation). By applying a deformable convolution kernel in the deformable convolution module 730, compensation features from different levels can be generated based on the feature maps of the levels and the corresponding generated offsets. The deformable convolution kernel can then be applied based on the offset map ΔP and one or more of the compensation features by upsampling and adding them to the upper-level compensation features to obtain the top-level alignment feature F. t-对准 .
[0076] Figure 8is a detailed block diagram of the TSA fusion module 620 within the detail enhancement module 420 during the testing phase according to an embodiment.
[0077] As shown in Figure 8 , the TSA fusion module 620 includes an activation module 810, an element-wise multiplication module 820, a fusion convolution module 830, a frame reconstruction module 840, and a frame synthesis module 850. The TSA fusion module 620 uses temporal and spatial attention. The goal of temporal attention is to compute the frame similarity in the embedding space. Intuitively, in the embedding space, more attention should be paid to the neighboring frames that are more similar to the reference frame. Given a feature map {F t-1 ,F t+1} and the aligned feature F t-对准 , the activation module 810 as input uses a sigmoid activation function to limit the input to [0, 1] to obtain the temporal attention map {F t-1 ,F t ,F t+1} as the output of the activation module 810. Note that for each spatial location, the temporal attention is spatial-specific.
[0078] Then, the element-wise multiplication module 820 multiplies the temporal attention map {F t-1 ,F t ,F t+1} with the aligned feature F t-对准 in a pixel-wise manner. An additional fusion convolution layer is employed in the fusion convolution module 830 to obtain the aggregated attention modulated feature F t-对准-TSA . Using the temporal attention map {F t-1 ,F t ,F t+1} and the attention modulated feature F t-对准-TSA , the frame reconstruction module 840 computes the aligned frame I t-对准 using a frame reconstruction DNN by feed-forward inference. Then, the aligned frame I t-对准 is passed through the frame synthesis module 850 to generate the synthesized predicted frame I p as the final output.
[0079] Figure 9 is a detailed block diagram of the TSA fusion module 620 within the detail enhancement module 420 during the testing phase according to another example embodiment. Figure 9 Modules with the same name and number convention in may be the same or one of multiple modules performing the described function.
[0080] As shown in Figure 9As shown, the TSA fusion module 620 according to the present embodiment includes an activation module 910, an element-wise multiplication module 920, a fusion convolution module 930, a down-sampling convolution (DSC) module 940, an up-sampling and addition module 950, a frame reconstruction module 960, and a frame synthesis module 970. Like the embodiments of Figure 8 , the TSA fusion module 620 uses temporal and spatial attention. Given a feature map {F t-1 ,F t+1} and an aligned feature F t-对准 of a center frame as inputs, the activation module 910 uses a sigmoid activation function to limit the inputs into [0, 1] as a simple convolution filter to obtain temporal attention maps {M t-1 ,M t ,M t+1} as outputs of the activation module 910.
[0081] Then, the element-wise multiplication module 920 pixel-wise multiplies the temporal attention maps {M t-1 ,M t ,M t+1} with the aligned feature F t-对准 . A fusion convolution is performed on the products in the fusion convolution module 930 to generate a fusion feature F 融合 . The fusion feature F 融合 is down-sampled and convolved by the DSC module 940. As Figure 9 shown, another convolution layer can be employed and processed by the DSC module 940. The output from each layer of the DSC module 940 is input to the up-sampling and addition module 950. As Figure 9 shown, additional layers of up-sampling and addition are applied to the additional layers with the fusion feature F 融合 as input. The up-sampling and addition module 950 generates a fusion attention map M t-融合 . The element-wise multiplication module 920 pixel-wise multiplies the fusion attention map M t-融合 with the fusion feature F 融合 to generate a TSA aligned feature F t-TSA . Using the temporal attention maps {M t-1 ,M t ,M t+1} and the TSA aligned feature F t-TSA , the frame reconstruction module 960 computes a TSA aligned frame I t-TSA using a frame reconstruction DNN by a feed-forward inference. Then, the aligned frame I t-TSA is passed through the frame synthesis module 970 to generate a synthesized predicted frame I p as a final output.
[0082] Figure 10is a flowchart of a method 1000 of video coding using neural network based inter prediction according to an embodiment.
[0083] In some implementations, Figure 10 One or more of the processing blocks can be performed by the platform 120. In some implementations, Figure 10 One or more of the processing blocks can be performed by another device or set of devices separate from or including the platform 120, such as the user device 110.
[0084] As Figure 10 shown, in operation 1001, the method 1000 includes generating an intermediate stream based on input frames. The intermediate stream can be further iteratively updated and corresponding pixels move from two input frames to the same position in a potential intermediate frame.
[0085] In operation 1002, the method 1000 includes generating a reconstructed frame by performing backward warping of the input frames with the intermediate stream.
[0086] In operation 1003, the method 1000 includes generating a blending map and a residual map based on the input frames, the intermediate stream, and the reconstructed frame.
[0087] In operation 1004, the method 1000 includes generating a feature map having a plurality of levels using a first neural network based on a current reference frame, a first reference frame, and a second reference frame. The current reference frame can be generated by linearly combining the reconstructed frame according to the blending map and adding the combined reconstructed frame to the residual map. Further, the first reference frame can be a reference frame preceding the current reference frame in output order, and the second reference frame can be a reference frame succeeding the current reference frame in output order.
[0088] In operation 1005, the method 1000 includes generating a predicted frame based on aligned features from the generated feature map by refining the current reference frame, the first reference frame, and the second reference frame. Specifically, the predicted frame can be generated by first performing a convolution to obtain an attention map, generating attention features based on the attention map and the aligned features, generating an aligned frame based on the attention map and the attention features, and then synthesizing the aligned frame to obtain the predicted frame.
[0089] The aligned features can be generated by computing offsets to the plurality of levels of the feature map generated in operation 1004 and performing deformable convolution to generate compensation features for the plurality of levels. The aligned features can then be generated based on at least one of the offsets and the generated compensation features.
[0090] In operation 1006, the method 1000 includes generating a final residual based on the predicted frame. Weights of the feature map generated in operation 1004 can also be generated to emphasize important features for generating a subsequent final residual.
[0091] At operation 1007, the method 1000 includes computing an enhanced frame as output by adding the final residual to the current reference frame.
[0092] Although Figure 10 Example blocks of the method are shown, but in some implementations, the method can include additional blocks, fewer blocks, different blocks, or differently arranged blocks than those depicted in Figure 10 the blocks depicted in FIG. 10. Additionally or alternatively, two or more of the blocks of the method can be performed in parallel.
[0093] Figure 11 is a block diagram of a device 1100 for video coding using neural network based inter prediction according to an embodiment.
[0094] As Figure 11 shown, the device includes a first generating code 1101, a second generating code 1102, a configured fusing code 1103, a third generating code 1104, a predicting code 1105, a residual code 1106, and a fourth generating code 1107.
[0095] The first generating code 1101 is configured to cause the at least one processor to generate an intermediate stream based on an input frame.
[0096] The second generating code 1102 is configured to cause the at least one processor to perform backward warping of the input frame with the intermediate stream to generate a reconstructed frame.
[0097] The configured fusing code 1103 is configured to cause the at least one processor to generate a fusion map and a residual map based on the input frame, the intermediate stream, and the reconstructed frame.
[0098] The third generating code 1104 is configured to cause the at least one processor to generate a feature map having a plurality of levels using a first neural network based on a current reference frame, a first reference frame, and a second reference frame.
[0099] The predicting code 1105 is configured to cause the at least one processor to predict a frame based on aligned features from the generated feature map by refining the current reference frame, the first reference frame, and the second reference frame.
[0100] The residual code 1106 is configured to cause the at least one processor to generate a final residual based on the predicted frame.
[0101] The fourth generating code 1107 is configured to cause the at least one processor to generate an enhanced frame as output by adding the final residual to the current reference frame.
[0102] Although Figure 11 Example blocks of the device are shown, but in some implementations, the device can include additional blocks, fewer blocks, different blocks, or differently arranged blocks than those depicted in Figure 11Outside the depicted boxes, the device can include additional boxes, fewer boxes, different boxes, or differently arranged boxes. Additionally or alternatively, two or more of the boxes of the device can be combined.
[0103] The device can further include update code configured to cause the at least one processor to iteratively update the intermediate stream and move corresponding pixels from the two input frames to the same location in the potential intermediate frame, reference frame code configured to cause the at least one processor to generate a current reference frame by linearly combining the reconstructed frames according to the fusion map and adding the combined reconstructed frames to the residual map, determination code configured to cause the at least one processor to determine weights of the feature maps, where the weights emphasize important features for generating subsequent final residuals, calculation code configured to cause the at least one processor to calculate offsets for a plurality of levels, compensation feature generation code configured to cause the at least one processor to perform deformable convolution to generate compensation features for the plurality of levels, alignment feature generation code configured to cause the at least one processor to generate alignment features based on at least one of the generated compensation features and the offsets, execution code configured to cause the at least one processor to perform convolution to obtain an attention map, attention feature generation code configured to cause the at least one processor to generate attention features based on the attention map and the alignment features, alignment frame generation code configured to cause the at least one processor to generate an alignment frame using a second neural network based on the attention map and the attention features, and synthesis code configured to cause the at least one processor to synthesize the alignment frame to obtain the predicted frame.
[0104] Compared with traditional inter-frame generation methods, the proposed method performs DNN-based network in video coding. The proposed method directly obtains data from a reference picture list (RPL) to generate a virtual reference frame, instead of calculating explicit motion vectors or motion streams that cannot handle complex motion or are prone to errors. And then, enhanced deformable convolution (DCN) is applied to capture pixel offsets and implicitly compensate for large complex motion for further detail enhancement. Finally, a high-quality enhanced frame is reconstructed by a DNN model.
[0105] The proposed methods can be used individually or in any combination. Furthermore, each of the methods (or implementations) can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program that is stored in a non-transitory computer-readable medium.
[0106] The present disclosure provides illustration and description, but is not intended to be exhaustive or to limit implementations to the precise form disclosed. Modifications and variations can be made in light of this disclosure, or can be acquired from practice of the implementations.
[0107] As used herein, the term component is intended to be broadly interpreted to include hardware, firmware, or a combination of hardware and software.
[0108] It will be apparent that systems and / or methods described herein can be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual specialized control hardware or software code used to implement these systems and / or methods is not limiting of the implementations. Thus, the operation and behavior of the systems and / or methods were described herein without reference to specific software code — it being understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0109] Even though combinations of features are recited in the claims and / or described herein, these combinations are not limited to cases where such features are used in combination. The combinations are intended to be freely combinable and can be used in various permutations. Although each of the dependent claims below can be directly dependent on only one claim, combinations of the dependent claims are also contemplated as being within the scope of the present disclosure.
[0110] No element, act, or instruction used herein should be construed as critical or essential unless explicitly described as such. Also, as used herein, the articles “a” and “an” are intended to include one or more items, and can be used interchangeably with “one or more.” Furthermore, as used herein, the term “set” is intended to include one or more items (e.g., related items, unrelated items, a combination of related and unrelated items, etc.), and can be used interchangeably with “one or more.” Where only one item is intended, the term “one” or similar language is used. Also, as used herein, the terms “has,” “have,” “having,” or the like are intended to be open-ended terms. Further, the phrase “based on” is intended to mean “based, at least in part, on” unless explicitly stated otherwise.
Claims
1. A method for video coding using inter-frame prediction based on neural networks, the method being executed by at least one processor, characterized in that, The method includes: Generate a list of reference images that includes multiple reference images; Based on the POC distance counted in the image sequence, it is determined that the first reference image included in the reference image list exists; Based on the determination that the first reference image exists, select one or more different network models; An intermediate flow is generated from the perspective of the first reference image, and the flow field of the intermediate flow is iteratively updated based on the first reference image. Based on the selected one or more different network models and the intermediate flow, an intermediate reference image is generated; By refining the intermediate reference images, images are predicted based on weights in the feature maps, which include multiple levels; and An enhanced image is calculated as the output by adding the final residual based on the predicted image to the first reference frame.
2. The method according to claim 1, characterized in that, The flow field of the intermediate flow is iteratively modified by moving corresponding pixels from two input reference images from the plurality of reference images to the same position in a potential intermediate image.
3. The method according to claim 1, characterized in that, The method further includes: A reconstructed frame is generated by performing a backward warp of the input reference image using the intermediate stream; Based on the reconstructed image, a fusion image and a residual image are generated; A current reference image is generated by linearly combining the reconstructed images according to the fusion map and adding the combined reconstructed images to the residual map; and The feature map is generated using a first neural network based on the current reference image and one or more of the plurality of reference images.
4. The method according to claim 1, characterized in that, The first reference image includes a first image and a second image, wherein the first image is a reference frame preceding the current reference frame in the output order, and the second reference image is a reference frame following the current reference frame in the output order.
5. The method according to claim 1, characterized in that, The weights in the feature map emphasize the subset of features in the feature map used to generate the subsequent final residual.
6. The method according to claim 1, characterized in that, The method further includes: Calculate the offsets for the multiple levels; Perform deformable convolutions to generate compensated features for the multiple levels; and The alignment feature is generated based on at least one of the generated compensation features and the offset.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Perform convolutions to obtain a fused attention map; Attention features are generated based on the fused attention map and the alignment features; Based on the fused attention map and the attention features, a second neural network is used to generate alignment frames; and The aligned frames are synthesized to obtain the predicted image.
8. A method for storing a bit stream, characterized in that, The method includes: generating a bit stream by performing the method of any one of claims 1-7; and storing the bit stream.
9. A method for transmitting a bit stream, characterized in that, The method includes: generating a bit stream by performing the method of any one of claims 1-7; and transmitting the bit stream.
10. An apparatus for video coding using inter-frame prediction based on neural networks, characterized in that, The device includes: The first generation unit is configured to generate a list of reference images that includes multiple reference images; A determining unit is configured to determine the existence of a first reference image included in the reference image list based on the POC distance counted in the image sequence. The selection unit is configured to select one or more different network models based on determining that the first reference image exists; The second generation unit is configured to generate an intermediate flow from the perspective of the first reference image and iteratively update the flow field of the intermediate flow based on the first reference image. The third generation unit is configured to generate intermediate reference images based on one or more different network models selected and the intermediate stream; A prediction unit, configured to predict an image based on weights in a feature map, which includes multiple levels, by refining the intermediate reference image; and A computing unit is configured to compute an enhanced image as output by adding the final residual based on the predicted image to the first reference frame.
11. A computer device, characterized in that, The computer device includes: One or more computer-readable non-transitory storage media configured to store computer program code; and One or more computer processors configured to access the computer program code and execute the method according to any one of claims 1 to 9 as instructed by the computer program code.
12. A non-transitory computer-readable medium storing instructions, characterized in that, The instructions, when executed by at least one processor for video coding using neural network-based inter-frame prediction, cause the at least one processor to perform the method according to any one of claims 1 to 9.
13. A computer-readable storage medium storing a computer program / instructions and a bit stream thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1-9 to generate the bit stream.