Method, computer device, equipment and computer-readable medium for video encoding
Through the inter-frame prediction method based on deep neural network, virtual reference data is generated, and efficiency and quality problems in complex motion scenarios in video encoding are solved, achieving more efficient video compression and lower artifact effects.
Patent Information
- Application Number
- CN202180030747.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-09-16
- Filing Date
- 2021-09-24
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-09-24
AI Technical Summary
The existing video encoding standards are difficult to effectively handle complex motion scenarios, resulting in low video encoding efficiency. Especially in the case of non-translational motion, traditional methods cannot accurately capture pixel-level motion information, and the distortion degree during lossy compression is difficult to control.
Using an inter-frame prediction method based on deep neural networks, the generation of intermediate streams, reconstruction frames, fusion maps and residual maps, and prediction is made using multi-level feature maps. Combined with optical flow estimation and detail enhancement modules, virtual reference data is generated to improve frame quality and reduce motion boundary artifacts.
It improves the compression efficiency and quality of video encoding, reduces artifacts on motion boundaries, enhances the accuracy and compression ratio of inter-frame prediction, adapts to complex motion scenarios, and reduces the distortion of lossy compression.
Smart Images

Figure CN115486068B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 131,625, filed on December 29, 2020, and U.S. Patent Application No. 17 / 476,928, filed on September 16, 2021. The disclosures of the said applications are incorporated herein by reference in their entireties. Technical field
[0003] The present disclosure relates to a method of video coding, a computer device, an apparatus, and a computer - readable medium. Background art
[0004] An uncompressed digital video can include a series of pictures, each picture having spatial dimensions such as 1920×1080 luminance samples and associated chrominance samples. The series of pictures can have a fixed or variable picture rate (also informally called frame rate), e.g., 60 pictures per second or 60 Hz. Uncompressed video has significant bit - rate requirements. For example, a 1080p60 4:2:0 video (1920×1080 luminance sample resolution at 60 Hz frame rate) with 8 bits per sample requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires more than 600 gigabytes (GByte) of storage space.
[0005] Traditional video coding standards such as H.264 / Advanced Video Coding (H.264 / AVC), High - Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC) share a similar (recursive) block - based hybrid prediction / transformation framework, where individual coding tools such as intra - frame / inter - frame prediction, integer transformation, and context - adaptive entropy coding are centrally hand - crafted to optimize overall efficiency. Summary of the invention
[0006] According to an embodiment, there is provided a method of video coding using neural - network - based inter - frame prediction, the method being executed by at least one processor, and the method includes: generating an intermediate stream based on an input frame; generating a reconstructed frame by performing backward warping of the input frame using the intermediate stream; generating a fusion map and a residual map based on the input frame, the intermediate stream, and the reconstructed frame; generating a feature map with multiple levels using a first neural network based on a current reference frame, a first reference frame, and a second reference frame; generating a predicted frame based on aligned features from the generated feature map by refining the current reference frame, the first reference frame, and the second reference frame; generating a final residual based on the predicted frame; and calculating an enhanced frame as an output by adding the final residual to the current reference frame.
[0007] According to an embodiment, a computer device is provided, characterized in that the computer device includes: one or more computer-readable non-transitory storage media configured to store computer program code; and one or more computer processors configured to access the computer program code and execute the method for video coding using inter-frame prediction based on a neural network as described above according to the instructions of the computer program code.
[0008] According to an embodiment, a device for video coding using inter-frame prediction based on a neural network is provided, characterized in that the device includes: a first generation unit configured to generate an intermediate stream based on an input frame; a second generation unit configured to perform backward warping of the input frame using the intermediate stream to generate a reconstructed frame; a fusion unit configured to generate a fusion map and a residual map based on the input frame, the intermediate stream, and the reconstructed frame; a third generation unit configured to generate feature maps with multiple levels using a first neural network based on a current reference frame, a first reference frame, and a second reference frame; a prediction unit configured to predict a frame based on aligned features from the generated feature maps by refining the current reference frame, the first reference frame, and the second reference frame; a residual unit configured to generate a final residual based on the predicted frame; and a fourth generation unit configured to generate an enhanced frame as an output by adding the final residual to the current reference frame.
[0009] According to an embodiment, a non-transitory computer-readable medium storing instructions is provided, and the instructions, when executed by at least one processor for video coding using inter-frame prediction based on a neural network, cause the at least one processor to execute the method for video coding using inter-frame prediction based on a neural network as described above.
[0010] The method, computer device, device, and computer-readable medium for video coding of the present invention describe a deep neural network (DNN)-based model related to video coding and decoding. More specifically, as a video frame interpolation (VFI) task and a detail enhancement module, the DNN-based model uses and generates virtual reference data based on adjacent reference frames for inter-frame prediction to further improve the quality of frames and reduce artifacts (such as noise, blur, blockiness, etc.) on motion boundaries. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is a diagram of an environment in which the methods, devices, and systems described herein can be implemented according to an embodiment.
[0012] Figure 2 is Figure 1 a block diagram of example components of one or more devices.
[0013] Figure 3 It is a schematic diagram showing an example in which a virtual reference picture is generated and inserted into a reference picture list.
[0014] Figure 4 It is a block diagram of a test device for a virtual reference generation process during a test phase according to an embodiment.
[0015] Figure 5 It is from during the test phase according to an embodiment Figure 4 of a detailed block diagram of an optical flow estimation and intermediate frame synthesis module of a test device.
[0016] Figure 6 It is from during the test phase according to an embodiment Figure 4 of a detailed block diagram of a detail enhancement module of a test device.
[0017] Figure 7 It is a detailed block diagram of a PCD alignment module during a test phase according to an embodiment.
[0018] Figure 8 It is a detailed block diagram of a TSA fusion module during a test phase according to an embodiment.
[0019] Figure 9 It is a detailed block diagram of a TSA fusion module during a test phase according to another embodiment.
[0020] Figure 10 It is a flowchart of a method for video coding using neural network-based inter-frame prediction according to an embodiment.
[0021] Figure 11 It is a block diagram of a device for video coding using neural network-based inter-frame prediction according to an embodiment. Detailed Embodiments
[0022] The present disclosure describes a deep neural network (DNN)-based model related to video coding and decoding. More specifically, as a video frame interpolation (VFI) task and a detail enhancement module, the DNN-based model uses and generates virtual reference data based on adjacent reference frames for inter-frame prediction to further improve the quality of frames and reduce artifacts (such as noise, blur, blockiness, etc.) on motion boundaries.
[0023] One purpose of video encoding and decoding can be to reduce redundancy in an input video signal through compression. Compression can help reduce the above bandwidth or storage space requirements, in some cases by up to two orders of magnitude or more. Both lossless compression and lossy compression, as well as combinations thereof, can be employed. Lossless compression refers to techniques in which an exact copy of the original signal can be reconstructed from the compressed signal. When lossy compression is used, the reconstructed signal may be different from the original signal, but the distortion between the original signal and the reconstructed signal is small enough such that the reconstructed signal is useful for the intended application. Lossy compression is widely employed in the case of video. The amount of distortion tolerated depends on the application; for example, users of certain consumer streaming applications may tolerate higher distortion than users of television contribution applications. The achievable compression ratio can reflect that higher allowable / tolerable distortion can result in a higher compression ratio.
[0024] Basically, spatio-temporal pixel neighborhoods are used for predicting signal structure to obtain corresponding residuals for subsequent transformation, quantization, and entropy coding. On the other hand, the essence of a deep neural network (DNN) is to extract different levels of spatio-temporal stimuli by analyzing spatio-temporal information from the receptive fields of adjacent pixels. The ability to highly explore non-linear and non-local spatio-temporal correlations provides promising opportunities for significantly improving compression quality.
[0025] One caveat in using information from multiple adjacent video frames is the complex motion caused by moving cameras and dynamic scenes. Traditional block-based motion vectors do not work well for non-translational motion. Learning-based optical flow methods can provide pixel-level accurate motion information, unfortunately, which is prone to errors, especially at the boundaries of moving objects. The present disclosure proposes using a DNN-based model to implicitly handle any complex motion in a data-driven manner without explicit motion estimation.
[0026] Figure 1 FIG. 100 is a diagram of an environment 100 in which the methods, devices, and systems described herein can be implemented according to an embodiment.
[0027] As Figure 1 shown, the environment 100 can include a user device 110, a platform 120, and a network 130. The devices in the environment 100 can be interconnected via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection.
[0028] The user device 110 includes one or more devices capable of receiving, generating, storing, processing, and / or providing information associated with the platform 120. For example, the user device 110 may include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smartphone, a wireless phone, etc.), a wearable device (e.g., a pair of smart glasses or a smart watch), or a similar device. In some implementations, the user device 110 may receive information from and / or send information to the platform 120.
[0029] The platform 120 includes one or more devices as described elsewhere herein. In some implementations, the platform 120 may include a cloud server or a group of cloud servers. In some implementations, the platform 120 may be designed to be modular such that software components can be swapped in or out. Thus, the platform 120 can be easily and / or quickly reconfigured for different uses.
[0030] In some implementations, as shown, the platform 120 may be hosted in a cloud computing environment 122. It is noted that while the implementations described herein describe the platform 120 as being hosted in the cloud computing environment 122, in some implementations, the platform 120 may not be cloud-based (i.e., may be implemented outside of a cloud computing environment) or may be partially cloud-based.
[0031] The cloud computing environment 122 includes an environment that hosts the platform 120. The cloud computing environment 122 may provide services such as computing, software, data access, storage, etc., which do not require an end user (e.g., the user device 110) to know the physical location and configuration of one or more systems and / or one or more devices hosting the platform 120. As shown, the cloud computing environment 122 may include a set of computing resources 124 (collectively referred to as "computing resources 124" and individually referred to as "computing resource 124").
[0032] The computing resources 124 include one or more personal computers, workstation computers, server devices, or other types of computing and / or communication devices. In some implementations, the computing resources 124 may host the platform 120. Cloud resources may include: computing instances executed in the computing resources 124, storage devices provided in the computing resources 124, data transmission devices provided by the computing resources 124, etc. In some implementations, the computing resources 124 may communicate with other computing resources 124 via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection.
[0033] Further asFigure 1 As shown, computing resources 124 include a set of cloud resources, such as one or more applications ("APP") 124-1, one or more virtual machines ("VM") 124-2, virtualized storage ("VS") 124-3, one or more hypervisors ("HYP") 124-4, and so on.
[0034] Application 124-1 includes one or more software applications that can be provided to and / or accessed by user device 110 and / or platform 120. Application 124-1 can eliminate the need to install and execute software applications on user device 110. For example, application 124-1 can include software associated with platform 120 and / or any other software that can be provided via cloud computing environment 122. In some implementations, one application 124-1 can send information to / receive information from one or more other applications 124-1 via virtual machine 124-2.
[0035] Virtual machine 124-2 includes a software implementation of a machine (e.g., a computer) that executes programs similar to a physical machine. Virtual machine 124-2 can be a system virtual machine or a process virtual machine, depending on the usage and correspondence to any real machine through virtual machine 124-2. A system virtual machine can provide a complete system platform that supports the execution of a full operating system ("OS"). A process virtual machine can execute a single program and can support a single process. In some implementations, virtual machine 124-2 can execute on behalf of a user (e.g., user device 110) and can manage the infrastructure of cloud computing environment 122, such as data management, synchronization, or long-duration data transfer.
[0036] Virtualized storage 124-3 includes one or more storage systems and / or one or more devices that use virtualization technology within the storage system or device of computing resources 124. In some implementations, in the case of a storage system, the types of virtualization can include block virtualization and file virtualization. Block virtualization can refer to the extraction (or separation) of logical storage from physical storage, such that the storage system can be accessed without considering the physical storage or heterogeneous structure. The separation can allow the administrator of the storage system flexibility in managing storage for end users. File virtualization can eliminate the dependence between the data accessed at the file level and the location where the file is physically stored. This can enable optimization of storage usage, server consolidation, and / or the performance of uninterrupted file migration.
[0037] The hypervisor 124-4 may provide hardware virtualization techniques that enable multiple operating systems (e.g., “guest operating systems”) to execute simultaneously on a host computer such as computing resource 124. The hypervisor 124-4 may present a virtual operating platform to the guest operating systems and may manage the execution of the guest operating systems. Multiple instances of various operating systems may share the virtualized hardware resources.
[0038] Network 130 includes one or more wired and / or wireless networks. For example, network 130 may include a cellular network (e.g., a fifth-generation (5G) network, a long-term evolution (LTE) network, a third-generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a public switched telephone network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-based network, etc., and / or a combination of these or other types of networks.
[0039] There is provided Figure 1 The number and arrangement of the devices and networks shown are by way of example. In fact, in addition to Figure 1 the devices and / or networks shown, there may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks in a different arrangement. Additionally, Figure 1 two or more of the devices shown may be implemented within a single device, or Figure 1 a single device shown may be implemented as multiple distributed devices. Additionally or alternatively, a set of devices (e.g., one or more devices) of environment 100 may perform one or more functions described as being performed by another set of devices of environment 100.
[0040] Figure 2 is Figure 1 a block diagram of example components of one or more of the devices.
[0041] Device 200 may correspond to user device 110 and / or platform 120. As Figure 2 shown, device 200 may include a bus 210, a processor 220, a memory 230, a storage component 240, an input component 250, an output component 260, and a communication interface 270.
[0042] The bus 210 includes components that permit communication among the components of the apparatus 200. The processor 220 is implemented in hardware, firmware, or a combination of hardware and software. The processor 220 is a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or another type of processing component. In some implementations, the processor 220 includes one or more processors that can be programmed to perform functions. The memory 230 includes random access memory (RAM), read only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, and / or optical memory) that stores information and / or instructions for use by the processor 220.
[0043] The storage component 240 stores information and / or software related to the operation and use of the apparatus 200. For example, the storage component 240 may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid state disk), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cassette tape, a magnetic tape, and / or another type of non-transitory computer-readable medium, as well as a corresponding drive.
[0044] The input component 250 includes components that permit the apparatus 200 to receive information, such as via a user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, and / or a microphone). Additionally or alternatively, the input component 250 may include sensors for sensing information (e.g., a global positioning system (GPS) component, an accelerometer, a gyroscope, and / or an actuator). The output component 260 includes components that provide output information from the apparatus 200 (e.g., a display, a speaker, and / or one or more light emitting diodes (LEDs)).
[0045] The communication interface 270 includes transceiver-like components (e.g., a transceiver and / or separate receiver and transmitter) that enable the apparatus 200 to communicate with other devices, such as via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection. The communication interface 270 may permit the apparatus 200 to receive information from and / or provide information to another device. For example, the communication interface 270 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, etc.
[0046] The apparatus 200 may perform one or more processes described herein. The apparatus 200 may perform these processes in response to a processor 220 executing software instructions stored by a non-transitory computer-readable medium such as a memory 230 and / or a storage component 240. The computer-readable medium is defined herein as a non-transitory memory device. The memory device includes a memory space within a single physical storage device or a memory space distributed over multiple physical storage devices.
[0047] The software instructions may be read into the memory 230 and / or the storage component 240 from another computer-readable medium or from another device via a communication interface 270. When the software instructions stored in the memory 230 and / or the storage component 240 are executed, the software instructions may cause the processor 220 to perform one or more processes described herein. Additionally or alternatively, hardwired circuitry may be used in place of or in combination with the software instructions to perform one or more processes described herein. Accordingly, the implementations described herein are not limited to any particular combination of hardware circuitry and software.
[0048] There is provided Figure 2 The number and arrangement of the components shown as an example. In fact, in addition to Figure 2 the components shown, the apparatus 200 may include additional components, fewer components, different components, or components arranged differently. Additionally or alternatively, a group of components (e.g., one or more components) of the apparatus 200 may perform one or more functions described as being performed by another group of components of the apparatus 200.
[0049] A typical video compression framework may be described as follows. In a first motion estimation step, given an input video x including a plurality of image frames x1, …, x T , the frames are partitioned into spatial blocks. Each block may be iteratively partitioned into smaller blocks, and a set of motion vectors m t between the current frame x and a set of previously reconstructed frames t is calculated for each block. Note that the subscript t represents the current t-th coding cycle, which may not match the timestamp of the image frame. Additionally, the set of previously reconstructed frames includes frames from multiple previous coding cycles. Then, in a second motion compensation step, a predicted frame t is obtained by copying corresponding pixels of the set of previously reconstructed frames based on the motion vectors m and a residual r t between the original frame x and the predicted frame t may be obtained (i.e., ). In a third motion compensation step, the residual rt Quantization is performed. The quantization step gives the quantized The motion vector m is encoded into a bitstream through entropy coding t and the quantized Both are encoded into a bitstream, and this bitstream is sent to the decoder. Then, on the decoder side, the quantized is dequantized (usually through an inverse transform similar to IDCT using dequantization coefficients) to obtain the recovered residual Then the recovered residual is added back to the prediction frame to obtain the reconstructed frame (i.e., ). Additional components are further used to improve the visual quality of the reconstructed frame One or more of the following enhancement modules can be selected to process the reconstructed frame The enhancement module includes a deblocking filter (DF), sample adaptive offset (SAO), adaptive loop filter (ALF), etc.
[0050] In HEVC, VVC or other video coding frameworks or standards, the decoded picture can be included in the reference picture list (RPL) and can be used as a reference picture for motion compensation prediction and other parameter predictions for encoding one or more pictures later in the encoding or decoding order, or can be used for intra prediction or block copy to encode different regions or blocks of the current picture.
[0051] In an embodiment, one or more virtual references can be generated in both the encoder and the decoder or only in the decoder and included in the RPL. The virtual reference picture can be generated through one or more processes, including signal processing, spatial or temporal filtering, scaling, weighted average, up / down sampling, pooling, memory recursive processing, linear system processing, nonlinear system processing, neural network processing, deep learning-based processing, artificial intelligence processing, pre-trained network processing, machine learning-based processing, online training network processing, or a combination thereof. For the process of generating one or more virtual references, zero or more forward reference pictures both before the current picture in the output / display order and in the encoding / decoding order and zero or more backward reference pictures both after the current picture in the output / display order but before the current picture in the encoding / decoding order are used as input data. The output of the process is the virtual / generated picture to be used as a new reference picture.
[0052] A DNN pre-trained network process for virtual reference picture generation is described according to an embodiment. Figure 3An example of virtual reference picture generation and insertion into a reference picture list is shown, including a hierarchical GOP structure 300, a reference picture list 310, and a virtual reference generation process 320.
[0053] As Figure 3 shown, considering the hierarchical GOP structure 300, when the picture order count (POC) of the current picture is equal to 3, generally, the decoded pictures with POC equal to 0, 2, 4, or 8 can be stored in the decoded picture buffer, and some of the decoded pictures are included in the reference picture list for decoding the current picture. As an example, the most recently decoded pictures with POC equal to 2 or 4 can be fed as input data into the virtual reference generation process 320. The virtual reference picture can be generated through one or more processes. The generated virtual reference picture can be stored in the decoded picture buffer and included in the reference picture list 310 of the current picture or one or more future pictures in decoding order. If the virtual reference picture is included in the reference picture list 310 of the current picture, when the virtual reference picture is indicated to be used through a reference index, the pixel data of the generated virtual reference picture can be used as reference data for motion compensation prediction.
[0054] In the same or another embodiment, the entire virtual reference generation process can include one of more signaling processing modules using one or more pre-trained neural network models or any predefined parameters. For example, as Figure 3 shown, the entire virtual reference generation process 320 can include an optical flow estimation module 330, an optical flow compensation module 340, and a detail enhancement module 350. The optical flow estimation module 330 can be an optical feature flow estimation process modeled using DNN. The optical flow compensation module 340 can be an optical flow compensation and rough intermediate frame synthesis process modeled using DNN. The detail enhancement module 350 can be an enhancement process modeled using DNN.
[0055] A method and apparatus for a DNN model for video frame interpolation (VFI) according to an embodiment will now be described in detail.
[0056] Figure 4 is a block diagram of a test apparatus for a virtual reference generation process 400 during a test phase according to an embodiment.
[0057] As Figure 4 shown, the virtual reference generation process 400 includes an optical flow estimation and intermediate frame synthesis module 410 and a detail enhancement module 420.
[0058] In the random access configuration of the VVC decoder, two reference frames are fed into the optical flow estimation and intermediate frame synthesis module 410 to generate a flow graph, and then a rough intermediate frame is generated and fed into the detail enhancement module 420 together with the forward / backward reference frames to further improve the quality of the frame. A more detailed description of the DNN module, the optical flow estimation and intermediate frame synthesis module 410, and the detail enhancement module 420 for performing the reference frame generation process will be described later with reference to Figure 5 and Figure 6 respectively.
[0059] Depending on the prediction structure and / or coding configuration, two or more different models can be selectively trained and used. In a random access configuration where both forward predicted reference pictures and backward predicted reference pictures can be used with a hierarchical prediction structure for motion compensation prediction, one or more forward reference pictures before the current picture in output (or display) order and one or more backward reference pictures after the current picture in output (or display) order are inputs to the network. In a low latency configuration where only forward predicted reference pictures can be used for motion compensation prediction, two or more forward reference pictures are inputs to the network. For each configuration, one or more different network models can be selected and used for inter-frame prediction.
[0060] Once one or more reference picture lists (RPLs) are constructed before processing network inference, the presence of appropriate reference pictures that can be used as inputs to network inference is checked in both the encoder and decoder. If there is a backward reference picture with the same POC distance from the current picture as the POC distance of the forward reference picture, the training model for the random access configuration can be selected and used in network inference to generate an intermediate (virtual) reference picture. If there is no appropriate backward reference picture, two forward reference pictures can be selected and used in network inference, where the POC distance of one of the two forward reference pictures is twice that of the other.
[0061] When the network model for inter-frame prediction explicitly uses the coding configuration, one or more syntax elements in the high-level syntax structure, such as parameter sets or headers, can indicate which models are used for the current sequence, picture, or slice. In an embodiment, all available network models in the current encoded video sequence are listed in the sequence parameter set (SPS), and the selected model for each encoded picture or slice is indicated by one or more syntax elements in the picture parameter set (PPS) or picture / slice header.
[0062] The network topology and parameters can be explicitly specified in the specification documents of video coding standards such as VVC, HEVC, or AV1. In this case, a predefined network model can be used for the entire encoded video sequence. If two or more network models are defined in the specification, one or more of these network models can be selected for each encoded video sequence, encoded picture, or slice.
[0063] In the case where a customized network is used for each encoded video sequence, picture, or slice, the network topology and parameters can be explicitly signaled in an advanced syntax structure, such as in a parameter set or SEI message in the base bitstream or in a parameter set or SEI message in a metadata track in a file format. In an embodiment, one or more network models with network topology and parameters are specified in one or more SEI messages, where these SEI messages can be inserted into encoded video sequences with different activation ranges. In addition, subsequent SEI messages in the encoded video bitstream can update previously activated SEI messages.
[0064] Figure 5 is a detailed block diagram of the optical flow estimation and intermediate frame synthesis module 410 during the test phase according to an embodiment.
[0065] As Figure 5 shown, the optical flow estimation and intermediate frame synthesis module 410 includes a flow estimation module 510, a backward warping module 520, and a fusion processing module 530.
[0066] Two reference frames {I0, I1} are used as inputs to the flow estimation module 510, which approximates the intermediate flow {F t , F 0->t , F 1->t} from the perspective of the frame I
[0067] to be synthesized. The flow estimation module 510 adopts a coarse-to-fine strategy with gradually increasing resolution: it iteratively updates the flow field. Conceptually, according to the iteratively updated flow field, the corresponding pixels are moved from the two input frames to the same positions in the potential intermediate frame. 0->t , F 1->t}, the backward warping module 520 generates a rough reconstructed frame or warped frame {I 0->t , I 1->t} by performing backward warping on the input frames {I0, I1}. Frame backward warping can be performed by inverse mapping and sampling the pixels of the potential intermediate frame to the input frames {I0, I1} to generate the warped frame {I 0->t , I 1->t}. The fusion processing module 530 combines the input frames {I0, I1}, the warped frame {I 0->t , I1->t} and the estimated intermediate flow {F 0->t , F 1->t} as inputs. The fusion processing module 530 estimates the fusion graph and another residual graph. The fusion graph and the residual graph are estimated graphs that map the feature changes in the input. Then, the deformed frames are linearly combined according to the fusion result and added to the residual graph for reconstruction to obtain the intermediate frame ( Figure 5 referred to as the reference frame I in t ).
[0068] Figure 6 is a detailed block diagram of the detail enhancement module 420 during the test phase according to an embodiment.
[0069] As Figure 6 shown, the detail enhancement module 420 includes a PCD alignment module 610, a TSA fusion module 620, and a reconstruction module 630.
[0070] Assume the reference frame I t is the reference to which all other frames will be aligned for use and the two reference frames {I t-1 , I t+1} are used as inputs together. The PCD (Pyramid, Cascade, and Deformable Convolution) alignment module 610 refines the (I t-1 , I t , I t+1 ) features of the reference frame into aligned features F t-对准 . Using the aligned features F t-对准 , the TSA (Temporal and Spatial Attention) fusion module 620 provides weights for the feature maps to focus attention on emphasizing important features for subsequent recovery and output prediction frame I p . A more detailed description of the PCD alignment module 610 and the TSA fusion module 620 will be described later with reference to Figure 7 and Figure 8 respectively.
[0071] Then, the reconstruction module formats the final residual of the reference frame I p based on the prediction frame I t . Finally, the addition module 640 adds the formatted final residual to the reference frame I t to obtain the enhanced frame I t-增强 as the final output of the detail enhancement module 420.
[0072] Figure 7 is a detailed block diagram of the PCD alignment module 610 within the detail enhancement module 420 during the test phase according to an embodiment. Figure 7 The modules with the same name and numbering convention in can be the same or one of multiple modules performing the functions (such as the deformable convolution module 730).
[0073] As shown Figure 7 , the PCD alignment module 610 includes a feature extraction module 710, an offset generation module 720, and a deformable convolution module 730. The PCD alignment module 610 calculates the alignment feature F between frames t-对准 .
[0074] First, using each reference frame in {I0, I1… I t , I t+1 …} as input, the feature extraction module 710 calculates feature maps {F0, F1… F t , F t+1 …} with three different levels by forward inference using the feature extraction DNN. For feature compensation, different levels have different resolutions to capture different levels of spatial / temporal information. For example, with larger motions in the sequence, smaller feature maps will be able to have larger corresponding fields and be able to handle various object offsets.
[0075] For feature maps of different levels, the offset generation module 720 calculates the offset map ΔP by concatenating the feature maps {F t , F t+1} and then passing the concatenated feature map through the offset generation DNN. Then, the deformable convolution module 730 calculates the new position of the DNN convolution kernel P 原始 (i.e., ΔP + P 原始 ) by adding the offset map ΔP to the original position P 新 . Note that the reference frame I t can be any frame within {I0, I1… I t , I t+1 …}. Without loss of generality, the frames can be arranged in an emphasized order based on their timestamps. In one embodiment, when the current target is to enhance the current reconstructed frame I' t , the reference frame I t is used as an intermediate frame.
[0076] Since the new position P 新 may be an irregular position and may not be an integer, the TDC operation can be performed by using interpolation (e.g., bilinear interpolation). By applying the deformable convolution kernel in the deformable convolution module 730, compensated features from different levels can be generated based on the level of the feature map and the corresponding generated offset. Then, the deformable convolution kernel can be applied based on one or more of the offset map ΔP and the compensated features by upsampling and adding to the upper-level compensated features to obtain the top-level alignment feature F t-对准 .
[0077] Figure 8Is a detailed block diagram of the TSA fusion module 620 within the detail enhancement module 420 during the test phase according to an embodiment.
[0078] As Figure 8 shown, the TSA fusion module 620 includes an activation module 810, an element-wise multiplication module 820, a fusion convolution module 830, a frame reconstruction module 840, and a frame synthesis module 850. The TSA fusion module 620 uses temporal and spatial attention. The goal of temporal attention is to calculate frame similarity in the embedding space. Intuitively, in the embedding space, more attention should be paid to adjacent frames that are more similar to the reference frame. Given the feature maps {F t-1 , F t+1} and the aligned feature F of the central frame t-对准 as inputs, the activation module 810 uses a sigmoid activation function as a simple convolutional filter to limit the inputs to [0,1] to obtain a temporal attention map {F' t-1 , F' t , F' t+1} as the output of the activation module 810. Note that for each spatial position, the temporal attention is spatially specific.
[0079] Then, the element-wise multiplication module 820 multiplies the temporal attention map {F' t-1 , F' t , F' t+1} with the aligned feature F t-对准 pixel by pixel. An additional fusion convolutional layer is employed in the fusion convolution module 830 to obtain the aggregated attention modulation feature F t-对准-TSA . Using the temporal attention map {F' t-1 , F' t , F' t+1} and the attention modulation feature F t-对准-TSA , the frame reconstruction module 840 calculates through feed-forward inference to generate the aligned frame I t-对准 using the frame reconstruction DNN. Then, the aligned frame I t-对准 is used by the frame synthesis module 850 to generate the synthesized predicted frame I p as the final output.
[0080] Figure 9 Is a detailed block diagram of the TSA fusion module 620 within the detail enhancement module 420 during the test phase according to another exemplary embodiment. Figure 9 The modules with the same name and numbering convention in can be the same or one of multiple modules performing the said functions.
[0081] As Figure 9As shown, the TSA fusion module 620 according to this embodiment includes an activation module 910, an element-wise multiplication module 920, a fusion convolution module 930, a downsampling convolution (DSC) module 940, an upsampling and addition module 950, a frame reconstruction module 960, and a frame synthesis module 970. Similar to the Figure 8 embodiment, the TSA fusion module 620 uses temporal and spatial attention. Given the feature maps {F t-1 , F t+1} and the aligned feature F t-对准 of the central frame as inputs, the activation module 910 uses a sigmoid activation function as a simple convolution filter to limit the inputs to [0, 1] to obtain the temporal attention maps {M t-1 , M t , M t+1} as the output of the activation module 910.
[0082] Then, the element-wise multiplication module 920 multiplies the temporal attention maps {M t-1 , M t , M t+1} with the aligned feature F t-对准 in a pixel-wise manner. A fusion convolution is performed on the product in the fusion convolution module 930 to generate the fused feature F 融合 . The fused feature F 融合 is downsampled and convolved by the DSC module 940. As Figure 9 shown, another convolutional layer can be adopted and processed by the DSC module 940. The output of each layer from the DSC module 940 is input to the upsampling and addition module 950. As Figure 9 shown, an additional layer of upsampling and addition is applied to the additional layer together with the fused feature F 融合 as the input. The upsampling and addition module 950 generates the fused attention map M t-融合 . The element-wise multiplication module 920 multiplies the fused attention map M t-融合 with the fused feature F 融合 in a pixel-wise manner to generate the TSA-aligned feature F t-TSA . Using the temporal attention maps {M t-1 , M t , M t+1} and the TSA-aligned feature F t-TSA , the frame reconstruction module 960 generates the TSA-aligned frame I t-TSA through feed-forward inference calculation using the frame reconstruction DNN. Then, the aligned frame I t-TSA generates the synthesized predicted frame I p as the final output through the frame synthesis module 970.
[0083] Figure 10It is a flowchart of a method 1000 for video encoding using neural network-based inter-frame prediction according to an embodiment.
[0084] In some implementations, Figure 10 one or more processing blocks of can be executed by platform 120. In some implementations, Figure 10 one or more processing blocks of can be executed by another device or a group of devices such as user device 110 that is separate from or includes platform 120.
[0085] As Figure 10 shown, in operation 1001, method 1000 includes generating an intermediate stream based on an input frame. The intermediate stream can be further iteratively updated and corresponding pixels move from two input frames to the same position in a potential intermediate frame.
[0086] In operation 1002, method 1000 includes generating a reconstructed frame by performing backward warping of the input frame using the intermediate stream.
[0087] In operation 1003, method 1000 includes generating a fusion map and a residual map based on the input frame, the intermediate stream, and the reconstructed frame.
[0088] In operation 1004, method 1000 includes generating a feature map with multiple levels using a first neural network based on a current reference frame, a first reference frame, and a second reference frame. The current reference frame can be generated by linearly combining the reconstructed frame according to the fusion map and adding the combined reconstructed frame to the residual map. In addition, the first reference frame can be a reference frame before the current reference frame in the output order, and the second reference frame can be a reference frame after the current reference frame in the output order.
[0089] In operation 1005, method 1000 includes generating a prediction frame based on aligned features from the generated feature map by refining the current reference frame, the first reference frame, and the second reference frame. Specifically, a prediction frame can be generated by first performing convolution to obtain an attention map, generating attention features based on the attention map and the aligned features, generating an aligned frame based on the attention map and the attention features, and then synthesizing the aligned frame to obtain the prediction frame.
[0090] The aligned features can be generated by calculating the offsets of the feature maps with multiple levels generated in operation 1004 and performing deformable convolution to generate compensation features for multiple levels. Then the aligned features can be generated based on at least one of the offsets and the generated compensation features.
[0091] In operation 1006, method 1000 includes generating a final residual based on the prediction frame. Weights of the feature maps generated in operation 1004 can also be generated to emphasize important features for generating subsequent final residuals.
[0092] In operation 1007, method 1000 includes calculating an enhanced frame as an output by adding a final residual to a current reference frame.
[0093] Although Figure 10 illustrates example blocks of the method, in some implementations, in addition to Figure 10 the blocks depicted in
[0094] Figure 11 is a block diagram of a device 1100 for video coding using neural network-based inter-frame prediction according to an embodiment.
[0095] As Figure 11 shown, the device includes a first generation code 1101, a second generation code 1102, a configured fusion code 1103, a third generation code 1104, a prediction code 1105, a residual code 1106, and a fourth generation code 1107.
[0096] The first generation code 1101 is configured to cause at least one processor to generate an intermediate stream based on an input frame.
[0097] The second generation code 1102 is configured to cause at least one processor to perform backward warping of the input frame using the intermediate stream to generate a reconstructed frame.
[0098] The configured fusion code 1103 is configured to cause at least one processor to generate a fusion map and a residual map based on the input frame, the intermediate stream, and the reconstructed frame.
[0099] The third generation code 1104 is configured to cause at least one processor to generate a feature map with multiple levels using a first neural network based on a current reference frame, a first reference frame, and a second reference frame.
[0100] The prediction code 1105 is configured to cause at least one processor to predict a frame by refining the current reference frame, the first reference frame, and the second reference frame based on aligned features from the generated feature map.
[0101] The residual code 1106 is configured to cause at least one processor to generate a final residual based on the predicted frame.
[0102] The fourth generation code 1107 is configured to cause at least one processor to generate an enhanced frame as an output by adding the final residual to the current reference frame.
[0103] Although Figure 11 illustrates example blocks of the device, in some implementations, in addition to Figure 11Outside the boxes depicted, the device may include additional boxes, fewer boxes, different boxes, or boxes in a different arrangement. Additionally or alternatively, two or more of the boxes of the device may be combined.
[0104] The device may further include: update code configured to cause at least one processor to iteratively update an intermediate stream and move corresponding pixels from two input frames to the same position in a potential intermediate frame; reference frame code configured to cause at least one processor to generate a current reference frame by linearly combining reconstructed frames according to a fusion map and adding the combined reconstructed frame to a residual map; determination code configured to cause at least one processor to determine weights of a feature map, wherein the weights emphasize important features for generating a subsequent final residual; calculation code configured to cause at least one processor to calculate offsets for multiple levels; compensation feature generation code configured to cause at least one processor to perform deformable convolution to generate compensation features for multiple levels; alignment feature generation code configured to cause at least one processor to generate alignment features based on at least one of the generated compensation features and an offset; execution code configured to cause at least one processor to perform convolution to obtain an attention map; attention feature generation code configured to cause at least one processor to generate attention features based on the attention map and the alignment features; alignment frame generation code configured to cause at least one processor to generate an alignment frame using a second neural network based on the attention map and the attention features; and synthesis code configured to cause at least one processor to synthesize the alignment frame to obtain a predicted frame.
[0105] Compared with traditional inter-frame generation methods, the proposed method performs a DNN-based network in video coding. The proposed method directly obtains data from a reference picture list (RPL) to generate a virtual reference frame, rather than calculating explicit motion vectors or motion flows that cannot handle complex motions or are error-prone. And then an enhanced deformable convolution (DCN) is applied to capture pixel offsets and implicitly compensate for large complex motions for further detail enhancement. Finally, a high-quality enhanced frame is reconstructed through a DNN model.
[0106] The proposed methods may be used alone or in any combination in any order. Additionally, each of the methods (or embodiments) may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.
[0107] This disclosure provides illustration and description, but is not intended to be exhaustive or to limit the implementations to the precise forms disclosed. Modifications and variations may be made in light of this disclosure, or may be obtained from practice of the implementations.
[0108] As used herein, the term component is intended to be broadly construed as hardware, firmware, or a combination of hardware and software.
[0109] It will be apparent that the systems and / or methods described herein can be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual special control hardware or software code used to implement these systems and / or methods does not limit the implementation. Accordingly, the operation and behavior of the systems and / or methods are described herein without reference to specific software code—it should be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0110] Even where combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features can be combined in ways not specifically recited in the claims and / or not disclosed in the specification. Although each dependent claim listed below can directly refer to only one claim, the disclosure of possible implementations includes each dependent claim in combination with every other claim in the claim group.
[0111] Any element, act, or instruction used herein is not to be construed as critical or essential unless explicitly so described. Additionally, as used herein, the articles “a” and “an” are intended to include one or more items and may be used interchangeably with “one or more.” Further, as used herein, the term “set” is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and may be used interchangeably with “one or more.” The term “one” or similar language is used where only one item is intended. Additionally, as used herein, the terms “having,” “comprising,” “containing,” etc. are intended to be open-ended terms. Further, unless otherwise explicitly stated, the phrase “based on” is intended to mean “at least partially based on.”
Claims
1. A method for video coding using inter-frame prediction based on a neural network, the method being executed by at least one processor, characterized in that, The method includes: generating an intermediate flow based on an input frame; generating a reconstructed frame by performing backward warping of the input frame using the intermediate flow; generating a fusion map and a residual map based on the input frame, the intermediate flow, and the reconstructed frame; generating feature maps with multiple levels using a first neural network based on a current reference frame, a first reference frame, and a second reference frame; generating a predicted frame by refining the current reference frame, the first reference frame, and the second reference frame based on aligned features from the generated feature maps; generating a final residual based on the predicted frame; and calculating an enhanced frame as an output by adding the final residual to the current reference frame.
2. The method according to claim 1, wherein The intermediate flow is iteratively updated and corresponding pixels are moved from two input frames to the same position in a potential intermediate frame.
3. The method according to claim 1, wherein The method further includes generating the current reference frame by linearly combining the reconstructed frame according to the fusion map and adding the combined reconstructed frame to the residual map.
4. The method according to claim 1, wherein The first reference frame is a reference frame before the current reference frame in the output order, and the second reference frame is a reference frame after the current reference frame in the output order.
5. The method according to claim 1, wherein The method further includes determining weights of features in the feature maps, where the weights emphasize a subset of features for generating a subsequent final residual.
6. The method according to claim 1, wherein The method further includes: calculating offsets for the multiple levels; performing deformable convolution to generate compensated features for the multiple levels; and generating the aligned features based on at least one of the generated compensated features and the offsets.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: performing convolution to obtain a fusion attention map; generating attention features based on the attention map and the aligned features; generating an aligned frame using a second neural network based on the attention map and the attention features; and synthesizing the aligned frame to obtain the predicted frame.
8. A computer device, characterized in that, The computer device includes: one or more computer-readable non-transitory storage media configured to store computer program code; and one or more computer processors configured to access the computer program code and execute the method according to any one of claims 1 to 7 as indicated by the computer program code.
9. A device for video coding using inter-frame prediction based on a neural network, characterized in that, The device includes: a first generation unit configured to generate an intermediate flow based on an input frame; a second generation unit configured to perform backward warping of the input frame using the intermediate flow to generate a reconstructed frame; a fusion unit configured to generate a fusion map and a residual map based on the input frame, the intermediate flow, and the reconstructed frame; a third generation unit configured to generate feature maps with multiple levels using a first neural network based on a current reference frame, a first reference frame, and a second reference frame; a prediction unit configured to predict a frame based on aligned features from the generated feature maps by refining the current reference frame, the first reference frame, and the second reference frame; a residual unit configured to generate a final residual based on the predicted frame; and a fourth generation unit configured to generate an enhanced frame as an output by adding the final residual to the current reference frame.
10. The device according to claim 9, characterized in that, The device further includes an update unit configured to iteratively update the intermediate stream and move corresponding pixels from two input frames to the same positions in a potential intermediate frame.
11. The device according to claim 9, characterized in that, The device further includes a reference frame unit configured to generate the current reference frame by linearly combining the reconstructed frames according to the fusion map and adding the combined reconstructed frames to the residual map.
12. The device according to claim 9, characterized in that, The first reference frame is a reference frame before the current reference frame in the output order, and the second reference frame is a reference frame after the current reference frame in the output order.
13. The device according to claim 9, characterized in that, The device further includes a determination unit configured to determine weights of features in the feature map, where the weights emphasize a subset of features for generating a subsequent final residual.
14. The device according to claim 9, characterized in that, The device further includes: a calculation unit configured to calculate offsets for the plurality of levels; a compensated feature generation unit configured to perform deformable convolution to generate compensated features for the plurality of levels; and an aligned feature generation unit configured to generate the aligned features based on at least one of the generated compensated features and the offsets.
15. The device according to any one of claims 9 to 14, characterized in that, The device further includes: an execution unit configured to perform convolution to obtain a fusion attention map; an attention feature generation unit configured to generate attention features based on the attention map and the aligned features; an aligned frame generation unit configured to generate an aligned frame using a second neural network based on the attention map and the attention features; and a synthesis unit configured to synthesize the aligned frame to obtain the predicted frame.
16. A non-transitory computer-readable medium storing instructions, characterized in that, When executed by at least one processor for video coding using neural network-based inter-frame prediction, the instructions cause the at least one processor to perform the method according to any one of claims 1 to 7.
17. A method for storing a bitstream, characterized in that, Performing the method according to any one of claims 1 - 7 generates a bitstream; and storing the bitstream.
18. A method for transmitting a bitstream, characterized in that Performing the method according to any one of claims 1 - 7 generates a bitstream; and transmitting the bitstream.
19. A computer-readable storage medium, on which a computer program / instructions and a bit stream are stored, characterized in that, When executed by a processor, the computer program / instructions implement the steps of the method according to any one of claims 1 - 7 to generate the bitstream.