Method, apparatus, and electronic device for permutation neural residual compression

By adopting a DNN-based permutation neural residual compression method in video encoding and decoding, the problem of low video residual compression efficiency in the prior art is solved, and better compression performance and flexible bit rate control are achieved.

CN114631101BActive Publication Date: 2025-06-13TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180006006.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-04-21
Filing Date
2021-05-27
Publication Date
2025-06-13
Estimated Expiration
2041-05-27

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies are not efficient when compressing video residuals, making it difficult to achieve optimized bit rate and compression quality.

Method used

The permutation neural residual compression method based on deep neural network (DNN) is adopted to estimate the motion vector, obtain predicted image frames, calculate the permutation residual, and use the first neural network to encode and compress the permutation residual.

Benefits of technology

Improves the compression efficiency of video encoding and decoding, provides better compression performance and flexible bit rate control, and enables adjustment of bit rate and target metrics without retraining the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114631101B_ABST
    Figure CN114631101B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method, apparatus, and electronic device for displacement neural residual compression. The method includes: estimating a motion vector based on a current image frame and a previously reconstructed image frame; obtaining a predicted image frame based on the estimated motion vector and the previously reconstructed image frame; and subtracting the obtained predicted image frame from the current image frame to obtain a displacement residual. The method further includes encoding the obtained displacement residual using a first neural network to obtain an encoded representation of the displacement residual, and compressing the encoded representation of the displacement residual to obtain a compressed representation of the encoded representation of the displacement residual.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 052,242, filed on July 15, 2020, and U.S. Patent Application No. 17 / 236,108, filed on April 21, 2021, the disclosures of which are hereby incorporated by reference in their entireties into this application. Technical field

[0003] The present disclosure relates to video coding and decoding, and more particularly to a method, apparatus, and electronic device for permutation neural residual compression. Background art

[0004] Video coding and decoding standards such as H.264 / Advanced Video Coding (H.264 / AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC) share a similar (recursive) block - based hybrid prediction / transformation framework. In this framework, individual coding and decoding tools such as intra - frame / inter - frame prediction, integer transformation, and context - adaptive entropy coding are centrally hand - designed to optimize overall efficiency. The prediction signal is constructed using spatio - temporal pixel neighborhoods to obtain the corresponding residual for subsequent transformation, quantization, and entropy coding.

[0005] On the other hand, the essence of a deep neural network (DNN) is to extract different levels of spatio - temporal excitation by analyzing spatio - temporal information from the receptive fields of adjacent pixels. The ability to explore highly non - linear and non - local spatio - temporal correlations provides a promising opportunity for significantly improving compression quality. Summary of the invention

[0006] According to an embodiment, a method for permutation neural residual compression, executed by at least one processor, and includes: estimating a motion vector based on a current image frame and a previously reconstructed image frame; obtaining a predicted image frame based on the estimated motion vector and the previously reconstructed image frame; and subtracting the obtained predicted image frame from the current image frame to obtain a permutation residual. The method further includes encoding the obtained permutation residual using a first neural network to obtain an encoded representation of the permutation residual, and compressing the encoded representation of the permutation residual to obtain a compressed representation of the encoded representation of the permutation residual.

[0007] According to an embodiment, an apparatus for displacement neural residual compression includes: at least one memory configured to store program code; and at least one processor configured to read the program code and operate according to the instructions of the program code. The program code includes: an estimation code configured to cause the at least one processor to estimate a motion vector based on a current image frame and a previously reconstructed image frame; a first obtaining code configured to cause the at least one processor to obtain a predicted image frame based on the estimated motion vector and the previously reconstructed image frame; and a subtraction code configured to cause the at least one processor to subtract the obtained predicted image frame from the current image frame to obtain a displacement residual. The program code further includes: an encoding code configured to cause the at least one processor to encode the obtained displacement residual using a first neural network to obtain an encoded representation of the displacement residual; and a compression code configured to cause the at least one processor to compress the encoded representation of the displacement residual to obtain a compressed representation of the encoded representation of the displacement residual.

[0008] According to an embodiment, a non-transitory computer-readable medium stores instructions that, when executed by at least one processor to perform displacement neural residual compression, cause the at least one processor to estimate a motion vector based on a current image frame and a previously reconstructed image frame, obtain a predicted image frame based on the estimated motion vector and the previously reconstructed image frame, subtract the obtained predicted image frame from the current image frame to obtain a displacement residual, encode the obtained displacement residual using a first neural network to obtain an encoded representation of the displacement residual, and compress the encoded representation of the displacement residual to obtain a compressed representation of the encoded representation of the displacement residual. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 is an environmental diagram according to an embodiment in which the methods, apparatuses, and systems described in the present disclosure can be implemented.

[0010] Figure 2 is Figure 1 a block diagram of example components of one or more devices.

[0011] Figure 3 is a block diagram of a training apparatus for displacement neural residual compression in a training phase according to an embodiment.

[0012] Figure 4 is a block diagram of a testing apparatus for displacement neural residual compression in a testing phase according to an embodiment.

[0013] Figure 5 is a flowchart of a displacement neural residual compression method according to an embodiment.

[0014] Figure 6 is a block diagram of a displacement neural residual compression apparatus according to an embodiment.

[0015] Figure 7 is a flowchart of a permutation neural residual compression method according to an embodiment.

[0016] Figure 8 is a block diagram of an apparatus for permutation neural residual compression according to an embodiment. Detailed implementation manners

[0017] The present disclosure relates to a method and an apparatus for compressing an input video by learning to use permutation residuals in a video codec framework based on DNN-based residual compression. The learned permutation residuals are better alternatives to the original residuals by being visually similar to the original residuals but better compressed. Moreover, the method and apparatus of the present disclosure have flexibility in controlling the bit rate in DNN-based residual compression.

[0018] The video compression framework can be described as follows. The input video x includes image frames x 1 ,..., x T . In the first motion estimation step, the image frames are divided into spatial blocks (e.g., 8×8 squares), and a motion vector m t for each block is calculated between the current image frame x ( which may include a set of previously reconstructed image frames) and the previously reconstructed image frame t set. Then, in the second motion compensation step, a predicted image frame t is obtained by replicating the corresponding pixels of the previously reconstructed image frame based on the motion vector m and the residual r t between the original image frame x and the predicted image frame t can be obtained: In the third step, after a linear transformation such as a discrete cosine transform (DCT) is performed, the residual r t is quantized. The DCT coefficients of r t are quantized to obtain better quantization performance. The quantization step provides a quantized value The motion vector m t and the quantized value are both encoded into a bitstream by entropy coding and sent to the decoder.

[0019] On the decoder side, first, the quantized value is dequantized by an inverse transformation such as an inverse discrete cosine transform (IDCT) with inverse quantization coefficients to obtain a recovered residual Then, the recovered residual is added back to the predicted image frame Obtain the reconstructed image frame

[0020] For the residual r t The efficiency of compressing the residual is a factor in video compression performance. A DNN-based method can be used to assist in residual compression. For example, a DNN can be used to learn highly non-linear transformations instead of linear transformations to improve quantization efficiency. The residual can also be encoded through an end-to-end (E2E) DNN, where the quantized representation is directly learned without explicit transformation.

[0021] Figure 1 is according to an embodiment and is a schematic diagram of an environment 100 in which the methods, apparatuses, and systems described herein can be implemented.

[0022] As Figure 1 shown, the environment 100 can include a user device 110, a platform 120, and a network 130. The devices in the environment 100 can be interconnected by wired connections, wireless connections, or a combination of wired and wireless connections.

[0023] The user device 110 includes one or more devices that are capable of receiving, generating, storing, processing, and / or providing information related to the platform 120. For example, the user device 110 can include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smartphone, a wireless phone, etc.), a wearable device (e.g., smart glasses or a smart watch), or a similar device. In some embodiments, the user device 110 can receive information from and / or send information to the platform 120.

[0024] The platform 120 includes one or more devices as described elsewhere herein. In some embodiments, the platform 120 can include a cloud server or a group of cloud servers. In some embodiments, the platform 120 can be designed to be modular such that software components can be swapped in or out. In this way, the platform 120 can be easily and / or quickly reconfigured to have different uses.

[0025] In some embodiments, as shown, the platform 120 can be hosted in a cloud computing environment 122. It is noted that while the embodiments described herein describe the platform 120 as being hosted in a cloud computing environment 122, in some embodiments, the platform 120 is not cloud-based (i.e., can be implemented outside of a cloud computing environment) or can be partially cloud-based.

[0026] The cloud computing environment 122 includes an environment that hosts the platform 120. The cloud computing environment 122 can provide services such as computing, software, data access, storage, etc., which do not require the end user (e.g., the user device 110) to know the physical location and configuration of the systems and / or devices of the hosted platform 120. As shown in the figure, the cloud computing environment 122 can include a set of computing resources 124 (collectively referred to as "computing resources 124" and individually referred to as "computing resource 124").

[0027] The computing resources 124 include one or more personal computers, workstation computers, server devices, or other types of computing and / or communication devices. In some embodiments, the computing resources 124 can host the platform 120. Cloud resources can include computing instances executed in the computing resources 124, storage devices provided in the computing resources 124, data transmission devices provided by the computing resources 124, etc. In some embodiments, the computing resources 124 can communicate with other computing resources 124 through a wired connection, a wireless connection, or a combination of wired and wireless connections.

[0028] Further as Figure 1 shown, the computing resources 124 include a set of cloud resources, such as one or more applications ("APP") 124-1, one or more virtual machines ("VM") 124-2, virtualized storage ("VS") 124-3, one or more hypervisors ("HYP") 124-4, etc.

[0029] The application 124-1 includes one or more software applications, which can be provided to and / or accessed by the user device 110 and / or the platform 120. The application 124-1 does not require the installation and execution of software applications on the user device 110. For example, the application 124-1 can include software related to the platform 120, and / or any other software that can be provided through the cloud computing environment 122. In some embodiments, an application 124-1 can send / receive information to / from one or more other applications 124-1 through the virtual machine 124-2.

[0030] The virtual machine 124-2 includes a software implementation of a machine (e.g., a computer) that executes programs, similar to a physical machine. The virtual machine 124-2 can be a system virtual machine or a process virtual machine, depending on the usage and correspondence of the virtual machine 124-2 to any real machine. A system virtual machine can provide a complete system platform that supports the execution of a complete operating system ("OS"). A process virtual machine can execute a single program and can support a single process. In some embodiments, the virtual machine 124-2 can execute on behalf of a user (e.g., the user device 110) and can manage the infrastructure of the cloud computing environment 122, such as data management, synchronization, or long-term data transfer.

[0031] The virtualized storage 124-3 includes one or more storage systems and / or one or more devices that use virtualization technology within the storage system or device of the computing resources 124. In some embodiments, within the context of a storage system, the types of virtualization can include block virtualization and file virtualization. Block virtualization can refer to the abstraction (or separation) of logical storage from physical storage so that the storage system can be accessed without considering the physical storage or heterogeneous structure. The separation can allow the administrator of the storage system to flexibly manage the storage of end users. File virtualization can eliminate the dependence between the data accessed at the file level and the location of the physical storage file. This can optimize the performance of storage usage, server consolidation, and / or uninterrupted file migration.

[0032] The hypervisor 124-4 can provide hardware virtualization technology that allows multiple operating systems (e.g., "guest operating systems") to execute simultaneously on a host computer such as the computing resources 124. The hypervisor 124-4 can provide a virtual operating platform to the guest operating systems and can manage the execution of the guest operating systems. Multiple instances of various operating systems can share the virtualized hardware resources.

[0033] Network 130 includes one or more wired and / or wireless networks. For example, network 130 may include a cellular network (e.g., a fifth generation (5G) network, a Long-Term Evolution (LTE) network, a third generation (3G) network, a Code Division Multiple Access (CDMA) network, etc.), a Public Land Mobile Network (PLMN), a Local Area Network (LAN), a Wide Area Network (WAN), a Metropolitan Area Network (MAN), a telephone network (e.g., a Public Switched Telephone Network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-optic based network, etc., and / or a combination of these or other types of networks.

[0034] Figure 1 The number and arrangement of the devices and networks shown are provided as examples. In fact, compared with Figure 1 the devices and / or networks shown, there may be more devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks with a different arrangement. Additionally, Figure 1 two or more of the devices shown may be implemented within a single device, or Figure 1 a single device shown may be implemented as multiple distributed devices. Additionally or alternatively, a set of devices (e.g., one or more devices) of environment 100 may perform one or more functions described as being performed by another set of devices of environment 100.

[0035] Figure 2 is Figure 1 a block diagram of example components of one or more of the devices in

[0036] Device 200 may correspond to user device 110 and / or platform 120. As Figure 2 shown, device 200 may include a bus 210, a processor 220, a memory 230, a storage component 240, an input component 250, an output component 260, and a communication interface 270.

[0037] The bus 210 includes components that allow communication between the components of the device 200. The processor 220 is implemented in hardware, firmware, or a combination of hardware and software. The processor 220 is a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or another type of processing component. In some embodiments, the processor 220 includes one or more processors that can be programmed to perform functions. The memory 230 includes random access memory (RAM), read only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, and / or optical memory) that stores information and / or instructions for use by the processor 220.

[0038] The storage component 240 stores information and / or software related to the operation and use of the device 200. For example, the storage component 240 may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid state disk), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cassette tape, a magnetic tape, and / or another type of non-volatile computer-readable medium, as well as corresponding drives.

[0039] The input component 250 includes components that allow the device 200 to receive information, for example, through user input, such as a touch screen display, a keyboard, a keypad, a mouse, buttons, switches, and / or a microphone. Additionally or alternatively, the input component 250 may include sensors for sensing information (e.g., a global positioning system (GPS) component, an accelerometer, a gyroscope, and / or an actuator). The output component 260 includes components that provide output information from the device 200, such as a display, a speaker, and / or one or more light emitting diodes (LEDs).

[0040] The communication interface 270 includes transceiver-like components (e.g., a transceiver and / or separate receiver and transmitter) that enable the device 200 to communicate with other devices, for example, through a wired connection, a wireless connection, or a combination of wired and wireless connections. The communication interface 270 may allow the device 200 to receive information from and / or provide information to another device. For example, the communication interface 270 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, etc.

[0041] Device 200 may perform one or more processes described herein. Device 200 may perform these processes in response to a processor 220 executing software instructions stored by a non-transitory computer-readable medium (e.g., memory 230 and / or storage component 240). A computer-readable medium is defined herein as a non-transitory memory device. A memory device includes storage space within a single physical storage device or storage space distributed across multiple physical storage devices.

[0042] The software instructions may be read into memory 230 and / or storage component 240 from another computer-readable medium or from another device via a communication interface 270. When executed, the software instructions stored in memory 230 and / or storage component 240 may cause processor 220 to perform one or more processes described herein. Additionally or alternatively, hardware wired circuitry may be used in place of or in combination with software instructions to perform one or more processes described herein. Accordingly, the embodiments described herein are not limited to any particular combination of hardware circuitry and software.

[0043] Figure 2 The number and arrangement of the components shown are provided as an example. In fact, compared with the Figure 2 components shown, device 200 may include more components, fewer components, different components, or components arranged differently. Additionally or alternatively, a set of components of device 200 (e.g., one or more components) may perform one or more functions described as being performed by another set of components of device 200.

[0044] A method and apparatus for permutation neural residual compression will now be described in detail.

[0045] The method and apparatus may improve the compression efficiency of DNN-based residuals. For each image frame having a residual compressed by an E2E DNN, a permutation residual is generated, and the permutation residual may provide better compression performance than the original residual.

[0046] Given a residual r of size (h, w, c) t , where h, w, and c are the height, width, and number of channels respectively, E2E residual compression includes computing a compressed representation from the compressed representation the recovered residual can be reconstructed A distortion loss function D(r t , ) is used to measure the reconstruction error (i.e., the distortion loss), such as mean squared error (MSE) or peak signal-to-noise ratio (PSNR). A rate loss function R( ) is used to measure the compressed representation Bit consumption. The compromise hyperparameter λ is used to balance the joint rate-distortion (R-D) loss L:

[0047]

[0048] Training with a larger compromise hyperparameter λ results in a compressed model with less distortion but more bit consumption, and vice versa. To achieve different bitrates in practice, the E2E residual compression method can train multiple model instances, one for each target compromise hyperparameter λ, and store all these model instances on the encoder and decoder sides. However, the embodiments described in this disclosure provide the ability to have a flexible bitrate control range without training and storing multiple model instances.

[0049] Figure 3 and Figure 4 The overall workflow of the encoder and decoder according to an embodiment is provided. There are two different stages: the training stage and the testing stage.

[0050] Figure 3 FIG. is a block diagram of a training apparatus 300 for permutation neural residual compression in the training stage according to an embodiment.

[0051] As Figure 3 shown, the training apparatus 300 includes a motion estimation component 305, a motion compensation component 310, a subtractor 315, a training DNN encoder 320, a training encoder 325, a training decoder 330, a training DNN decoder 335, an adder 340, a rate loss component 345, a distortion loss component 350, and a data update component 355.

[0052] The goal of the training stage is to generate a permutation residual r t ′ for each image frame x t Given the current image frame x t and the previously reconstructed image frame the motion estimation component 305 obtains the motion vector m t The previously reconstructed image frame can include a set of image frames. For example, when the current image frame x t is a P-frame, all previously reconstructed image frames are before the current image frame x t When the current image frame x t is a B-frame, the previously reconstructed image frames contain the image frames before and after the current image frame x t When the current frame x t is a low-latency B-frame, the previously reconstructed image frame is the one before the current image frame x tBefore. The motion estimation component 305 can use one or more block-based motion estimators or DNN-based optical flow estimators.

[0053] The motion compensation component 310 obtains a predicted image frame based on the previously reconstructed image frame and the motion vector m t to obtain a predicted image frame In an embodiment, a DNN can be used to simultaneously calculate the motion vector m t and the predicted image frame

[0054] The subtractor 315 obtains the residual r between the original image frame x t and the predicted image frame to obtain the residual r t :

[0055] Using the residual r t as input, the trained DNN encoder 320 obtains a DNN-encoded representation, and based on this DNN-encoded representation, the trained encoder 325 obtains a compressed representation The trained DNN encoder 320 and the trained encoder 325 are the encoder part of the E2E DNN encoder-decoder.

[0056] Based on the compressed representation the trained decoder 330 obtains a decompressed representation, which is used as input for the trained DNN decoder 335 to obtain the recovered residual The trained decoder 330 and the trained DNN decoder 335 are the decoder part of the E2E DNN encoder-decoder.

[0057] In the present disclosure, there are no restrictions on the network structure of the E2E DNN encoder-decoder as long as it can learn using gradient backpropagation. In an embodiment, the trained encoder 325 and the trained decoder 330 use a differentiable statistical sampler to approximate the true quantization and dequantization effects.

[0058] The adder 340 adds the recovered residual back to the predicted image frame to obtain the reconstructed image frame

[0059] In the training phase, first, the trained DNN encoder 320 and the trained DNN decoder 335 are initialized, that is, based on a predetermined DNN encoder and a predetermined DNN decoder, the model weights of the trained DNN encoder 320 and the trained DNN decoder 335 are set.

[0060] Then, a retraining / fine-tuning process is performed to calculate the permutation residual rt ′ such that the total loss L(r t ′, )(where the permutation residual r t ′ is used to replace the residual r t ) can be optimized or reduced.

[0061] Specifically, the rate loss component 345 uses a predetermined rate loss estimator to obtain the rate loss R( ). For example, an entropy estimation method can be used to estimate the rate loss R( ).

[0062] The distortion loss component 350 obtains the distortion loss D(r , t , ) based on the reconstructed image frames.

[0063] The data update component 355 obtains the gradient of the total loss L(r t ′, ) to update the permutation residual r t ′ through backpropagation. Iterate this backpropagation process to continue updating the permutation residual r t ′ until a stopping criterion is reached, such as when the optimization converges or when the maximum number of iterations is reached.

[0064] In an embodiment, based on a pre-training dataset, the weights of a predetermined DNN encoder, a predetermined DNN decoder, and a predetermined rate loss estimator are pre-trained, where for each training image frame the same process is performed to calculate the residual Then the same forward calculation is performed on the residual through DNN encoding, encoding, decoding, and DNN decoding to generate a compressed representation and a recovered residual Then, given a hyperparameter λ pre-train , a pre-training loss is obtained, and the pre-training loss includes both the rate loss and the distortion loss L pre-train ( ), similar to Equation (1), and the gradients of the rate loss and the distortion loss are used to update the weights of the predetermined DNN encoder, the predetermined DNN decoder, and the predetermined rate loss estimator through iterative backpropagation. The pre-training dataset can be the same as or different from the dataset based on which the training DNN encoder 320 and the training DNN decoder 335 are trained.

[0065] Figure 4 is a block diagram of a test device 400 for permutation neural residual compression in the test phase according to an embodiment.

[0066] AsFigure 4 As shown in Figure 4 , the test apparatus 400 includes a motion estimation component 305, a motion compensation component 310, a subtractor 315, a test DNN encoder 405, a test encoder 410, a test decoder 415, a test DNN decoder 420, and an adder 340.

[0067] During the test phase of the encoders (i.e., the test DNN encoder 405 and the test encoder 410), after learning the permutation residual r t ′, the permutation residual r t ′ undergoes the forward inference process of the test DNN encoder 405 to generate a DNN-encoded representation. Based on this DNN-encoded representation, the test encoder 410 obtains the final compressed representation.

[0068] Then, the test decoder 415 obtains a decompressed representation, which is used as the input to the test DNN decoder 420 to obtain the reconstructed residual.

[0069] In an embodiment, the test DNN encoder 405 and the test DNN decoder 420 are respectively the same as the training DNN encoder 320 and the training DNN decoder 335, while the test encoder 410 and the test decoder 415 are respectively different from the training encoder 325 and the training decoder 330. As described above, the quantization and dequantization processes are replaced by differentiable statistical samplers in the training encoder 325 and the training decoder 330. During the test phase, true quantization and dequantization are respectively performed in the test encoder 410 and the test decoder 415. In the present disclosure, there is no limitation on the quantization method and the dequantization method used for the test encoder 410 and the test decoder 415.

[0070] After reconstructing the reconstructed residual , the adder 340 adds the reconstructed residual back to the predicted image frame to obtain the reconstructed image frame and the test apparatus 400 continues to process the next image frame x t+1 .

[0071] The motion vector m t and the compressed representation can both be sent to the decoder. They can be further encoded into a bitstream through entropy coding.

[0072] During the test phase of the decoder (i.e., the test decoder 415 and the test DNN decoder 420), after obtaining the motion vector m t and the compressed representation After (e.g., by decoding from the encoded bitstream), given a previously reconstructed image frame The test decoder 415 obtains the decompressed representation, and the test DNN decoder 420 obtains the reconstructed residual based on this decompressed representation In the motion compensation component on the decoder side, in the same manner as the motion compensation component 310 on the encoder side, based on the previously reconstructed image frame and the motion vector m t , the predicted image frame is obtained Then, the adder 340 adds the reconstructed residual back to the predicted image frame to obtain the reconstructed image frame And the test device 400 continues to process the next image frame

[0073] The above embodiments provide flexibility in bitrate control and target metric control. When the target bitrate of the compressed representation changes, without retraining / fine-tuning the trained DNN encoder 320, trained encoder 325, trained decoder 330, trained DNN decoder 335, test DNN encoder 405, test encoder 410, test decoder 415, and test DNN decoder 420, only change Figure 3 the hyperparameter λ in the training phase of the encoding process described in . Similarly, to obtain the compressed residuals that are optimal for different target metrics (e.g., PSNR or structural similarity (SSIM)) of the compressed representation Figure 3 , the way of obtaining the distortion loss in the training phase of the encoding process described in

[0074] Figure 5 is a flowchart of a method for permutation neural residual compression according to an embodiment.

[0075] In some embodiments, Figure 5 one or more processing blocks of Figure 5 can be executed by the platform 120. In some implementations,

[0076] As Figure 5As shown in FIG. 5, in operation 510, method 500 includes estimating a motion vector based on a current image frame and a previously reconstructed image frame.

[0077] In operation 520, method 500 includes obtaining a predicted image frame based on the estimated motion vector and the previously reconstructed image frame.

[0078] In operation 530, method 500 includes subtracting the obtained predicted image frame from the current image frame to obtain a displacement residual.

[0079] In operation 540, method 500 includes encoding the obtained displacement residual using a first neural network to obtain an encoded representation.

[0080] In operation 550, method 500 includes compressing the encoded representation.

[0081] The estimating the motion vector, obtaining the predicted image frame, obtaining the displacement residual, obtaining the encoded representation, and compressing the encoded representation may be performed by an encoding processor.

[0082] Although Figure 5 example blocks of method 500 are shown, in some embodiments, method 500 may include Figure 5 blocks other than those depicted in FIG. 5, fewer blocks than those depicted in FIG. 5, different blocks than those depicted in FIG. 5, or blocks arranged differently than those depicted in FIG. 5. Additionally or alternatively, two or more of the blocks of method 500 may be performed in parallel.

[0083] Figure 6 FIG. 5 is a flowchart of a method of displacement neural residual compression according to an embodiment.

[0084] In some implementations, Figure 6 one or more of the processing blocks of FIG. 5 may be performed by platform 120. In some embodiments, Figure 6 one or more of the processing blocks of FIG. 5 may be performed by another device or group of devices (such as user device 110) separate from or including platform 120.

[0085] As Figure 6 shown in FIG. 6, in operation 610, method 600 includes decompressing the compressed representation.

[0086] In operation 620, method 600 includes decoding the decompressed representation using a second neural network to obtain a recovered residual.

[0087] In operation 630, method 600 includes adding the obtained recovered residual to the obtained predicted image frame to obtain a reconstructed image frame.

[0088] The first neural network and the second neural network can be trained through the following operations: determining a distortion loss based on the obtained recovered residual and the obtained permutation residual; determining a rate loss based on the compressed representation; determining a gradient of the rate-distortion loss based on the determined distortion loss, the determined rate loss, and hyperparameters; and updating the obtained permutation residual to reduce the determined gradient of the rate-distortion loss.

[0089] The hyperparameters can be set based on the target bitrate of the compressed representation without retraining the first neural network and the second neural network.

[0090] The distortion loss can be determined based on a function that is set based on the type of the target metric of the compressed representation without retraining the first neural network and the second neural network.

[0091] The decompressing of the compressed representation, decoding the decompressed representation, and obtaining the reconstructed image frame can be performed by a decoding processor. The method can further include obtaining a predicted image frame by the decoding processor based on the estimated motion vector and the previously reconstructed image frame.

[0092] Although Figure 6 illustrates example blocks of method 600, in some embodiments, method 600 can include Figure 6 blocks other than those depicted in, fewer blocks than, different blocks from, or a different arrangement of those depicted in. Additionally or alternatively, two or more of the blocks of method 600 can be executed in parallel.

[0093] Figure 7 is a block diagram of an apparatus 700 for permutation neural residual compression according to an embodiment.

[0094] As Figure 7 shown, apparatus 700 includes an estimation code 710, a first acquisition code 720, a subtraction code 730, an encoding code 740, and a compression code 750.

[0095] The estimation code 710 is configured to cause at least one processor to estimate a motion vector based on a current image frame and a previously reconstructed image frame.

[0096] The first acquisition code 720 is configured to cause at least one processor to obtain a predicted image frame based on the estimated motion vector and the previously reconstructed image frame;

[0097] The subtraction code 730 is configured to cause at least one processor to subtract the obtained predicted image frame from the current image frame to obtain a permutation residual;

[0098] The encoding code 740 is configured to cause at least one processor to encode the obtained permutation residuals using a first neural network to obtain an encoded representation; and

[0099] The compression code 750 is configured to cause at least one processor to compress the encoded representation.

[0100] The estimation code, the first acquisition code, the subtraction code, the encoding code, and the compression code may be configured to cause an encoding processor to perform functions.

[0101] Figure 8 is a block diagram of an apparatus 800 for permutation neural residual compression according to an embodiment.

[0102] As Figure 8 shown, the apparatus 800 includes a decompression code 810, a decoding code 820, and an addition code 830.

[0103] The decompression code 810 is configured to cause at least one processor to decompress the compressed representation;

[0104] The decoding code 820 is configured to cause at least one processor to decode the decompressed representation using a second neural network to obtain a recovered residual.

[0105] The addition code 830 is configured to cause at least one processor to add the obtained recovered residual to the obtained predicted image frame to obtain a reconstructed image frame.

[0106] The first neural network and the second neural network may be trained by: determining a distortion loss based on the obtained recovered residual and the obtained permutation residual; determining a rate loss based on the compressed representation; determining a gradient of the rate-distortion loss based on the determined distortion loss, the determined rate loss, and hyperparameters; and updating the obtained permutation residual to reduce the gradient of the determined rate-distortion loss.

[0107] The hyperparameters may be set based on the target bit rate of the compressed representation without retraining the first neural network and the second neural network.

[0108] The distortion loss may be determined based on a function that is set based on the type of the target metric of the compressed representation without retraining the first neural network and the second neural network.

[0109] The decompression code 810, the decoding code 820, and the addition code 830 may be configured to cause a decoding processor to perform functions. The apparatus 800 may further include a second acquisition code, which is configured to cause the decoding processor to obtain a predicted image frame based on an estimated motion vector and a previously reconstructed image frame.

[0110] Compared with previous video compression methods, the above embodiments have the following advantages. These embodiments can be regarded as general modules applicable to any E2E residual compression DNN method. For each individual image frame, its permutation residual can be optimized through a separate retraining / fine-tuning process based on the feedback of its loss, which can improve the compression performance.

[0111] In addition, the above embodiments can achieve flexible bitrate control without retraining / fine-tuning the E2E residual compression model or using multiple models. The embodiments can also change the target compression metric without retraining / fine-tuning the E2E residual compression model.

[0112] These methods can be used alone or in any combination in any order. In addition, each of the methods (or embodiments), encoders, and decoders can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-volatile computer-readable medium.

[0113] The above disclosure provides illustration and description, but is not intended to be exhaustive or limit the implementation to the exact form disclosed. Modifications and variations are possible in light of the above disclosure, or may be obtained from the practice of the implementation.

[0114] As used herein, the term component is intended to be broadly interpreted as hardware, firmware, or a combination of hardware and software.

[0115] Obviously, the systems and / or methods described herein can be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual specific control hardware or software code for implementing these systems and / or methods is not a limitation on the implementation. Therefore, the operation and behavior of the systems and / or methods are described herein without reference to specific software code—it should be understood that the software and hardware can be designed to implement the systems and / or methods based on the description herein.

[0116] Even if combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features can be combined in ways not specifically recited in the claims and / or disclosed in the specification. Although each of the dependent claims listed below may directly depend on only one claim, the disclosure of possible implementations includes each dependent claim in combination with all other claims in the claim set.

[0117] The elements, acts, or instructions used herein cannot be construed as critical or essential unless expressly described as such. Also, as used herein, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." Additionally, as used herein, the term "set" is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and may be used interchangeably with "one or more." The term "one" or similar language is used where only one item is intended. Also, as used herein, the terms "has," "have," "having," etc. are intended to be open-ended terms. Additionally, unless expressly stated otherwise, the phrase "based on" is intended to mean "at least partially based on."

Claims

1. A method for permutation neural residual compression, characterized in that, the method includes: estimating a motion vector based on a current image frame and a previously reconstructed image frame; obtaining a predicted image frame based on the estimated motion vector and the previously reconstructed image frame; subtracting the obtained predicted image frame from the current image frame to obtain a permutation residual; encoding the obtained permutation residual using a first neural network to obtain an encoded representation of the permutation residual; and compressing the encoded representation of the permutation residual to obtain a compressed representation of the encoded representation of the permutation residual, wherein the first neural network is trained by updating the obtained permutation residual to reduce the gradient of the determined rate-distortion loss.

2. The method according to claim 1, characterized in that, further includes: decompressing the compressed representation of the encoded representation of the permutation residual to obtain a decompressed representation of the encoded representation of the permutation residual; decoding the decompressed representation of the encoded representation of the permutation residual using a second neural network to obtain a recovered permutation residual; and adding the obtained recovered permutation residual to the obtained predicted image frame to obtain a reconstructed image frame.

3. The method according to claim 2, characterized in that, the first neural network and the second neural network are trained by the following operations: determining a distortion loss based on the obtained recovered permutation residual and the obtained permutation residual; determining a rate loss based on the compressed representation of the encoded representation of the permutation residual; determining the gradient of the rate-distortion loss based on the determined distortion loss, the determined rate loss and hyperparameters.

4. The method according to claim 3, characterized in that, the hyperparameters are set based on the target bit rate of the compressed representation of the encoded representation of the permutation residual without retraining the first neural network and the second neural network.

5. The method according to claim 3, characterized in that, the distortion loss is determined based on a function without retraining the first neural network and the second neural network, and the function is set based on the type of the target metric of the compressed representation of the encoded representation of the permutation residual.

6. The method according to claim 1, characterized in that, the estimating of the motion vector, the obtaining of the predicted image frame of the current image frame, the obtaining of the permutation residual, the obtaining of the encoded representation of the permutation residual and the compressing of the encoded representation of the permutation residual are performed by an encoding processor.

7. The method according to claim 6, characterized in that, further includes: decompressing the compressed representation of the encoded representation of the permutation residual by a decoding processor; decoding the decompressed representation of the encoded representation of the permutation residual by the decoding processor using a second neural network to obtain a recovered permutation residual; obtaining the predicted image frame by the decoding processor based on the estimated motion vector and the previously reconstructed image frame; and The obtained restored displacement residual is added to the obtained predicted image frame by the decoding processor to obtain a reconstructed image frame.

8. An apparatus for displacement neural residual compression, characterized in that the apparatus includes: at least one memory configured to store program code; and at least one processor configured to read the program code and operate according to the instructions of the program code, the program code including: estimation code configured to cause the at least one processor to estimate a motion vector based on a current image frame and a previously reconstructed image frame; first acquisition code configured to cause the at least one processor to obtain a predicted image frame based on the estimated motion vector and the previously reconstructed image frame; subtraction code configured to cause the at least one processor to subtract the obtained predicted image frame from the current image frame to obtain a displacement residual; encoding code configured to cause the at least one processor to encode the obtained displacement residual using a first neural network to obtain an encoded representation of the displacement residual; and compression code configured to cause the at least one processor to compress the encoded representation of the displacement residual to obtain a compressed representation of the encoded representation of the displacement residual, wherein the first neural network is trained by updating the obtained displacement residual to reduce the gradient of a determined rate-distortion loss.

9. The apparatus according to claim 8, characterized in that the program code further includes: decompression code configured to cause the at least one processor to decompress the compressed representation of the encoded representation of the displacement residual to obtain a decompressed representation of the encoded representation of the displacement residual; decoding code configured to cause the at least one processor to decode the decompressed representation of the encoded representation of the displacement residual using a second neural network to obtain a restored displacement residual; and addition code configured to cause the at least one processor to add the obtained restored displacement residual to the obtained predicted image frame to obtain a reconstructed image frame.

10. The apparatus according to claim 9, characterized in that the first neural network and the second neural network are trained by the following operations: determining a distortion loss based on the obtained restored displacement residual and the obtained displacement residual; determining a rate loss based on the compressed representation of the encoded representation of the displacement residual; determining a gradient of the rate-distortion loss based on the determined distortion loss, the determined rate loss, and hyperparameters.

11. The apparatus according to claim 10, characterized in that the hyperparameters are set based on a target bit rate of the compressed representation of the encoded representation of the displacement residual without retraining the first neural network and the second neural network.

12. The apparatus according to claim 10, characterized in that The distortion loss is determined based on a function without retraining the first neural network and the second neural network, and the function is set based on the type of the target metric of the compressed representation of the encoded representation of the displacement residual.

13. The apparatus according to claim 8, wherein, the estimated code, the first obtained code, the subtraction code, the encoding code, and the compression code are configured to enable an encoding processor to perform functions.

14. The apparatus according to claim 13, wherein, the program code further includes: a decompression code configured to enable a decoding processor to decompress the compressed representation of the encoded representation of the displacement residual; a decoding code configured to enable the decoding processor to decode the decompressed representation of the encoded representation of the displacement residual using a second neural network to obtain a recovered displacement residual; a second obtained code configured to enable the decoding processor to obtain the predicted image frame based on the estimated motion vector and the previously reconstructed image frame; and an addition code configured to enable the decoding processor to add the obtained recovered displacement residual to the obtained predicted image frame to obtain a reconstructed image frame.

15. An electronic device, wherein, comprising: a memory for storing computer-readable instructions; a processor for reading the computer-readable instructions and executing the method according to any one of claims 1 to 7 according to the instructions of the computer-readable instructions.

16. A method for displacement neural residual decompression, wherein, the method includes: estimating a motion vector based on a current image frame and a previously reconstructed image frame; obtaining a predicted image frame based on the estimated motion vector and the previously reconstructed image frame; subtracting the obtained predicted image frame from the current image frame to obtain a displacement residual; encoding the obtained displacement residual using a first neural network to obtain an encoded representation of the displacement residual; compressing the encoded representation of the displacement residual to obtain a compressed representation of the encoded representation of the displacement residual; decompressing the compressed representation to obtain a decompressed representation of the encoded representation of the displacement residual; decoding the decompressed representation using a second neural network to obtain a recovered displacement residual; adding the recovered displacement residual to the obtained predicted image frame to obtain a reconstructed image frame; wherein, the first neural network and the second neural network are trained by updating the obtained displacement residual to reduce the gradient of the determined rate-distortion loss.

17. A device for displacement neural residual decompression, wherein, the device includes: an estimation code configured to enable at least one processor to estimate a motion vector based on a current image frame and a previously reconstructed image frame; a first obtained code configured to enable at least one processor to obtain a predicted image frame based on the estimated motion vector and the previously reconstructed image frame; a subtraction code configured to enable at least one processor to subtract the obtained predicted image frame from the current image frame to obtain a displacement residual; Encoding code, configured to cause at least one processor to encode the obtained displacement residual using a first neural network to obtain an encoded representation of the displacement residual; Compression code, configured to cause at least one processor to compress the encoded representation of the displacement residual to obtain a compressed representation of the encoded representation of the displacement residual; Decompression code, for causing at least one processor to decompress the compressed representation to obtain a decompressed representation of the encoded representation of the displacement residual; Decoding code, for causing at least one processor to decode the decompressed representation using a second neural network to obtain a recovered displacement residual; Addition code, for causing at least one processor to add the obtained recovered displacement residual to the obtained predicted image frame to obtain a reconstructed image frame, wherein the first neural network and the second neural network are trained by updating the obtained displacement residual to reduce the gradient of the determined rate-distortion loss.

18. A method for storing or transmitting a video bitstream, characterized in that, the video bitstream is encoded according to the method according to any one of claims 1-7, or decoded according to the method according to claim 16.

Citation Information

Patent Citations

  • Video compression method and device and terminal equipment

    CN110753225A