Machine-based video coding enhancement process

The method enhances machine-based video coding by optimizing decoding and enhancement processes for human and machine tasks, addressing the variability in video content characteristics and improving viewing performance.

JP7742891B2Active Publication Date: 2025-09-22TENCENT AMERICA LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023564596
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-12-06
Filing Date
2023-01-05
Publication Date
2025-09-22
Estimated Expiration
2043-01-05

AI Technical Summary

Technical Problem

Existing machine-based video coding methods are optimized for specific types of video content and may not adequately address the unique characteristics of individual images/videos, requiring further enhancement for optimal performance in both human and machine viewing tasks.

Method used

A method and device for machine-based video coding (VCM) image enhancement that includes obtaining encoded images, decoding them using VCM decoding modules, and generating enhanced images using optimization parameters tailored for human, machine, or combined human-machine viewing tasks, with enhancement modules optimized for specific tasks.

Benefits of technology

Enhances decoded images for improved performance in both human and machine viewing tasks, addressing the variability in video content characteristics and optimizing for specific viewing requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007742891000009
    Figure 0007742891000009
  • Figure 0007742891000010
    Figure 0007742891000010
  • Figure 0007742891000011
    Figure 0007742891000011
Patent Text Reader

Abstract

A system, device and method for performing machine video coding (VCM) image enhancement, comprising: obtaining an encoded image from an encoded bitstream; obtaining enhancement parameters corresponding to the encoded image; decoding the encoded image using a VCM decoding module to generate a decoded image; generating an enhanced image using an enhancement module based on the decoded image and the enhancement parameters, wherein the enhancement parameters are optimized for one of a human-viewed VCM task, a machine-viewed VCM task, and a combined human-machine view VCM task; and providing at least one of the decoded image and the enhanced image to at least one of a human view module and a machine view module performing one of the human view VCM task, the machine view VCM task, and the combined human-machine view VCM task.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 313,616, filed February 24, 2022, and U.S. Application No. 18 / 076,020, filed December 6, 2022, both of which are incorporated herein by reference in their entireties.

[0002] SUMMARY OF THE INVENTION Embodiments of the present disclosure are directed to video coding, and in particular to image enhancement compatible with machine-based video coding (VCM). [Background technology]

[0003] Video and images may be used by human users for a variety of purposes (e.g., entertainment, education, etc.), and therefore video and image coding often exploits the properties of the human visual system to achieve high compression efficiency while maintaining high subjective quality.

[0004] With the increasing use of machine learning, many intelligent platforms have adopted video for machine vision tasks (e.g., object detection, segmentation, and tracking) in addition to numerous sensors. As a result, the coding of videos or images used in machine tasks has become an interesting and challenging issue. As a result, the idea of ​​video coding for machines (VCM) has been introduced. To achieve this goal, the international standards organization MPEG established an ad hoc group called "Video coding for machines (VCM)" to standardize related technologies to increase interoperability between different devices.

[0005] Existing VCM methods may be optimized for specific types of video content. For example, some implementations of VCM, either learning-based or hand-crafted, may be trained and optimized using a collection of image / video datasets. However, in actual encoding operations, individual images / videos may have their own unique characteristics that may deviate from the characteristics of the training image / video dataset. Therefore, further enhancement of the decoded images / videos may be required. Summary of the Invention [Means for solving the problem]

[0006] According to an aspect of the present disclosure, a method for performing machine video coding (VCM) image enhancement is executed by at least one processor and includes the steps of obtaining an encoded image from an encoded bitstream; obtaining enhancement parameters corresponding to the encoded image; decoding the encoded image using a VCM decoding module to generate a decoded image; generating an enhanced image using an enhancement module based on the decoded image and the enhancement parameters, wherein the enhancement parameters are optimized for one of a human-viewed VCM task, a machine-viewed VCM task, and a combined human-machine view VCM task; and providing at least one of the decoded image and the enhanced image to at least one of a human-viewed module and a machine-viewed module performing one of the human-viewed VCM task, the machine-viewed VCM task, and the combined human-machine view VCM task.

[0007] According to an aspect of the present disclosure, a device for performing machine-based video coding (VCM) image enhancement includes at least one memory configured to store program code; and at least one processor configured to load the program code and operate as directed by the program code, the program code including: first retrieval code configured to cause the at least one processor to retrieve an encoded image from an encoded bitstream; second retrieval code configured to cause the at least one processor to retrieve enhancement parameters corresponding to the encoded image; and second retrieval code configured to cause the at least one processor to decode the encoded image using a VCM decoding module to generate a decoded image. and at least one processor including: a decoding code configured to generate an enhanced image using an enhancement module based on the decoded image and enhancement parameters, the enhancement parameters being optimized for one of a human-viewed VCM task, a machine-viewed VCM task, and a combined human-machine view VCM task; and a providing code configured to cause the at least one processor to provide at least one of the decoded image and the enhanced image to at least one of a human-viewed module and a machine-viewed module that performs one of the human-viewed VCM task, the machine-viewed VCM task, and the combined human-machine view VCM task.

[0008] According to an aspect of the present disclosure, a non-transitory computer-readable medium includes one or more instructions that, when executed by one or more processors of a device for machine-based video coding (VCM) image enhancement, cause the one or more processors to obtain an encoded image from an encoded bitstream, obtain enhancement parameters corresponding to the encoded image, decode the encoded image with a VCM decoding module to generate a decoded image, generate an enhanced image with the enhancement module based on the decoded image and the enhancement parameters, the enhancement parameters being optimized for one of a human-viewed VCM task, a machine-viewed VCM task, and a combined human-machine view VCM task, and provide at least one of the decoded image and the enhanced image to at least one of a human-viewed module and a machine-viewed module performing one of the human-viewed VCM task, the machine-viewed VCM task, and the combined human-machine view VCM task.

[0009] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a diagram of an environment in which the methods, apparatus, and systems described in the present application may be implemented, according to an embodiment. [Figure 2] 2 is a block diagram of an example of elements of one or more devices of FIG. 1. [Figure 3] FIG. 1 is a block diagram of an example architecture for performing video encoding, according to an embodiment. [Figure 4] FIG. 1 is a block diagram of an example architecture for performing video encoding, including an enhancement module, according to an embodiment. [Figure 5] FIG. 1 is a block diagram of an example architecture for performing video encoding, including an enhancement module, according to an embodiment. [Figure 6]FIG. 1 is a block diagram of a module for determining mean squared error, according to an embodiment. [Figure 7A] FIG. 2 is a diagram illustrating an example of a bitstream according to an embodiment. [Figure 7B] FIG. 2 is a diagram illustrating an example of a bitstream according to an embodiment. [Figure 7C] FIG. 2 is a diagram illustrating an example of a bitstream according to an embodiment. [Figure 7D] FIG. 2 is a diagram illustrating an example of a bitstream according to an embodiment. [Figure 8] FIG. 2 is a block diagram of an example enhancement module, according to an embodiment. [Figure 9] 1 is a flowchart of an example process for performing feature compression, according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] Figure 1 is a diagram of an environment 100 in which the methods, apparatus, and systems described herein may be implemented, according to an embodiment. As shown in Figure 1, environment 100 may include a user device 110, a platform 120, and a network 130. The devices in environment 100 may be connected via wired connections, wireless connections, or a combination of wired and wireless connections.

[0012] User device 110 includes one or more devices capable of receiving, generating, storing, processing, and / or providing information related to platform 120. For example, user device 110 may include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smartphone, a wireless phone, etc.), a wearable device (e.g., smart glasses or a smart watch), or the like. In some implementations, user device 110 may receive information from and / or transmit information to platform 120.

[0013] Platform 120 includes one or more devices as described elsewhere herein. In some implementations, platform 120 may include a cloud server or a group of cloud servers. In some implementations, platform 120 may be designed to be modular such that software components can be swapped in and out depending on what is needed at the time. Thus, platform 120 can be easily and / or quickly reconfigured for different uses.

[0014] In some implementations, as shown, platform 120 may be hosted in a cloud computing environment 122. In particular, although the implementations described herein are described as platform 120 being hosted in a cloud computing environment 122, in some implementations platform 120 may be non-cloud based (i.e., implemented outside of a cloud computing environment) or may be partially cloud based.

[0015] Cloud computing environment 122 includes an environment that hosts platform 120. Cloud computing environment 122 may provide services such as computing, software, data access, storage, etc. that do not require end-user (e.g., user device 110) information about the physical location and configuration of one or more systems and / or one or more devices that host platform 120. As shown, cloud computing environment 122 may include a collection of computing resources 124 (collectively referred to as “computing resources 124” and individually referred to as “computing resource 124”).

[0016] Computational resources 124 may include one or more personal computers, workstation computers, server devices, or other types of computing and / or communication devices. In some implementations, computational resources 124 may provide hosting for platform 120. Cloud resources may include computing instances running on computational resources 124, storage devices provided by computational resources 124, data transfer devices provided by computational resources 124, etc. In some implementations, computational resources 124 may communicate with other computational resources 124 via wired connections, wireless connections, or a combination of wired and wireless connections.

[0017] As further shown in FIG. 1 , the computational resources 124 include a group of cloud resources such as one or more applications (APPs) 124-1, one or more virtual machines (VMs) 124-2, virtual storage (VS) 124-3, and one or more hypervisors (HYPs) 124-4.

[0018] Application 124-1 may include one or more software applications that may be provided to or accessed by user device 110 and / or platform 120. Application 124-1 may eliminate the need for software applications to be installed on user device 110 and run on user device 110. For example, application 124-1 may include software associated with platform 120 and / or any other software that may be provided via cloud computing environment 122. In some implementations, one application 124-1 may send information to one or more other applications 124-1 via virtual machine 124-2, or one application 124-1 may receive information from one or more other applications 124-1 via virtual machine 124-2.

[0019] Virtual machine 124-2 includes a software implementation of a machine (e.g., a computer) that executes programs like a physical machine. Virtual machine 124-2 may be either a system virtual machine or a process virtual machine, depending on the application and the degree to which virtual machine 124-2 matches any real-world machine. A system virtual machine can provide a complete system platform that supports the execution of a complete operating system (OS). A process virtual machine can execute only one program and support only one process. In some implementations, virtual machine 124-2 may execute on behalf of a user (e.g., user device 110) and manage the infrastructure of cloud computing environment 122, such as data management, synchronization, and long-term data transfers.

[0020] Virtual storage 124-3 includes one or more storage systems and / or one or more devices that employ virtualization techniques within the storage systems or devices of computing resources 124. In some implementations, with respect to storage systems, types of virtualization may include block virtualization and file virtualization. Block virtualization refers to the abstraction (or separation) of logical storage from physical storage, allowing access to storage systems whether they are physical or heterogeneous. Separation allows storage system administrators greater flexibility in how they manage storage for end users. File virtualization removes the dependency between data being accessed at a specific file level and where the file is physically stored. This may enable optimization of storage usage, server consolidation, and / or non-disruptive file migration.

[0021] The hypervisor 124-4 may provide hardware virtualization technology that allows multiple operating systems (e.g., "guest operating systems") to run simultaneously on a host computer, such as the computing resource 124. The hypervisor 124-4 may provide a virtual operating platform for the guest operating systems and may manage the execution of the guest operating systems. Multiple instances of different operating systems may share virtualized hardware resources.

[0022] Network 130 may include one or more wired and / or wireless networks. For example, network 130 may include a cellular network (e.g., a fifth-generation (5G) network, a long-term evolution (LTE) network, a third-generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a public switched telephone network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber optic-based network, etc., and / or a combination of these or other types of networks.

[0023] The number and arrangement of devices and networks shown in Figure 1 are shown by way of example. In practice, there may be more devices and / or networks, fewer devices and / or networks, different devices and / or networks, or different arrangements of devices and / or networks compared to those shown in Figure 1. Furthermore, two or more devices shown in Figure 1 may be implemented in a single device, or a single device shown in Figure 1 may be implemented as multiple distributed devices. Additionally or alternatively, a collection of devices (e.g., one or more devices) in environment 100 may perform one or more functions that are described as being performed by another collection of devices in environment 100.

[0024] Figure 2 is a block diagram of example elements of one or more devices of Figure 1. Device 200 may correspond to user device 110 and / or platform 120. As shown in Figure 2, device 200 may include a bus 210, a processor 220, a memory 230, a storage element 240, an input element 250, an output element 260, and a communication interface 270.

[0025] Bus 210 includes elements that enable communication between elements of device 200. Processor 220 may be implemented in hardware, firmware, or a combination of hardware and software. Processor 220 may be a central processing unit (CPU), graphics processing unit (GPU), accelerated processing unit (APU), microprocessor, microcontroller, digital signal processor (DSP), field programmable gate array (FPGA), application specific integrated circuit (ASIC), or another type of processor. In some implementations, processor 220 includes one or more processors that can be programmed to perform functions. Memory 230 includes random access memory (RAM), read-only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, and / or optical memory) that stores information and / or instructions used by processor 220.

[0026] Storage element 240 stores information and / or software related to the operation and use of device 200. For example, storage element 240 may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid-state disk), a compact disk (CD), a digital versatile disk (DVD), a floppy disk, a cartridge, a magnetic tape, and / or another type of non-transitory computer-readable medium with a corresponding drive.

[0027] Input elements 250 include elements that allow device 200 to receive information, such as via user input (e.g., a touchscreen display, a keyboard, a keypad, a mouse, buttons, switches, and / or a microphone). Additionally or alternatively, input elements 250 may include sensors that sense information (e.g., a global positioning system (GPS) element, an accelerometer, a gyroscope, and / or an actuator). Output elements 260 include elements that provide output information from device 200 (e.g., a display, a speaker, and / or one or more light-emitting diodes (LEDs)).

[0028] Communication interface 270 includes transceiver-like elements (e.g., a transceiver and / or separate receiver and transmitter) that enable device 200 to communicate with other devices, such as via a wired connection, a wireless connection, or a combination of wired and wireless connections. Communication interface 270 may be used to enable device 200 to receive information from and / or provide information to another device. For example, communication interface 270 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, etc.

[0029] Device 200 may perform one or more of the processes described herein. Device 200 may perform the processes in response to processor 220 executing software instructions stored by a non-transitory computer-readable medium, such as memory 230 and / or storage element 240. A computer-readable medium is defined in this application as a non-transitory memory device. A memory device includes storage space in one physical storage device or across multiple physical storage devices.

[0030] The software instructions may be read from another computer-readable medium or from another device via communications interface 270 and loaded into memory 230 and / or storage element 240. When executed, the software instructions stored in memory 230 and / or storage element 240 may be used to cause processor 220 to perform one or more processes described herein. Additionally or alternatively, hardwired circuitry may be used in place of or in combination with software instructions to perform one or more processes described herein. Thus, implementations described herein are not limited to any specific combination of hardware circuitry and software.

[0031] The number and arrangement of elements shown in Figure 2 are provided as an example. In practice, device 200 may include more, fewer, different, or differently arranged elements than those shown in Figure 2. Additionally or alternatively, a collection of elements (e.g., one or more elements) of device 200 may perform one or more functions that are described as being performed by another collection of elements of device 200.

[0032] 3 is a block diagram of an example architecture 300 for performing video coding, according to an embodiment. In an embodiment, architecture 300 may be a machine-based video coding (VCM) architecture or any other architecture compatible with or configured to perform VCM coding. For example, architecture 300 may be compatible with "Use cases and requirements for Video Coding for Machines" (ISO / IEC JTC 1 / SC 29 / WG 2 N18), "Draft of Evaluation Framework for Video Coding for Machines" (ISO / IEC JTC 1 / SC 29 / WG 2 N19), and "Call for Evidence for Video Coding for Machines" (ISO / IEC JTC 1 / SC 29 / WG 2 N20), the disclosures of which are incorporated herein by reference in their entireties.

[0033] In an embodiment, one or more of the elements illustrated in FIG. 3 may correspond to or be implemented by one or more of the elements described above with respect to FIGS. 1-2, for example, one or more of user device 110, platform 120, device 200, or any of the elements included therein.

[0034] As can be seen in FIG. 3 , architecture 300 may include a VCM encoder 310 and a VCM decoder 320. In an embodiment, the VCM encoder may receive a sensor input 301, which may include, for example, one or more input images or input videos. The sensor input 301 may be provided to a feature extraction module 311, which may extract features from the sensor input, and the extracted features may be transformed using a feature transformation module 312 and encoded using a feature encoding module 313. In an embodiment, the term “encoding” may include, correspond to, or be used instead of the term “compression.” Architecture 300 may include an interface 302, which may enable feature extraction module 311 to interface with a neural network (NN), which may assist in performing feature extraction.

[0035] The sensor input 301 may be provided to a video encoding module 314, which may generate coded video. In embodiments, after feature extraction, transformation, and encoding, the coded features may be provided to the video encoding module 314, which may use the coded features to help generate coded video. In embodiments, the video encoding module 314 may output the coded video as a coded video bitstream, and the feature encoding module 313 may output the coded features as a coded feature bitstream. In embodiments, the VCM encoder 310 may provide both the coded video bitstream and the coded feature bitstream to a bitstream multiplexer 315, which may generate the coded bitstream by combining the coded video bitstream and the coded feature bitstream.

[0036] In embodiments, the encoded bitstream may be received by a bitstream demultiplexer (demux), which may separate the encoded bitstream into an encoded video bitstream and an encoded feature bitstream, which may be provided to a VCM decoder 320. The encoded feature bitstream may be provided to a feature decoding module 322, which may generate decoded features, and the encoded video bitstream may be provided to a video decoding module, which may generate decoded video. In embodiments, the decoded features may be further provided to a video decoding module 323, which may use the decoded features to help generate the decoded video.

[0037] In embodiments, the outputs of the video decoding module 323 and the feature decoding module 322 may be primarily used for machine use, e.g., by a machine vision module 332. In embodiments, the outputs may also be used for human use, as illustrated in FIG. 3 as a human vision module 331. A VCM system, e.g., architecture 300, from the client side, e.g., VCM decoder 320, may first perform video decoding to obtain a sample domain video. One or more machine tasks of understanding the video content may then be performed, e.g., by the machine vision module 332. In embodiments, the architecture 300 may include an interface 303 that may enable the machine vision module 332 to interface with a NN that may assist in performing the one or more machine tasks.

[0038] As can be seen from FIG. 3, in addition to the video encoding / decoding path including the video encoding module 314 and the video decoding module 323, another path included in the architecture 300 may be a feature extraction, feature encoding, feature decoding path, which includes a feature extraction module 311, a feature transformation module 312, a feature encoding module 313, and a feature decoding module 322.

[0039] Embodiments may relate to methods for enhancing decoded video for machine viewing, human viewing, or combined human / machine viewing. In embodiments, each decoded image may be generated by, for example, a VCM decoder 320, and each image may be enhanced for machine viewing or human viewing using an enhancement module and metadata provided by the encoder. In embodiments, these methods may be applied to any VCM codec. While some embodiments may be described using broad terms such as "image / video" or more specific terms such as "image" or "video," it is understood that the embodiments may be applied.

[0040] FIG. 4 is a block diagram of an example architecture 400 for performing video encoding, according to an embodiment. As shown in FIG. 4, the architecture 400 may include an enhancement module 402. The output of the VCM decoder 320 for the decoded image / video and metadata generated by the VCM encoder 310 may be provided to the enhancement module 402 to generate enhanced images and / or enhanced video, which may be used for machine or human vision tasks. In an embodiment, the metadata may include parameters of the enhancement module 402, which may be referred to as enhancement parameters, for example. In an embodiment, the enhancement parameters may be, for example, parameters used to configure the enhancement module 402 to perform an enhancement process, such as those described below. In an embodiment, the enhancement module 402 may be an image / video processing module. In an embodiment, the enhancement module 402 may be a neural network-based processing module. Depending on the viewing task, either the decoded images / video or the enhanced images / video may be selected and provided to the human viewing module 331 and the machine viewing module 332. While this selection is illustrated in Figure 4 using a switch 404, embodiments are not so limited and other techniques for selectively providing the decoded images / video and the enhanced images / video may be used.

[0041] In embodiments, the transmission of metadata may be performed as appropriate, for example, if the decoded image / video is to be used by the machine vision module 332, the decoder may inform the VCM encoder 310 not to send metadata since the decoder will not use it.

[0042] In an embodiment, the enhancement parameters may be fixed, thus eliminating the need to send metadata.

[0043] In embodiments, VCM encoder 310 and VCM decoder 320 may be optimized for a machine task, e.g., a task corresponding to machine vision module 332. In embodiments, enhancement module 402 may be designed to refine decoded images / video for a human vision task, e.g., a task corresponding to human vision module 331. In embodiments, enhancement module 402 may be designed to further refine decoded images / video for a machine vision task. In embodiments, enhancement module 402 may be designed to refine decoded images / video for a combined machine / human vision task, e.g., a task corresponding to both machine vision module 332 and human vision module 331. The enhancement parameters may be different for different tasks.

[0044] In embodiments, the enhancement module 402 may be or include a neural network, and the VCM encoder 310 may optimize the parameters of the neural network to perform well on a machine vision task, a human vision task, or a combined machine / human vision task. In embodiments, a rate-distortion optimization technique may be used. In embodiments, the parameters of the neural network may be optimized based on enhancement parameters, e.g., metadata, provided by the VCM encoder 310. In embodiments, the parameters of the neural network may be included directly in the enhancement parameters. In embodiments, the enhancement parameters may specify modifications to the neural network parameters or may include information that allows the neural network parameters to be derived.

[0045] 5 is a block diagram of an example architecture 500 for performing video encoding, according to an embodiment. As shown in FIG. 5, the architecture 500 may include an enhancement module 402, and may be configured to perform image enhancement and / or video enhancement using rate-distortion optimization techniques.

[0046] In embodiments, the rate-distortion optimization process may be performed on the encoder side, e.g., by the VCM encoder 310 or by other elements associated with the VCM encoder 310. In embodiments, a distortion metric D between an input image and its corresponding enhanced image may be calculated, and a parameter size R for the enhancement parameters may be determined. An overall loss function L loss may be expressed using the following Equation 1: L loss =R+λD (Equation 1)

[0047] In Equation 1, λ can be used to set the tradeoff (distortion D vs. rate R). While Figure 5 is illustrated with respect to images, it will be appreciated that embodiments may also be applied to videos, e.g., by treating the video as a sequence of images when calculating the distortion metric.

[0048] In embodiments, the VCM encoder 310 may optimize the enhancement parameters using gradient descent or its variations. In embodiments, the optimized enhancement parameters may be obtained for each image, and the optimized enhancement parameters may be metadata sent to the decoder side, or the optimized enhancement parameters may be included in the metadata sent to the decoder side. In embodiments, the enhancement parameters may be constant for multiple images, such as a group of images (e.g., a group of pictures (GOP)). For example, the distortion metric may be set as the average distortion of the GOP or group of images. Metadata, such as the enhancement parameters, may be shared across the GOP. Therefore, the metadata size can be reduced.

[0049] For human viewing, the distortion metrics may include one or more of mean squared error (MSE), 1-ssim, or 1-ms_ssim, where ssim denotes the structural similarity metric between the input image and the enhanced image (SSIM), and ms_ssim denotes the multi-scale structural similarity metric between the input image and the enhanced image (MS-SSIM).

[0050] In the case of machine viewing, ssim or ms_ssim may be strongly correlated with better performance in machine viewing tasks, so 1-ssim or 1-ms_ssim may also be used in this case.

[0051] FIG. 6 is a block diagram of an error determination module 600, according to an embodiment. In an embodiment, the error determination module 600 may be used to determine the MSE in feature space between the enhanced image and the input image. In an embodiment, the MSE in feature space may be used as a distortion metric for machine vision. As shown in FIG. 6 , the error determination module may include a feature extraction neural network 602 that may be used to extract features from the input image and a feature extraction neural network 604 that may be used to extract features from the enhanced image. In an embodiment, one or more of the feature extraction neural network 602 and the feature extraction neural network 604 may correspond to one or more elements included in the VCM encoder 310, such as the feature extraction module 311. In an embodiment, the MSE may be calculated by the MSE module 606, for example, according to Equation 2 below:

number

[0052] In the above equation 2, f(c,h,w) represents the features of the input image,

number

[0053] 6, feature extraction neural networks 602 and 604 may be simple, for example, the first few layers of a machine task network. In embodiments, if the machine analysis network is known on the encoder side, it may be effective to use the first few layers of a given machine analysis network as at least one of feature extraction neural networks 602 and 604. In embodiments, if the machine analysis network is unknown on the encoder side, the first few layers of a commonly used machine analysis network such as Faster R-CNN, Mask R-CNN, or VGG-16 may be used as at least one of feature extraction neural networks 602 and 604.

[0054] In an embodiment, there may be multiple ways to send metadata representing parameters of enhancement module 402. Figures 7A-7D illustrate example bitstreams including metadata representing enhancement parameters, according to an embodiment.

[0055] In an embodiment, metadata for a particular image may be included in a bitstream containing encoded image data corresponding to the particular image, e.g., a bitstream containing image data for image 1 through image k may also include at least one of the corresponding metadata, e.g., metadata 1 through metadata k.

[0056] In an embodiment, as shown in FIG. 7A, in the bitstream, the portion of the bitstream corresponding to image 1 may be attached to, adjacent to, or otherwise associated with the portion of the bitstream corresponding to the metadata for image 1, and similarly for images 2 through k.

[0057] In embodiments, metadata may be selectively included. As an example, a flag F may be used to indicate whether metadata is attached for a particular image, e.g., as shown in FIG. 7B. In embodiments, if no metadata is attached for a given image, the enhancement parameters may remain unchanged. If metadata is attached for a given image, the enhancement parameters may be modified according to information in the metadata. An example of such an arrangement is shown in FIG. 7B. As can be seen in FIG. 7B, each image may have a corresponding flag F that indicates whether the bitstream includes metadata, such as enhancement parameters for the image.

[0058] In embodiments, flag F can be represented by one bit because flag F has a value of 0 or 1, thereby reducing the overhead introduced by flag F. To further reduce overhead, flag F may be entropy coded with or without a context model. In embodiments, flag F may be represented by a single byte or multiple bits that indicate which set of known parameters the decoder should use. This may be useful if the decoder, e.g., VCM decoder 320 or enhancement module 402, stores or receives one or more sets of enhancement parameters.

[0059] In an embodiment, if a GOP contains one set of metadata and the GOP size is freely selectable, selective metadata attachment may be used to carry the metadata. In an embodiment, if the GOP size is constant, for example, if every GOP contains K images, metadata may be attached to the beginning or end of every Kth image without using flag F.

[0060] As shown in Figures 7A and 7B, the metadata and flag F may be appended to or included in the bitstream of the corresponding image, but embodiments are not limited in this respect. For example, in other embodiments, flag F may be placed before the corresponding bitstream, or both flag F and the associated metadata may be placed before the corresponding bitstream.

[0061] In an embodiment, the metadata may be sent separately from the main bitstream of the coded image / video, e.g., in a separate bitstream. For example, Figure 7C shows example metadata bitstreams for each of images 1 to k, and Figure 7D shows an example metadata bitstream for flag F, which is used to indicate whether metadata is present for a particular image.

[0062] As explained above, according to an embodiment, the enhancement module 402 may be a neural network. Depending on the implementation complexity and performance requirements, the enhancement module 402 may be simple or complex.

[0063] Figure 8 shows a diagram of an example, relatively simple enhancement module 402. For example, as shown in Figure 8, the enhancement module 402 may include a simple convolutional layer with a kernel size of 3x3 and a stride of 1.

[0064] The enhancement module 402 shown in Figure 8 may use a color image, and therefore the size of both the decoded image and the enhancement image is 3 x W x H. The 3 indicates three color channels, e.g., the three color channels of RGB, YCrCb, or other color formats. The enhancement module 402 shown in Figure 8 may use three filters, and each filter may include 3 x 3 x 3 weight values ​​and one bias value. Therefore, a total of 84 parameters (e.g., 3 x 3 x 3 x 3 + 3 = 84) may be used.

[0065] In an embodiment, when the bit rate of the bitstream generated by VCM encoder 310 is high, the parameter size of the enhancement parameters may be larger compared to when the rate is low. For example, as shown in FIG. 8, when the rate is low, the convolution kernel may be 3×3, whereas when the bit rate is high, the convolution kernel may be 5×5 or 7×7.

[0066] Typically, neural network parameters may be represented as 32-bit floating point numbers. In embodiments, the enhancement parameters may be represented with a smaller bit depth precision, such as 16-bit floating point numbers, to reduce metadata size. In embodiments, the enhancement parameters of the kth image may be expressed as

number

[0067] In an embodiment, the N numbers may be transmitted as metadata, along with the enhancement parameters of the kth image, e.g.

number

number

[0068] In an embodiment, the difference between the new enhancement parameters and one of the sets of known parameters may be transmitted as metadata. For example, the difference between the parameters of the kth image and the previous image may be determined according to Equation 4 below and transmitted as metadata.

number

[0069] While the above equations 3 and 4 may be considered to correspond to an embodiment in which metadata is sent for each image, the embodiment is not limited to this. For example, similar methods can be applied to an example in which metadata is selected and attached to an image, or an example in which metadata is shared across a GOP.

[0070] As shown in FIG. 9, process 900 may include a block (block 902) for obtaining an encoded image from an encoded bitstream.

[0071] 9, process 900 may include a block (block 904) of obtaining enhancement parameters corresponding to the encoded image. In an embodiment, the enhancement parameters may correspond to the enhancement parameters described above. In an embodiment, the enhancement parameters may be optimized for at least one of a human-vision VCM task, a machine-vision VCM task, and a combined human-machine vision VCM task.

[0072] 9, process 900 may include a block (block 906) that decodes the encoded image using a VCM decoding module to generate a decoded image. In an embodiment, the VCM decoding module may correspond to VCM decoder 320 described above.

[0073] 9, process 900 may include a block (block 908) for generating an enhanced image using an enhancement module based on the decoded image and the enhancement parameters. In an embodiment, the enhancement module may correspond to enhancement module 402 described above.

[0074] 9, process 900 may include a block (block 910) that provides the decoded image and / or the enhanced image to a human-vision module and / or a machine-vision module that performs a VCM task, e.g., one of a human-vision VCM task, a machine-vision VCM task, and a combined human-machine vision VCM task. In an embodiment, the human-vision module may correspond to human-vision module 331 described above, and the machine-vision module may correspond to machine-vision module 332 described above.

[0075] In embodiments, the enhancement module may include a neural network, and the enhancement parameters may include neural network parameters corresponding to the neural network.

[0076] In an embodiment, the enhanced image may be the product of rate-distortion optimization, and neural network parameters may be selected based on a distortion metric and parameter size.

[0077] In embodiments, the distortion metric may include at least one of a mean squared error, a structural similarity metric, and a multi-scale structural similarity metric associated with the enhanced image and the input image.

[0078] In an embodiment, the mean squared error may be calculated using Equation 2 described above.

[0079] In an embodiment, the decoded pictures may be included in a group of pictures (GOP) corresponding to the encoded bitstream, and all pictures included in the GOP share enhancement parameters.

[0080] In an embodiment, the enhancement parameters may be included in the encoded bitstream.

[0081] In an embodiment, the encoded bitstream may include a flag corresponding to the encoded image, and the flag may indicate whether enhancement parameters corresponding to the encoded image are included in the encoded bitstream.

[0082] In an embodiment, the enhancement parameters may be included in a metadata bitstream that is separate from the encoded bitstream.

[0083] In an embodiment, the metadata bitstream may include a flag corresponding to the encoded image, and the flag may indicate whether enhancement parameters corresponding to the encoded image are included in the metadata bitstream.

[0084] Although Figure 9 illustrates example blocks of process 900, in some implementations, process 900 may include additional blocks, fewer blocks, different blocks, or blocks in a different arrangement than those shown in Figure 9. Additionally or alternatively, two or more of the blocks of process 900 may be performed in parallel.

[0085] Additionally, the proposed methods may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, a program stored on a non-transitory computer-readable medium is executed by one or more processors to perform one or more of the proposed methods.

[0086] The techniques described above can be implemented using computer-readable instructions and as computer software physically stored on one or more computer-readable mediums.

[0087] The embodiments of the present disclosure may be used individually or combined in any order. Furthermore, each of the embodiments (and methods thereof) may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.

[0088] While the foregoing disclosure has been illustrated and described, it is not intended to be exhaustive or to limit implementations to the precise forms disclosed. Modifications and variations are possible in light of the above disclosure and may be acquired from practicing the implementations.

[0089] The term element as used in this application is intended to be broadly construed as hardware, firmware, or a combination of hardware and software.

[0090] Although a combination of features may be recited in a claim and / or disclosed in this specification, such combination is not intended to limit the disclosure of possible implementations. In fact, many of these features can be combined in ways not explicitly recited in the claims and / or disclosed in the specification. Each dependent claim listed below may depend directly on only one claim, but each dependent claim in combination with any other claim in the set of claims will be included in the disclosure of possible implementations.

[0091] No element, act, or instruction used in this application should be construed as essential or required unless expressly stated otherwise. Additionally, as used in this application, the articles "a" and "an" are intended to include one or more items and can also be used to mean "one or more." Additionally, as used in this application, the term "collection" is intended to include one or more items (e.g., related items, unrelated items, or a combination of related and unrelated items) and can also be used to mean "one or more." Where only one item is intended, the term "one" or similar description is used. Additionally, as used in this application, the terms "has," "have," "having," etc. are intended to be open-ended terms. Additionally, the phrase "based on" is intended to mean "based at least in part on," unless expressly stated otherwise. [Explanation of symbols]

[0092] 100 Environment 110 User Devices 120 Platform 122 Cloud Computing Environment 124-1 Application 124-4 Hypervisor 124-3 Virtual Storage 124-2 Virtual Machine 124 Computational Resources 130 Network 200 devices 210 Bus 220 processors 230 memory 240 Memory Elements 250 input elements 260 Output Elements 270 Communication Interface 300 Architecture 301 Sensor Input 302 Interface to NN 303 Interface to NN 310 VCM Encoder 311 Feature Extraction Module 312 Feature Transformation Module 313 Feature Encoding Module 314 Video Encoding Module 315 Bitstream Multiplexer 320 VCM decoder 322 Feature Decoding Module 323 Video Decoding Module 331 Human Vision Module 332 Machine Vision Module 400 Architecture 402 Enhancement Module 404 Switch 500 Architecture 600 Error Determination Module 602 Feature Extraction Neural Network 604 Feature Extraction Neural Network 606 MSE Module 900 processes

Claims

1. 1. A method for performing machine-based video coding (VCM) image enhancement, the method being performed by at least one processor and comprising: obtaining an encoded image from the encoded bitstream; obtaining enhancement parameters corresponding to the encoded image; decoding the encoded image using a VCM decoding module to generate a decoded image; generating an enhanced image using an enhancement module based on the decoded image and the enhancement parameters, wherein the enhancement parameters are optimized for one of a human-vision VCM task, a machine-vision VCM task, and a combined human-machine vision VCM task; providing at least one of the decoded image and the enhanced image to at least one of a human-viewed module and a machine-viewed module that executes the one of the human-viewed VCM task, the machine-viewed VCM task, and the combined human-machine-viewed VCM task; the enhancement parameters are included in a metadata bitstream; the enhancement parameters are included in the encoded bitstream; the encoded bitstream comprises flags corresponding to the encoded images; the flag indicates whether the enhancement parameters corresponding to the coded image are included in the coded bitstream; method.

2. the enhancement module comprises a neural network; The method of claim 1 , wherein the enhancement parameters comprise neural network parameters corresponding to the neural network.

3. the enhanced image is a product of rate-distortion optimization; The method of claim 2 , wherein the neural network parameters are selected based on a distortion metric and a parameter size.

4. The method of claim 3 , wherein the distortion metric comprises at least one of a mean squared error, a structural similarity metric, and a multi-scale structural similarity metric associated with the enhanced image and an input image.

5. The mean square error is expressed by the formula [Equation 1] is calculated using MSE represents the mean square error, f(c,h,w) represents the features of the input image, [Equation 2] 5. The method of claim 4, wherein C represents the number of channels in a feature map, H represents the number of rows in the feature map, W represents the number of columns in the feature map, c represents a channel index, h represents a row, and w represents a column position.

6. the decoded image is included in a group of pictures (GOP) corresponding to the encoded bitstream; The method of claim 1 , wherein all images included in the GOP share the enhancement parameters.

7. The method of claim 1 , wherein the enhancement parameters are included in a metadata bitstream that is separate from the encoded bitstream.

8. the metadata bitstream comprises flags corresponding to the encoded images; The method of claim 7 , wherein the flag indicates whether the enhancement parameters corresponding to the encoded image are included in the metadata bitstream.

9. A device configured to perform the method according to any one of claims 1 to 7.

10. A computer program product for causing one or more processors to carry out the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method for improving image quality

    JP2020010331A

  • Backward-compatible HDR codecs with temporal scalability

    US20170295382A1

  • Feature-Domain Residual for Video Coding for Machines

    US20210314573A1