Neural network in-loop filter for machine tasks

A neural network in-loop filter (NNLF) is used to dynamically enhance video coding for machine tasks, addressing the challenge of balancing compression efficiency and quality by applying filtering based on VCM processing outputs, improving machine vision performance.

WO2025221773A1PCT designated stage Publication Date: 2025-10-23TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/024747
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-04-14
Filing Date
2025-04-15
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Existing video coding technologies struggle to effectively balance compression efficiency and subjective quality for both human and machine tasks, particularly in video coding for machines, where machine vision tasks require specialized filtering techniques.

Method used

Implementing a neural network in-loop filter (NNLF) that can be dynamically enabled or disabled based on output from VCM processing modules, with parameters signaled in the bitstream or derived during decoding, to enhance reconstructed video for machine tasks.

Benefits of technology

The NNLF improves the quality of reconstructed video for machine tasks by selectively applying filtering techniques, optimizing compression efficiency and subjective quality based on Region of Interest (Rol) space thresholds, thereby enhancing machine vision performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025024747_23102025_PF_FP_ABST
    Figure US2025024747_23102025_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure relates generally to video coding, and more particularly to in-loop filtering for video coding for machine tasks based on neural networks. For example, utilization of one or more neural network in-loop filter (NNLF) may be determined (either enabled or disabled) at various coding level. Such determination may be based on output of one or more other VCM- codec-related processed. The usages of the NNLF may be signaled in the bitstream or may be implicitly derived from the output of the one or more other VCM-coded-related processes during decoding process in a decoder or during in-loop decoding process of an encoder.
Need to check novelty before this filing date? Find Prior Art

Description

NEURAL NETWORK IN-LOOP FILTER FOR MACHINE TASKS

[0001] This PCT application is based on and claims the benefit of priority to U.S. NonProvisional Patent Application No. 19 / 178,394 filed on April 14, 2025 and U.S. Provisional Patent Application No. 63 / 634,864 filed on April 16, 2024, both entitled “NEURAL NETWORK IN-LOOP FILTER FOR MACHINE TASKS,” which are herein incorporated by reference in their entireties.TECHNICAL FIELD

[0002] This disclosure relates generally to video coding, and more particularly to inloop filtering for video coding for machine tasks based on neural networks.BACKGROUND

[0003] Video or images may be consumed by human users for a variety of purposes, for example entertainment, education, etc. Thus, video coding or image coding may often utilize characteristics of human visual systems for better compression efficiency while maintaining good subjective quality.

[0004] With the rise of machine learning applications, along with the abundance of sensors, many intelligent platforms have utilized video for machine vision tasks such as object detection, video / image segmentation or object tracking. As a result, encoding video or images for consumption by machine tasks has become an interesting and challenging problem. This has led to the introduction of Video Coding for Machines (VCM) studies.

[0005] While the various embodiments below are described in the context of VCM, the underlying principles are generally applicable to other video coding systems.SUMMARY

[0006] This disclosure relates generally to video coding, and more particularly to inloop filtering for video coding for machine tasks based on neural networks. For example, utilization of one or more neural network in-loop filter (NNLF) may be determined (either enabled or disabled) at various coding level. Such determination may be based on output of one or more other VCM-codec-related processed. The usages of the NNLF may be signaledin the bitstream or may be implicitly derived from the output of the one or more other VCM- coded-related processes during decoding process in a decoder or during in-loop decoding process of an encoder.

[0007] In some example implementations, a method for decoding a video is disclosed. The method may include receiving a bitstream of the video; decoding the bitstream to obtain a reconstructed video; when it is determined that a neural network in-loop filter (NNLF) is to be applied to a portion of the reconstructed video, identifying the NNLF and extracting from the reconstructed video a set of filter parameters associated with the NNLF; and applying the NNLF with the set of filter parameters to the portion of the reconstructed video to generate a filtered video portion.

[0008] In the example implementations above, the NNLF is determined always applied to an entirety of the video.

[0009] In any one of the example implementations above, the method may further include extracting a signaling information item from the bitstream which indicates whether the NNLF is to be applied to the portion of the reconstructed video.

[0010] In any one of the example implementations above, the signaling information item is included in a Sequence Parameter Set (SPS), a Picture Parameter Set (PPS), a Sequence Header (SH), a Picture Header (PH), or Supplemental Enhancement Information (SEI) in the bitstream.

[0011] In any one of the example implementations above, whether the NNLF is to be applied to the portion of the reconstructed video is controlled by an output of at least one VCM (Video Coding for Machine) processing module applied to the bitstream.

[0012] In any one of the example implementations above, the at least one VCM processing module comprises one or more of a Region of Interest (Rol) processing module, a bit depth truncation module, a spatial sampling module, or temporal sampling module.

[0013] In any one of the example implementations above, whether the NNLF is to be applied to the portion of the reconstructed video is determined at a sequence or sub-sequence level.

[0014] In any one of the example implementations above, the at least one VCM processing module comprises an Rol processing module and the NNLF is or is not applied to a sequence or subsequence when a total Rol space of each frame within the sequence or subsequence is greater or less than a predetermined Rol space threshold, respectively.

[0015] In any one of the example implementations above, the at least one VCM processing module comprises an Rol processing module and the NNLF is or is not applied toa sequence or subsequence when Rol space of majority of frames within the sequence or subsequence is greater or less than a predetermined Rol space threshold, respectively.

[0016] In any one of the example implementations above, the at least one VCM processing module comprises an Rol processing module and the NNLF is or is not applied to a sequence or subsequence when Rol space of at least one of frames within the sequence or subsequence is greater or smaller than a predetermined Rol space threshold, respectively.

[0017] In any one of the example implementations above, whether the NNLF is to be applied to the portion of the reconstructed video is determined at a frame or slice level.

[0018] In any one of the example implementations above, the at least one VCM processing module comprises an Rol processing module and the NNLF is or is not applied to a frame or slice when Rol space of the frame or slice is greater or smaller than a predetermined Rol space threshold, respectively.

[0019] In some other example implementations, a method for processing a video is disclosed. The method may include encoding a portion of the video as part of an encoded bitstream; reconstructing the encoded portion of the video in an in-loop process to generate a reconstructed portion of the video; when it is determined that an NNLF is to be applied to the reconstructed portion of the video, identifying the NNLF and determining a set of filter parameters associated with the NNLF; and applying the NNLF with the set of filter parameters to the reconstructed portion of the video to generate a filtered video portion in the in-loop process.

[0020] In the example implementations above, the NNLF is determined as always applied to an entirety of the video.

[0021] In any one of the example implementations above, the method may further comprise including in the encoded bitstream a signaling information item for indicating whether the NNLF is applied in the in-loop process.

[0022] In any one of the example implementations above, the signaling information item is included in a Sequence Parameter Set (SPS), a Picture Parameter Set (PPS), a Sequence Header (SH), a Picture Header (PH), or Supplemental Enhancement Information (SEI) in the bitstream.

[0023] In any one of the example implementations above, whether the NNLF is applied to the reconstructed portion of the video is controlled by an output of at least one VCM processing module applied to the reconstructed portion of the video in the in-loop process.

[0024] In any one of the example implementations above, whether the NNLF is applied to the reconstructed portion of the video is determined at a sequence or sub-sequence level;and the method further comprising determining to apply the NNLF and including a corresponding signaling indicator in the encoded bitstream for the sequence or subsequence when: a total Rol space of each frame within the sequence or subsequence is greater than a predetermined Rol space threshold; when Rol space of majority of frames within the sequence or subsequence is greater than a predetermined Rol space threshold; or when Rol space of at least one of frames within the sequence or subsequence is greater than a predetermined Rol space threshold.

[0025] In any one of the example implementations above, whether the NNLF is applied to the reconstructed portion of the video is determined at a level of a frame or slice; and the method further comprising determining to apply the NNLF and including a corresponding signaling indicator in the encoded bitstream for the frame or slice when a total Rol space of the frame or slice is greater than a predetermined Rol space threshold.

[0026] In some other example implementations above, a non-transitory computer readable storage medium for storing a bitstream of a video is disclosed. The bitstream may include an encoded portion of the video; an indicator for signaling to a decoder whether an NNLF is to be applied to a reconstruction of the encoded portion of the video; and when the indicator signals that the NNLF is to be applied, a set of filter parameters associated with the NNLF.

[0027] Aspects of the disclosure also provide an electronic device or apparatus function as encoder or decoder including a circuitry configured to carry out any of the method implementations above.

[0028] Aspects of the disclosure also provide non-transitory computer-readable medium for storing computer instructions which when executed by at least one processor of a video processing device, cause the video processing device to perform any one of the method implementations above.BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Further features, the nature, and various advantages of the disclosed subject matter will be more apparent from the following detailed description and the accompanying drawings in which:

[0030] FIG. 1 is a diagram of an environment in which methods, apparatuses, and systems described herein may be implemented, according to embodiments.

[0031] FIG. 2 is a schematic illustration of an example computer system in accordance with an embodiment.

[0032] FIG. 3 is a block diagram of an example architecture for performing video coding, according to embodiments.

[0033] FIG. 4 illustrates various decoding modules that may be utilized in a VCM decoder.

[0034] FIG. 5 illustrates an example process for deblocking filtering or post filtering in video coding utilizing NNLF.

[0035] FIG. 6 is a flowchart of an example process for video decoding.

[0036] FIG. 7 is a flowchart of an example process for encoding a video.DETAILED DESCRIPTION OF EMBODIMENTS

[0037] Throughout this specification and claims, terms may have nuanced meanings suggested or implied in contexts beyond an explicitly stated meaning. The phrase “in one embodiment” or “in some embodiments” as used herein does not necessarily refer to the same embodiment and the phrase “in another embodiment” or “in other embodiments” as used herein does not necessarily refer to a different embodiment. Likewise, the phrase “in one implementation” or “in some implementations” as used herein does not necessarily refer to the same implementation and the phrase “in another implementation” or “in other implementations” as used herein does not necessarily refer to a different implementation. It is intended, for example, that claimed subject matter includes combinations of exemplary embodiments / implementations in whole or in part.

[0038] In general, terminology may be understood at least in part from usage in context. For example, terms, such as “and”, “or”, or “and / or,” as used herein may include a variety of meanings that may depend at least in part upon the context in which such terms are used. Typically, “or” if used to associate a list, such as A, B or C, is intended to mean A, B, and C, here used in the inclusive sense, as well as A, B or C, here used in the exclusive sense. In addition, the term “one or more” or “at least one” as used herein, depending at least in part upon context, may be used to describe any feature, structure, or characteristic in a singular sense or may be used to describe combinations of features, structures or characteristics in a plural sense. Similarly, terms, such as “a”, “an”, or “the”, again, may be understood to convey a singular usage or to convey a plural usage, depending at least in part upon context. In addition, the term “based on” or “determined by” may be understood as not necessarilyintended to convey an exclusive set of factors and may, instead, allow for existence of additional factors not necessarily expressly described, again, depending at least in part on context.

[0039] FIG. 1 is a diagram of an application environment 100 in which methods, apparatuses, and systems described herein may be implemented, according to the example embodiments. As shown in FIG. 1, the environment 100 may include a user device 110, a platform 120, and a network 130. Devices of the environment 100 may interconnect via wired connections, wireless connections, or a combination of wired and wireless connections.

[0040] The user device 110 may include one or more devices capable of receiving, generating, storing, processing, and / or providing information associated with platform 120. For example, the user device 110 may include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smart phone, a radiotelephone, etc.), a wearable device (e.g., a pair of smart glasses or a smart watch), or a similar device. In some implementations, the user device 110 may receive information from and / or transmit information to the platform 120.

[0041] The platform 120 includes one or more devices as described elsewhere herein. In some implementations, the platform 120 may include a cloud server or a group of cloud servers. In some implementations, the platform 120 may be designed to be modular such that software components may be swapped in or out depending on a particular need. As such, the platform 120 may be easily and / or quickly reconfigured for different uses.

[0042] In some implementations, as shown in FIG. 1, the platform 120 may be hosted in a cloud computing environment 122. Notably, while implementations described herein describe the platform 120 as being hosted in the cloud computing environment 122, in some implementations, the platform 120 may not be cloud-based (i.e., may be implemented outside of a cloud computing environment) or may be partially cloud-based.

[0043] The cloud computing environment 122 includes an environment that hosts the platform 120. The cloud computing environment 122 may provide computation, software, data access, storage, etc. services that do not require end-user (e.g. the user device 110) knowledge of a physical location and configuration of system(s) and / or device(s) that hosts the platform 120. As shown, the cloud computing environment 122 may include a group of computing resources 124 (referred to collectively as “computing resources 124” and individually as “computing resource 124”).

[0044] The computing resource 124 includes one or more personal computers, workstation computers, server devices, or other types of computation and / or communicationdevices. In some implementations, the computing resource 124 may host the platform 120. The cloud resources may include compute instances executing in the computing resource 124, storage devices provided in the computing resource 124, data transfer devices provided by the computing resource 124, etc. In some implementations, the computing resource 124 may communicate with other computing resources 124 via wired connections, wireless connections, or a combination of wired and wireless connections.

[0045] As further shown in FIG. 1, the computing resource 124 includes a group of cloud resources, such as one or more applications (“APPs”) 124-1, one or more virtual machines (“VMs”) 124-2, virtualized storage (“VSs”) 124-3, one or more hypervisors (“HYPs”) 124-4, or the like.

[0046] The application 124-1 includes one or more software applications that may be provided to or accessed by the user device 110 and / or the platform 120. The application 124-1 may eliminate a need to install and execute the software applications on the user device 110. For example, the application 124-1 may include software associated with the platform 120 and / or any other software capable of being provided via the cloud computing environment 122. In some implementations, one application 124-1 may send / receive information to / from one or more other applications 124-1, via the virtual machine 124-2.

[0047] The virtual machine 124-2 includes a software implementation of a machine (e.g. a computer) that executes programs like a physical machine. The virtual machine 124-2 may be either a system virtual machine or a process virtual machine, depending upon use and degree of correspondence to any real machine by the virtual machine 124-2. A system virtual machine may provide a complete system platform that supports execution of a complete operating system (“OS”). A process virtual machine may execute a single program, and may support a single process. In some implementations, the virtual machine 124-2 may execute on behalf of a user (e.g. the user device 110), and may manage infrastructure of the cloud computing environment 122, such as data management, synchronization, or long- duration data transfers.

[0048] The virtualized storage 124-3 includes one or more storage systems and / or one or more devices that use virtualization techniques within the storage systems or devices of the computing resource 124. In some implementations, within the context of a storage system, types of virtualizations may include block virtualization and file virtualization. Block virtualization may refer to abstraction (or separation) of logical storage from physical storage so that the storage system may be accessed without regard to physical storage orheterogeneous structure. The separation may permit administrators of the storage system flexibility in how the administrators manage storage for end users. File virtualization may eliminate dependencies between data accessed at a file level and a location where files are physically stored. This may enable optimization of storage use, server consolidation, and / or performance of non-disruptive file migrations.

[0049] The hypervisor 124-4 may provide hardware virtualization techniques that allow multiple operating systems (e.g. “guest operating systems”) to execute concurrently on a host computer, such as the computing resource 124. The hypervisor 124-4 may present a virtual operating platform to the guest operating systems, and may manage the execution of the guest operating systems. Multiple instances of a variety of operating systems may share virtualized hardware resources.

[0050] The network 130 includes one or more wired and / or wireless networks. For example, the network 130 may include a cellular network (e.g. a fifth generation (5G) network, a long-term evolution (LTE) network, a third generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g. the Public Switched Telephone Network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber optic-based network, or the like, and / or a combination of these or other types of networks.

[0051] The number and arrangement of devices and networks shown in FIG. 1 are provided as an example. In practice, there may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or differently arranged devices and / or networks than those shown in FIG. 1. Furthermore, two or more devices shown in FIG. 1 may be implemented within a single device, or a single device shown in FIG. 1 may be implemented as multiple, distributed devices. Additionally, or alternatively, a set of devices (e.g. one or more devices) of the environment 100 may perform one or more functions described as being performed by another set of devices of the environment 100.

[0052] The techniques and implementations described below can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, FIG. 2 shows a computer system (200) suitable for implementing certain embodiments of the disclosed subject matter.

[0053] The computer software can be coded using any suitable machine code or computer language, that may be subject to assembly, compilation, linking, or like mechanisms to create code comprising instructions that can be executed directly, or throughinterpretation, micro-code execution, and the like, by one or more computer central processing units (CPUs), Graphics Processing Units (GPUs), and the like.

[0054] The instructions can be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, internet of things devices, and the like.

[0055] The components shown in FIG. 2 for computer system (200) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of the present disclosure. Neither should the configuration of components be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary embodiment of a computer system (200).

[0056] Computer system (200) may include certain human interface input devices.Such a human interface input device may be responsive to input by one or more human users through, for example, tactile input (such as: keystrokes, swipes, data glove movements), audio input (such as: voice, clapping), visual input (such as: gestures), olfactory input (not depicted). The human interface devices can also be used to capture certain media not necessarily directly related to conscious input by a human, such as audio (such as: speech, music, ambient sound), images (such as: scanned images, photographic images obtain from a still image camera), video (such as two-dimensional video, three-dimensional video including stereoscopic video).

[0057] Input human interface devices may include one or more of (only one of each depicted): keyboard (201), mouse (202), trackpad (203), touch screen (210), data-glove (not shown), joystick (205), microphone (206), scanner (207), camera (208).

[0058] Computer system (200) may also include certain human interface output devices. Such human interface output devices may be stimulating the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (for example tactile feedback by the touch-screen (210), data-glove (not shown), or joystick (205), but there can also be tactile feedback devices that do not serve as input devices), audio output devices (such as: speakers (209), headphones (not depicted)), visual output devices (such as screens (210) to include CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input capability, each with or without tactile feedback capability — some of which may be capable to output two dimensional visual output or more than three dimensional outputthrough means such as stereographic output; virtual-reality glasses (not depicted), holographic displays and smoke tanks (not depicted)), and printers (not depicted).

[0059] Computer system (200) can also include human accessible storage devices and their associated media such as optical media including CD / DVD ROM / RW (220) with CD / DVD or the like media (221), thumb-drive (222), removable hard drive or solid state drive (223), legacy magnetic media such as tape and floppy disc (not depicted), specialized ROM / ASIC / PLD based devices such as security dongles (not depicted), and the like.

[0060] Those skilled in the art should also understand that term “computer readable media” as used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.

[0061] Computer system (200) can also include an interface (254) to one or more communication networks (255). Networks can for example be wireless, wireline, optical. Networks can further be local, wide-area, metropolitan, vehicular and industrial, real-time, delay-tolerant, and so on. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks to include GSM, 3G, 4G, 5G, LTE and the like, TV wireline or wireless wide area digital networks to include cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial to include CANBus, and so forth. Certain networks commonly require external network interface adapters that attached to certain general-purpose data ports or peripheral buses (249) (such as, for example USB ports of the computer system (200)); others are commonly integrated into the core of the computer system (200) by attachment to a system bus as described below (for example Ethernet interface into a PC computer system or cellular network interface into a smartphone computer system). Using any of these networks, computer system (200) can communicate with other entities. Such communication can be uni-directional, receive only (for example, broadcast TV), uni-directional send-only (for example CANbus to certain CANbus devices), or bidirectional, for example to other computer systems using local or wide area digital networks. Certain protocols and protocol stacks can be used on each of those networks and network interfaces as described above.

[0062] Aforementioned human interface devices, human-accessible storage devices, and network interfaces can be attached to a core (240) of the computer system (200).

[0063] The core (240) can include one or more Central Processing Units (CPU) (241), Graphics Processing Units (GPU) (242), specialized programmable processing units in the form of Field Programmable Gate Areas (FPGA) (243), hardware accelerators for certain tasks (244), graphics adapters (250), and so forth. These devices, along with Read-onlymemory (ROM) (245), Random-access memory (246), internal mass storage such as internal non-user accessible hard drives, SSDs, and the like (247), may be connected through a system bus (248). In some computer systems, the system bus (248) can be accessible in the form of one or more physical plugs to enable extensions by additional CPUs, GPU, and the like. The peripheral devices can be attached either directly to the core’s system bus (248), or through a peripheral bus (249). In an example, the screen (210) can be connected to the graphics adapter (250). Architectures for a peripheral bus include PCI, USB, and the like.

[0064] CPUs (241), GPUs (242), FPGAs (243), and accelerators (244) can execute certain instructions that, in combination, can make up the aforementioned computer code. That computer code can be stored in ROM (245) or RAM (246). Transitional data can also be stored in RAM (246), whereas permanent data can be stored for example, in the internal mass storage (247). Fast storage and retrieve to any of the memory devices can be enabled through the use of cache memory, that can be closely associated with one or more CPU (241), GPU (242), mass storage (247), ROM (245), RAM (246), and the like.

[0065] The computer readable media can have computer code thereon for performing various computer-implemented operations. The media and computer code can be those specially designed and constructed for the purposes of the present disclosure, or they can be of the kind well known and available to those having skill in the computer software arts.

[0066] As an example and not by way of limitation, the computer system having architecture (200), and specifically the core (240) can provide functionality as a result of processor(s) (including CPUs, GPUs, FPGA, accelerators, and the like) executing software embodied in one or more tangible, computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage as introduced above, as well as certain storage of the core (240) that are of non-transitory nature, such as core-internal mass storage (247) or ROM (245). The software implementing various embodiments of the present disclosure can be stored in such devices and executed by core (240). A computer- readable medium can include one or more memory devices or chips, according to particular needs. The software can cause the core (240) and specifically the processors therein (including CPU, GPU, FPGA, and the like) to execute particular processes or particular parts of particular processes described herein, including defining data structures stored in RAM (246) and modifying such data structures according to the processes defined by the software. In addition, or as an alternative, the computer system can provide functionality as a result of logic hardwired or otherwise embodied in a circuit (for example: accelerator (244)), which can operate in place of or together with software to execute particular processes or particularparts of particular processes described herein. Reference to software can encompass logic, and vice versa, where appropriate. Reference to a computer-readable media can encompass a circuit (such as an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware and software.

[0067] The number and arrangement of components shown in FIG. 2 are provided as an example. In practice, the device 200 may include additional components, fewer components, different components, or differently arranged components than those shown in FIG. 2. Additionally, or alternatively, a set of components (e.g. one or more components) of the device 200 may perform one or more functions described as being performed by another set of components of the device 200.

[0068] FIG. 3 is a block diagram of an example architecture 300 for performing video coding, according to some example embodiments. In some example implementations, the architecture 300 may be used as a video coding for machines (VCM) architecture, or an architecture that is otherwise compatible with or configured to perform VCM coding. For example, architecture 300 may be compatible with “Use cases and requirements for Video Coding for Machines” (ISO / IEC JTC 1 / SC 29 / WG 2 N18), “Draft of Evaluation Framework for Video Coding for Machines” (ISO / IEC JTC 1 / SC 29 / WG 2 N19), and “Call for Evidence for Video Coding for Machines” (ISO / IEC JTC 1 / SC 29 / WG 2 N20), the disclosures of which are incorporated by reference herein in their entireties.

[0069] In some example implementations, one or more of the elements illustrated in FIG. 3 may correspond to, or be implemented by, one or more of the elements discussed above with respect to FIGs. 1-2, for example, one ore more of the user device 110, the platform 120, the device 200, or any of the elements included therein.

[0070] As can be seen in FIG. 3, the architecture 300 may include a VCM encoder 310 and a VCM decoder 320. In some example embodiments, the VCM encoder may receive sensor input 301, which may include for example one or more input images, or an input video. The sensor input 301 may be provided to a feature extraction module 311 which may extract features from the sensor input, and the extracted features may be converted using feature a conversion module 312, and encoded using a feature encoding module 313. In In some example implementations, the term “encoding” may include, may correspond to, or may be used interchangeably with, the term “compressing”. The architecture 300 may include an interface 302, which may allow the feature extraction module 311 to interface with a neural network (NN) which may assist in performing the feature extraction.

[0071] The sensor input 301 may be provided to a video encoding module 314, which may generate an encoded video. In some example embodiments, after the features are extracted, converted, and encoded, the encoded features may be provided to the video encoding module 314, which may use the encoded features to assist in generating the encoded video. In In some example implementations, the video encoding module 314 may output the encoded video as an encoded video bitstream, and the feature encoding module 313 may output the encoded features as an encoded feature bitstream. In some example implementations, the VCM encoder 310 may provide both the encoded video bitstream and the encoded feature bitstream to a bitstream multiplexer 315, which may generate an encoded bitstream by combining the encoded video bitstream and the encoded feature bitstream.

[0072] In some example implementations, the encoded bitstream may be received by a bitstream demultiplexer (demux), which may separate the encoded bitstream into the encoded video bitstream and the encoded feature bitstream, which may be provided to the VCM decoder 320. The encoded feature bitstream may be provided to the feature decoding module 322, which may generate decoded features, and the encoded video bitstream may be provided to the video decoding module, which may generate a decoded video. In some example implementations, the decoded features may also be provided to the video decoding module 323, which may use the decoded features to assist in generating the decoded video.

[0073] In some example implementations, the output of the video decoding module 323 and the feature decoding module 322 may be used mainly for machine consumption, for example, for use by a machine vision module 332. In some example implementations, the output can also be used for human consumption, as illustrated in FIG. 3 as a human vision module 331. A VCM system, for example, according to the architecture 300, from the client end, for example, from the side of the VCM decoder 320, may perform video decoding to obtain the video in the sample domain first. Then one or more machine tasks to understand the video content may be performed, for example, by the machine vision module 332. In some example implementations, the architecture 300 may include an interface 303, which may allow the machine vision module 332 to interface with an NN which may assist in performing the one or more machine tasks.

[0074] As can be seen in FIG. 3 , in addition to a video encoding and decoding path, which includes the video encoding module 314 and the video decoding module 323, another path included in the architecture 300 may be a feature extraction, feature encoding, and feature decoding path, which includes the feature extraction module 311, the featureconversion module 312, the feature encoding module 313, and the feature decoding module 322.

[0075] Example embodiments below may relate to methods for enhancing decoded video for machine vision, human vision, or human / machine hybrid vision. In some example implementations, each decoded image, which may be generated, for example, by the VCM decoder 320, may be enhanced for machine vision or human vision using an enhancement module and metadata sent from the encoder side. In some example implementations, these methods can be applied to any VCM codec. Although the various example embodiments disclosed herein may be described using broader terms such as “image / video,” or using more specific terms such as “image” and “video”, it may be understood that such embodiments may be applied to other types of data.

[0076] In some example implementations, during and after reconstruction of the input video, various encoding / decoding tools may be utilized to refine or repurpose the video. These tools may be referred to as encoding / decoding modules. From a decoding standpoint, for example, various decoding modules may be executed by a decoder after reconstruction of the image / video to refine or repurpose the image / video. For example, reconstructed image / video may be processed by various decoding module for machine vision in VCM. These tools may be selectively invoked and executed by the decoder. Correspondingly, these tools may be used in the encoder in its encoding process and decoding loop. These modules, in the context of VCM decoder, is shown in FIG. 4. Example decoding modules are shown as Region of Interest (Rol) module 410, temporal resampling module 420 (e.g., temporal upsampling module, temporal interpolation module, temporal extrapolation module, etc.), spatial resampling module 430 (e.g., spatial upsampling module), post filter module 440, bit depth restoration module 450, format adapter module 460, and the like. The post filtering module 440, for example, may include a deblocking filter for removing or reducing artifacts at block boundaries of the reconstructed video. A post filter or deblocking filter may be alternatively referred to as in-loop filters, in that they are also part of an in-loop process after reconstruction in an encoder so that a filtered version of the reconstructed samples of a previous coding unit are used for coding next coding unit. These various modules above may be invoked in certain order during or post reconstruction of the encoded video.

[0077] In some example implementations, a neural network-based post filter or neural network based in-loop filter (NNLF) may be used for in-loop deblocking / post filtering. Such NNLF may be used by itself. Alternatively, such NNLF may be used in conjunction with or in addition to other deblocking / post filters to process the reconstructed samples. An exampleis shown in FIG. 5, where both an NNLF 504 and a non-NN deblocking filters are applied to reconstructed samples of the video and the filtered results are combined (e.g., averaged, weight averaged, or in other manners).

[0078] An NNLF, for example, may include a neural network with model parameters (e.g., weights, bias, and other model parameters). The NNLF may be characterized by a set of hyper parameters that define the neural network architecture (neural network layer structures and connectivity between the neural network layers). The NNLF may be pretrained. There may be multiple pretrained NNLF candidates for an encoder and a decoder to use. The choice of which NNLF from the candidate NNLFs to use, if the encoder determines to use an NNLF, may be made by the encoder at various levels (e.g., sequence, subsequence, picture, slice, frame levels), and the choice may be provided to the decoder in the encoded bitstream as an index of chosen NNLF among the candidates NNLFs. The NNLF models, may be stored in or pre-fetched to the encoder and decoder. Alternatively, they may be stored in a repository (local or remote) for the encoder and the decoder to fetch when needed. In some other implementations, the trained model parameters may be provided from the encoder to the decoder, for example, as part of supplemental enhanced information (SEI) of the bitstream. In some example implementations, in the case that the NNLF is to be fetched from or be run at a remote location, a link to the NNLF may be included in the bitstream by the encoder.

[0079] An NNLF may also be associated with filter parameters in addition to the neural network model parameters discussed above. As discussed above, these additional parameters may be predefined or may otherwise been made know to both the encoder and the decoder. Alternatively, such filter parameters may be transmitted from the encoder to the decoder in the bitstream in syntax element(s) or in the SEI of the bitstream.

[0080] In some example implementations, the usage of the NNLF to refine reconstructed samples may always be enabled. An NNLF may be applied to every block of every frame of the reconstructed video to preform post filtering in the in-loop processes in the encoder. The decoder, after reconstructing each block, would also apply the NNLF to refine the reconstructed samples. If the encoder has made a choice of an applicable NNLF (e.g., for improving filtering quality) among a plurality of NNLFs at a particular coding unit level (e.g., sequence, subsequence, picture, slice, frame, and the like), the NNLF chosen for the particular coding unit may be signaled in the bit stream. Further, the filtering parameters, if unknow to the decoder, may also be signaled in the bitstream.

[0081] In some other example implementations, whether NNLF is to be used may be expressly indicated by one or more syntax elements as included by the encoder in the bitstream. The decoder, correspondingly, may reconstruct the encoded sample, and apply the NNLF appropriately according to the one or more syntax elements in the bitstream. The filter parameters of the NNLF needed to in order to apply the post filtering of the reconstructed samples and as determined by the encoder, may also be signaled in the bitstream. The signaling of whether the NNLF usage is enabled may be provided at any signaling levels, including but not limited to the sequence, subsequence, picture, slice, frame levels, and the like. As such, the signaling of whether to enable NNLF may be included in on or more of Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Slice Header (SH), Picture Header (PH), or SEI, or the like.

[0082] In some other example implementations, the usage of NNLF at various levels may be implicit in the bitstream. For example, whether the reconstructed samples of a particular unit at a particular coding level are to be processed by an NNLF may be controlled by output(s) of one or more other VCM codec related processing tools. Such VCM codec related processing tools, for example, my include but are not limited to modules for Region of Interest (Rol) processing, bit depth truncation processing, spatial sampling processing, temporal sampling processing, and the like).

[0083] Merely as an example, a VCM codec related processing tool that may be relied on for determining the usage of NNLF may include an Rol processing module. In one particular example, whether the NNLF usage may be enabled may be determined, implicitly and without any additional express signaling, on a sequence (or sub-sequence) level. Also as a mere example, one or more output and an Rol processing module may be used to determine whether NNLF should be applied.

[0084] For example, the decoder may determine to use or enable an NNLF in a sequence or subsequence of video frames when an Rol space of each frame within the sequence (or sub sequence) is more than some predefined or signaled Rol space threshold. The decoder may determine to not use or disable the NNLF in an entirety of a sequence or subsequence of video frames when the Rol space of any frame within the sequence (or sub sequence) is less than the predefined or signaled Rol space threshold. If the decoder determines that the NNLF is enabled for a particular coding unit, it would then go to the encoded bitstream to further extract one or more filtering parameters for performing the NNLF. The encoder side may use a similar decision process in terms of calculating the Rolspace and determining whether to use the NNLF based on the calculated Rol space in its inloop filtering process.

[0085] For another example, the decoder may determine to use or enable an NNLF in a sequence or subsequence of video frames when an Rol space of a majority of frames within the sequence (or sub sequence) is more than some predefined or signaled Rol space threshold. The term “majority” means more than half, or half or more than half. Likewise, the decoder may determine not to use or disable the NNLF usage in an entirety of a sequence or subsequence of video frames in post filtering of reconstructed samples when the Rol space of a majority of frames within the sequence (or sub sequence) is less than the predefined or signaled Rol space. If the decoder determines that the NNLF is enabled for a particular coding unit, it would then go to the encoded bitstream to further extract one or more filtering parameters for performing the NNLF. The encoder side may use a similar decision process in terms of calculating the Rol space and determine whether to use the NNLF based on the calculated Rol space in its in-loop process.

[0086] For yet another example, the decoder may determine to use or enable an NNLF in a sequence or subsequence of video frames when an Rol space of some or any of the frames within the sequence (or sub sequence) is more than some predefined or signaled Rol space threshold. Likewise, the decoder may determine not to use or disable the NNLF in the sequence of subsequence of the video frames when the Rol space of none of the frames within the sequence (or subsequence) is greater than the predefined or signaled Rol space threshold If the decoder determines that the NNLF is enabled for a particular coding unit, it would then go to the encoded bitstream to further extract one or more filtering parameters for performing the NNLF. The encoder side may use a similar decision process in terms of calculating the Rol space and determining whether to use the NNLF based on the calculated Rol space in its in-loop process.

[0087] In some other example implementations, whether to apply an NNLF may be determined on a sequence or subsequence level by the encoder and such decision may be explicitly signaled in the bitstream as, for example, a control flag in a Sequence Parameter Set (SPS). For example, the flag may be set at “1” for enabling the NNLF and “0” for disabling the NNLF. Also as a mere example, one or more output and an Rol process may be used to determine whether NNLF should be applied.

[0088] For example, the encoder may determine to use or enable an NNLF in a sequence or subsequence of video frames when an Rol space of each frame within the sequence (or sub sequence) is more than some predefined or signaled Rol space threshold.The encoder may determine to not use or disable the NNLF usage in an entirety of a sequence or subsequence of video frames in its in-loop process when the Rol space of any frame within the sequence (or sub sequence) is less than the predefined or signaled Rol space. The encoder may explicitly signal the decision at the sequence level in the bitstream. The decoder may determine whether to apply the NNLF according the explicit signaling. If the decoder determines that the NNLF is enabled for a particular coding unit from the signaling in the bitstream, it would then go to the encoded bitstream to further extract one or more filtering parameters for performing the NNLF.

[0089] For another example, the encoder may determine to use or enable an NNLF in a sequence or subsequence of video frames when an Rol space of a majority of frames within the sequence (or sub sequence) is more than some predefined or signaled Rol space threshold. The term “majority” means more than half, or half or more than half. Likewise, the encoder may determine not to use or disable the NNLF usage in an entirety of a sequence or subsequence of video frames in post filtering of reconstructed samples when the Rol space of a majority of frames within the sequence (or sub sequence) is less than the predefined or configured Rol space. The encoder may explicitly signal the decision at the sequence level in the bitstream. The decoder may determine whether to apply the NNLF according the explicit signaling. If the decoder determines that the NNLF is enabled for a particular coding unit from the signaling in the bitstream, it would then go to the encoded bitstream to further extract one or more filtering parameters for performing the NNLF.

[0090] For yet another example, the encoder may determine to use or enable an NNLF in a sequence or subsequence of video frames when an Rol space some or any of the frames within the sequence (or sub sequence) is more than some predefined or signaled Rol space threshold. Likewise, the encoder may determine not to use or disable the NNLF in the sequence of subsequence of the video frames when the Rol space of none of the frames within the sequence (or subsequence) is greater than the predefined or signaled Rol space threshold. The encoder may explicitly signal the decision at the sequence level in the bitstream. If the decoder determines that the NNLF is enabled for a particular coding unit from the signaling in the bitstream, it would then go to the encoded bitstream to further extract one or more filtering parameters for performing the NNLF.

[0091] In some other example implementations, whether to apply an NNLF may be determined on a frame or slice level by the decoder according to, for example, output of the Rol processing without any explicit signaling from the bitstream.

[0092] For example, the decoder may determine to use or enable an NNLF in a frame or slice of a video when an Rol space of the frame or the slice is more than some predefined or signaled Rol space threshold. The decoder may determine to not use or disable the NNLF usage in the frame or slice when the Rol space of the frame or slide is less than the predefined or signaled Rol space threshold. If the decoder determines that the NNLF is enabled for the frame or slice, it would then go to the encoded bitstream to further extract one or more filtering parameters for performing the NNLF. The encoder side may use a similar decision process in terms of calculating the Rol space and determine whether to use the NNLF based on the calculated Rol space in its in-loop process.

[0093] In some other example implementations, whether to apply an NNLF may be determined for a frame or slice by the encoder and such decision may be explicitly signaled in the bitstream as, for example, a control flag in a Picture Header (PH) or Slice Header (SH). For example, the flag may be set at “1” for enabling the NNLF and “0” for disabling the NNLF. Also as a mere example, one or more output and an Rol process may be used to determine whether NNLF should be applied.

[0094] For example, the encoder may determine to use or enable an NNLF in a frame or slice of a video when an Rol space of the frame or slice is more than some predefined or signaled Rol space threshold. The encoder may determine to not use or disable the NNLF usage in the frame or slice in its in-loop process when the Rol space of the frame or slide is less than the predefined or signaled Rol space. The encoder may explicitly signal the decision at the frame or slice level in the bitstream. The decoder may determine whether to apply the NNLF according the explicit signaling. If the decoder determines that the NNLF is enabled for a particular frame or slice from the signaling in the bitstream, it would then go to the encoded bitstream to further extract one or more filtering parameters for performing the NNLF.

[0095] FIG. 6 shows a flow chart for an example process (600) according to an embodiment of the disclosure. The process (600) starts at step (S601). In Step (S610), a bitstream of a video is received. In Step (S620), the bitstream is decoded to obtain a reconstructed video. In Step (S630), when it is determined that a neural network in-loop filter (NNLF) is to be applied to a portion of the reconstructed video, the NNLF is identified and a set of filter parameters associated with the NNLF are extracted from the reconstructed video. In Step (S640), the NNLF is applied with the set of filter parameters to the portion of the reconstructed video to generate a filtered video portion. The procedure (600) stops at (S699).

[0096] FIG. 7 shows a flow chart for an example process (700) according to an embodiment of the disclosure. The process (700) starts at step (S701). In Step (S710), a portion of the video is encoded as part of an encoded bitstream. In Step (S720), the encoded portion of the video in an in-loop process to generate a reconstructed portion of the video is reconstructed. In Step (S730), when it is determined that an NNLF is to be applied to the reconstructed portion of the video, the NNLF is identified and a set of filter parameters associated with the NNLF is determined. In Step (S740), the NNLF is applied with the set of filter parameters to the reconstructed portion of the video to generate a filtered video portion in the in-loop process. The procedure (700) stops at (S799).

[0097] The processes (600), and (700) can be suitably adapted. Step(s) in the processes (600) and (700) can be modified and / or omitted. Additional step(s) can be added. Any suitable order of implementation can be used.

[0098] The techniques disclosed in the present disclosure may be used separately or combined in any order. Further, each of the techniques (e.g., methods, embodiments), encoder, and decoder may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In some examples, the one or more processors execute a program that is stored in a non-transitory computer-readable medium.

[0099] While this disclosure has described several exemplary embodiments, there are alterations, permutations, and various substitute equivalents, which fall within the scope of the disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods which, although not explicitly shown or described herein, embody the principles of the disclosure and are thus within the spirit and scope thereof.

Claims

WHAT IS CLAIMED IS:

1. A method for decoding a video, comprising: receiving a bitstream of the video; decoding the bitstream to obtain a reconstructed video; when it is determined that a neural network in-loop filter (NNLF) is to be applied to a portion of the reconstructed video, identifying the NNLF and extracting from the reconstructed video a set of filter parameters associated with the NNLF; and applying the NNLF with the set of filter parameters to the portion of the reconstructed video to generate a filtered video portion.

2. The method of claim 1, wherein the NNLF is determined always applied to an entirety of the video.

3. The method of clam 1, further comprising extracting a signaling information item from the bitstream which indicates whether the NNLF is to be applied to the portion of the reconstructed video.

4. The method of claim 3, wherein the signaling information item is included in a Sequence Parameter Set (SPS), a Picture Parameter Set (PPS), a Sequence Header (SH), a Picture Header (PH), or Supplemental Enhancement Information (SEI) in the bitstream.

5. The method of claim 1, wherein whether the NNLF is to be applied to the portion of the reconstructed video is controlled by an output of at least one VCM (Video Coding for Machine) processing module applied to the bitstream.

6. The method of claim 5, wherein the at least one VCM processing module comprises one or more of a Region of Interest (Rol) processing module, a bit depth truncation module, a spatial sampling module, or temporal sampling module.

7. The method of claim 5, wherein whether the NNLF is to be applied to the portion of the reconstructed video is determined at a sequence or sub-sequence level.

8. The method of claim 7, wherein the at least one VCM processing module comprises an Rol processing module and the NNLF is or is not applied to a sequence or subsequence when a total Rol space of each frame within the sequence or subsequence is greater or less than a predetermined Rol space threshold, respectively.

9. The method of claim 7, wherein the at least one VCM processing module comprises an Rol processing module and the NNLF is or is not applied to a sequence or subsequence when Rol space of majority of frames within the sequence or subsequence is greater or less than a predetermined Rol space threshold, respectively.

10. The method of claim 7, wherein the at least one VCM processing module comprises an Rol processing module and the NNLF is or is not applied to a sequence or subsequence when Rol space of at least one of frames within the sequence or subsequence is greater or smaller than a predetermined Rol space threshold, respectively.

11. The method of claim 5, wherein whether the NNLF is to be applied to the portion of the reconstructed video is determined at a frame or slice level.

12. The method of claim 11, wherein the at least one VCM processing module comprises an Rol processing module and the NNLF is or is not applied to a frame or slice when Rol space of the frame or slice is greater or smaller than a predetermined Rol space threshold, respectively.

13. A method for processing a video, comprising: encoding a portion of the video as part of an encoded bitstream; reconstructing the encoded portion of the video in an in-loop process to generate a reconstructed portion of the video; when it is determined that an NNLF is to be applied to the reconstructed portion of the video, identifying the NNLF and determining a set of filter parameters associated with the NNLF; and applying the NNLF with the set of filter parameters to the reconstructed portion of the video to generate a filtered video portion in the in-loop process.

14. The method of claim 13, wherein the NNLF is determined as always applied to an entirety of the video.

15. The method of claim 13, further comprising including in the encoded bitstream a signaling information item for indicating whether the NNLF is applied in the in-loop process.

16. The method of claim 15, wherein the signaling information item is included in a Sequence Parameter Set (SPS), a Picture Parameter Set (PPS), a Sequence Header (SH), a Picture Header (PH), or Supplemental Enhancement Information (SEI) in the bitstream.

17. The method of claim 15, wherein whether the NNLF is applied to the reconstructed portion of the video is controlled by an output of at least one VCM processing module applied to the reconstructed portion of the video in the in-loop process.

18. The method of claim 17, wherein: whether the NNLF is applied to the reconstructed portion of the video is determined at a sequence or sub-sequence level; and the method further comprising determining to apply the NNLF and including a corresponding signaling indicator in the encoded bitstream for the sequence or subsequence when: a total Rol space of each frame within the sequence or subsequence is greater than a predetermined Rol space threshold; when Rol space of majority of frames within the sequence or subsequence is greater than a predetermined Rol space threshold; or when Rol space of at least one of frames within the sequence or subsequence is greater than a predetermined Rol space threshold.

19. The method of claim 17, wherein: whether the NNLF is applied to the reconstructed portion of the video is determined at a level of a frame or slice; and the method further comprising determining to apply the NNLF and including a corresponding signaling indicator in the encoded bitstream for the frame or slice when a total Rol space of the frame or slice is greater than a predetermined Rol space threshold.

20. A non-transitory computer readable storage medium for storing a bitstream of a video, the bitstream comprising: an encoded portion of the video; an indicator for signaling to a decoder whether an NNLF is to be applied to a reconstruction of the encoded portion of the video; and when the indicator signals that the NNLF is to be applied, a set of filter parameters associated with the NNLF.

Citation Information

Patent Citations

  • High-level syntax for signaling neural networks within a media bitstream

    US20220256227A1

  • Systems and methods for signaling neural network-based in-loop filter parameter information in video coding

    US20220321919A1