Video decoding method, video encoding method, and storage medium

By adaptively adjusting the order of encoding and decoding modules, and flexibly selecting the execution order of modules according to the video content characteristics, the problems of efficiency and performance limitations in the prior art are solved, and more efficient video encoding and decoding effects are achieved.

CN120475159APending Publication Date: 2025-08-12TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510150194.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-01-23
Filing Date
2025-02-11
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

When the existing video encoding and decoding technology is oriented towards machine vision tasks, it is difficult to adaptively adjust the order of encoding and decoding modules according to the characteristics of the video content, resulting in limited efficiency and performance.

Method used

By adaptively determining the order of the encoding and decoding modules, using the content characteristics and additional information of the video, the execution order of multiple decoding modules is flexibly selected, including the ROI module, the time upsampling module, the spatial upsampling module, the post-filtering module and the bit-depth recovery module, etc.

Benefits of technology

Improves the efficiency and performance of video encoding and decoding, especially in machine vision tasks, providing higher encoding gain and less encoding loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120475159A_ABST
    Figure CN120475159A_ABST
Patent Text Reader

Abstract

The invention relates to a video decoding method, a video encoding method, and a storage medium. The encoding / decoding order may be adaptively determined by content characteristics of the video and / or additional information / parameters associated with various encoding / decoding modules.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is based upon and claims the benefit of U.S. Provisional Patent Application No. 63 / 552,606, filed on February 12, 2024, entitled “Flexibility of module order in video coding systems,” and U.S. Patent Application No. 19 / 034,911, filed on January 23, 2025, the contents of which are incorporated herein by reference in their entirety. Technical Field

[0002] The present application mainly relates to embodiments of video encoding and decoding, and specifically to adaptive sorting of encoding modules and decoding modules. Background Art

[0003] Videos or images may be used by human users for various purposes, such as entertainment, education, etc. Therefore, video codecs or image codecs may generally utilize the characteristics of the human visual system to achieve better compression efficiency while maintaining good subjective quality.

[0004] With the rise of machine learning applications and the abundance of sensors, many intelligent platforms have begun using video for machine vision tasks such as object detection, segmentation, and tracking. Therefore, encoding videos or images for machine vision tasks has become an interesting and challenging problem. This has led to the research on video coding for machines (VCM).

[0005] Although the embodiments are described in the context of VCM, the underlying principles are generally applicable to other video codec systems. Summary of the Invention

[0006] The present application generally relates to embodiments of video coding and decoding, and more particularly to adaptive ordering of encoding and decoding modules. The encoding / decoding order can be adaptively determined by content characteristics of the video and / or additional information / parameters associated with various encoding / decoding modules.

[0007] In some exemplary embodiments, a video decoding method is disclosed. The method may include: receiving an encoded bitstream of a video; determining a sequential order for executing multiple decoding modules based on an indication extracted from the encoded bitstream; and decoding the encoded bitstream by executing the multiple decoding modules according to the sequential order. The multiple decoding modules include at least one of the following: a region of interest (ROI) module, a temporal upsampling module, a spatial upsampling module, a post-filtering module, a bit depth recovery module, or a format adapter module.

[0008] In the above exemplary embodiment, the sequential order is selected from a set of predefined sequential orders based on the indication.

[0009] In any of the above exemplary embodiments, if the plurality of decoding modules are executed, the set of redefined sequential orders of the plurality of decoding modules includes at least one of the following: a RoI module, followed by a spatial upsampling module, followed by a temporal upsampling module, followed by a post-filter module, followed by a bit depth recovery module, and followed by a format adapter module; a bit depth recovery module, followed by the RoI module, followed by a spatial upsampling module, followed by a temporal upsampling module, followed by a post-filter module, and followed by a format adapter module; and a RoI module, followed by a temporal upsampling module, followed by a spatial upsampling module, followed by a post-filter module, followed by a bit depth recovery module, and followed by a format adapter module.

[0010] In any of the above exemplary embodiments, the sequential order is adaptively determined.

[0011] In any of the above exemplary embodiments, the indication is derived from at least one content characteristic of the video, wherein the at least one content characteristic is extracted from the encoded bitstream; and the at least one content characteristic includes at least one of the following: time dependency, content complexity, color complexity, motion complexity, texture characteristics, dynamic range change, foreground-background segmentation, ratio of high-frequency areas, and ratio of RoI areas.

[0012] In any of the above exemplary embodiments, the indication is based on inter-frame content movement, and a higher degree of inter-frame content movement indicates an earlier execution of the temporal upsampling module.

[0013] In any of the above exemplary embodiments, the inter-frame content motion is determined by: the number of pixels having an inter-frame change; the percentage of pixels having an inter-frame value change; and / or the average inter-frame pixel value change.

[0014] In any of the above exemplary embodiments, the indication is derived based on content complexity, and a higher level of content complexity indicates an earlier execution of the spatial upsampling module.

[0015] In any of the above exemplary embodiments, the indication is derived based on the ratio of the high frequency region, and a higher level of the ratio of the high frequency region indicates an earlier execution of the spatial upsampling module.

[0016] In any of the above exemplary embodiments, the indication is derived based on the RoI region, and a higher level of the ratio of the RoI region indicates an earlier execution of the ROI module.

[0017] In any of the above exemplary embodiments, the sequential order is selected from a plurality of sequential orders based on a set of threshold levels of at least one content characteristic.

[0018] In any of the above exemplary embodiments, the indication is obtained by processing a set of video characteristics derived from the encoded bitstream using a pre-trained neural network.

[0019] In any of the above exemplary embodiments, the indication is derived from a set of additional parameters, wherein the set of additional parameters is extracted from an encoded code stream associated with at least one decoding module of the plurality of decoding modules.

[0020] In any of the above exemplary embodiments, the set of additional parameters includes at least one of a temporal resampling rate or a spatial resampling rate.

[0021] In any of the above exemplary embodiments, the method further includes: determining the successive order is adaptively determined based on a flag in the encoded code stream; and determining the successive order based on an index represented by a signal, wherein the index is an index of the successive order in multiple predefined successive orders.

[0022] In any of the above exemplary embodiments, the sequential order is asymmetric to the encoding order of the plurality of encoding modules used to generate the encoded code stream, and the plurality of encoding modules correspond to the plurality of decoding modules.

[0023] In some other exemplary embodiments, a video encoding method is disclosed. The method may include: adaptively determining a sequential order for executing multiple encoding modules; encoding the video by executing the multiple encoding modules in the sequential order to generate an encoded bitstream; and adding an indication to the encoded bitstream to indicate the sequential order to a decoder. The multiple encoding modules include at least one of the following: a region of interest (ROI) module, a temporal upsampling module, a spatial upsampling module, a post-filtering module, a bit depth recovery module, or a format adapter module.

[0024] In the above exemplary embodiment, the sequential order is determined based on at least one content characteristic of the video; and the at least one content characteristic includes at least one of the following: time dependency, content complexity, color complexity, motion complexity, texture characteristics, dynamic range variation, foreground-background segmentation, ratio of high-frequency areas, and ratio of RoI areas.

[0025] In any of the above exemplary embodiments, the indication includes: a flag for indicating that the sequential order is adaptively determined; and an index of the sequential order in a plurality of predefined sequential orders.

[0026] In some other exemplary embodiments, a non-transitory computer-readable storage medium for storing an encoded bitstream of a video is disclosed. The encoded bitstream may include: a flag for indicating whether a decoder uses an adaptive sequential order to execute multiple decoding modules; and an indication for, when the flag indicates the use of the adaptive sequential order, selecting the adaptive sequential order from among a plurality of predefined sequential orders for executing the multiple decoding modules.

[0027] Aspects of the present application further provide an electronic device or apparatus used as an encoder or a decoder, which includes a circuit configured to execute any of the above method embodiments.

[0028] Aspects of the present application also provide a non-volatile computer-readable medium for storing computer instructions, which, when executed by a computer, causes the computer to perform any one of the above method embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Further features, properties and various advantages of the subject matter of the present application will become more apparent from the following detailed description and accompanying drawings, in which:

[0030] Figure 1 is a schematic diagram of an example application environment of various embodiments, in which the methods, devices, and systems described herein can be implemented.

[0031] Figure 2 is a schematic diagram of an example computer system for various embodiments.

[0032] Figure 3 is a block diagram of an example architecture for performing video encoding and decoding according to various embodiments.

[0033] Figure 4 Various decoding modules that may be utilized in a VCM decoder are shown.

[0034] Figure 5 is a flow chart of an example method for utilizing multiple decoding modules.

[0035] Figure 6 is a flow chart of an example method for encoding video. DETAILED DESCRIPTION

[0036] Throughout the specification and claims, terms may have subtle meanings that are suggested or implied by the context, not just the explicitly stated meanings. The phrases "in one embodiment" or "in some embodiments" as used herein do not necessarily refer to the same embodiment, and the phrases "in another embodiment" or "in other embodiments" as used herein do not necessarily refer to different embodiments. Similarly, the phrases "in one implementation" or "in some implementations" as used herein do not necessarily refer to the same implementation, and the phrases "in another implementation" or "in other implementations" as used herein do not necessarily refer to different implementations. For example, the claimed subject matter is intended to include, in whole or in part, a combination of exemplary embodiments / implementations.

[0037] Typically, terms can be understood at least in part from how they are used in context. For example, terms such as "and," "or," or "and / or," as used herein, can include multiple meanings that can depend, at least in part, on the context in which the terms are used. Typically, "or," if used in an associative list such as A, B, or C, is intended to mean A, B, and C, used herein in an inclusive sense, and A, B, or C, used herein in an exclusive sense. Additionally, the terms "one or more" or "at least one," as used herein, can be used to describe any feature, structure, or characteristic in the singular, or can be used to describe a combination of features, structures, or characteristics in the plural, depending, at least in part, on the context. Similarly, terms such as "a," "an," or "the" can also be understood to convey singular usage or plural usage, depending, at least in part, on the context. Additionally, the terms "based on" or "determined by..." can be understood to not necessarily be intended to convey a set of exclusive factors, but rather can allow for the presence of additional factors that are not necessarily explicitly described, again, depending, at least in part, on the context.

[0038] Figure 1 FIG. 1 is a schematic diagram of an application environment 100 of various embodiments, in which the methods, devices, and systems described herein may be implemented. Figure 1 As shown, environment 100 may include user device 110, platform 120, and network 130. The devices of environment 100 may be interconnected via wired connections, wireless connections, or a combination of wired and wireless connections.

[0039] User device 110 includes at least one device capable of receiving, generating, storing, processing, and / or providing information related to platform 120. For example, user device 110 may include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smartphone, a wireless phone, etc.), a wearable device (e.g., smart glasses or a smart watch), or the like. In some embodiments, user device 110 can receive information from platform 120 and / or send information to platform 120.

[0040] As described elsewhere herein, platform 120 includes at least one device. In some embodiments, platform 120 may include a cloud server or a group of cloud servers. In some embodiments, platform 120 may be designed to be modular so that software components can be swapped in or out based on specific needs. This allows platform 120 to be easily and / or quickly reconfigured for different uses.

[0041] In some embodiments, as Figure 1 As shown, the platform 120 can be hosted in a cloud computing environment 122. It is worth noting that although the embodiments described herein describe the platform 120 as being hosted in the cloud computing environment 122, in some embodiments, the platform 120 is not cloud-based (i.e., can be implemented outside of the cloud computing environment) or can be partially cloud-based.

[0042] Cloud computing environment 122 includes an environment hosting platform 120. Cloud computing environment 122 can provide computing, software, data access, storage, and other services without requiring end users (e.g., user devices 110) to be aware of the physical location and configuration of the systems and / or devices hosting platform 120. As shown, cloud computing environment 122 can include a set of computing resources 124 (collectively, "computing resources 124," and each resource individually, "computing resource 124").

[0043] Computing resources 124 include at least one personal computer, workstation computer, server device, or other type of computing and / or communication device. In some embodiments, computing resources 124 may host platform 120. Cloud resources may include computing instances executed in computing resources 124, storage devices provided in computing resources 124, data transmission devices provided by computing resources 124, and the like. In some embodiments, computing resources 124 may communicate with other computing resources 124 via wired connections, wireless connections, or a combination of wired and wireless connections.

[0044] Further Figure 1 As shown, the computing resources 124 include a set of cloud resources, such as at least one application ("APP") 124-1, at least one virtual machine ("VM") 124-2, virtualized storage ("VS") 124-3, at least one hypervisor ("HYP") 124-4, etc.

[0045] Application 124-1 includes at least one software application that can be provided to or accessed by user device 110 and / or platform 120. Application 124-1 eliminates the need to install and execute software applications on user device 110. For example, application 124-1 can include software associated with platform 120 and / or any other software that can be provided through cloud computing environment 122. In some embodiments, one application 124-1 can send / receive information to / from at least one other application 124-1 via virtual machine 124-2.

[0046] The virtual machine 124-2 comprises a software implementation of a machine (e.g., a computer) that executes programs in a manner similar to a physical machine. The virtual machine 124-2 can be a system virtual machine or a process virtual machine, depending on the use and correspondence of the virtual machine 124-2 to any real machine. A system virtual machine can provide a complete system platform that supports the execution of a complete operating system ("OS"). A process virtual machine can execute a single program and can support a single process. In some embodiments, the virtual machine 124-2 can execute on behalf of a user (e.g., a user device 110) and can manage the infrastructure of the cloud computing environment 122, such as data management, synchronization, or long-term data transfer.

[0047] Virtualized storage 124-3 includes at least one storage system and / or at least one device that uses virtualization technology within the storage system or device of the computing resource 124. In some embodiments, within the context of the storage system, the types of virtualization may include block virtualization and file virtualization. Block virtualization may refer to abstracting (or separating) logical storage from physical storage so that the storage system can be accessed without considering the physical storage or heterogeneous structure. The above separation may allow administrators of the storage system to flexibly manage the storage of end users. File virtualization may eliminate the dependency between data accessed at the file level and the location of the physical storage file. This may optimize the use of storage, server consolidation, and / or the performance of non-disruptive file migration.

[0048] Hypervisor 124-4 can provide hardware virtualization technology that allows multiple operating systems (e.g., "guest operating systems") to execute simultaneously on a host computer such as computing resource 124. Hypervisor 124-4 can provide a virtual operating platform to the guest operating systems and can manage the execution of the guest operating systems. Multiple instances of various operating systems can share virtualized hardware resources.

[0049] The network 130 includes at least one wired network and / or wireless network. For example, the network 130 may include a cellular network (e.g., a fifth generation (5G) network, a long-term evolution (LTE) network, a third generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a public switched telephone network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-optic-based network, etc., and / or a combination of these or other types of networks.

[0050] Figure 1 The number and deployment of devices and networks shown are provided as examples. Figure 1 There may be more devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks deployed in different ways than those shown. Figure 1 At least two of the devices shown may be implemented in a single device, or Figure 1 The single device shown may be implemented by multiple distributed devices. Additionally or alternatively, one set of devices (eg, at least one device) of environment 100 may perform at least one function described as being performed by another set of devices of environment 100.

[0051] The techniques and implementations described below may be implemented using computer-readable instructions by computer software and physically stored in at least one computer-readable medium. Figure 2 A computer system (200) suitable for implementing certain embodiments of the presently disclosed subject matter is shown.

[0052] Computer software may be encoded using any suitable machine code or computer language that may be assembled, compiled, linked, or the like to create code comprising instructions that may be executed directly by at least one computer central processing unit (CPU), graphics processing unit (GPU), or the like, or through interpretation, microcode execution, or the like.

[0053] These instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, IoT devices, and the like.

[0054] Figure 2 The components shown for the computer system (200) are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing embodiments of the present application. The configuration of components should not be interpreted as having any dependency or requirement on any one or combination of components shown in the exemplary embodiment of the computer system (200).

[0055] The computer system (200) may include certain human-computer interface input devices. Such human-computer interface input devices may be responsive to input from at least one human user, such as tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, taps), visual input (e.g., gestures), and olfactory input (not shown). The human-computer interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0056] The human-machine interface input device may include at least one of the following (only one of each is depicted): keyboard (201), mouse (202), touchpad (203), touch screen (210), data gloves (not shown), joystick (205), microphone (206), scanner (207), camera (208).

[0057] The computer system (200) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate at least one sense of a human user through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback through a touch screen (210), a data glove (not shown), or a joystick (205), but there may also be tactile feedback devices that are not used as input devices), audio output devices (e.g., speakers (209), headphones (not shown)), visual output devices (e.g., screens (210), virtual reality glasses (not shown), holographic displays and smoke boxes (not shown), and printers (not shown). The screens (210) include CRT screens, LCD screens, plasma screens, and OLED screens, each of which may or may not have touch screen input capabilities and may or may not have tactile feedback capabilities - some of which are capable of outputting two-dimensional visual output or three or more-dimensional output in a manner such as stereo output.

[0058] The computer system (200) may also include human-accessible storage devices and their associated media, such as optical media (221) (such as CD / DVD ROM / RW (220) with CD / DVD, etc.), thumb drives (222), removable hard disks or solid-state drives (223), traditional magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), etc.

[0059] Those skilled in the art will also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other volatile signals.

[0060] The computer system (200) may also include an interface (254) to at least one communication network (255). The network may be, for example, wireless, wired, or optical. The network may also be local, wide, metropolitan, vehicular, and industrial, real-time, delay-tolerant, and the like. Examples of networks include local area networks (e.g., Ethernet, WLAN), cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), TV wired or wireless wide area digital networks (including cable TV, satellite TV, and terrestrial broadcast TV), vehicular and industrial networks (including CAN buses), and the like. Some networks typically require an external network interface adapter to attach to some common data port or peripheral bus (249) (e.g., a USB port of the computer system (200)); other networks are typically integrated into the kernel of the computer system (200) by connecting to the system bus (e.g., an Ethernet interface connected to a PC computer system or a cellular network interface connected to a smartphone computer system) as described below. Using any of these networks, the computer system (200) can communicate with other entities. This communication can be one-way receive-only (e.g., broadcast TV), one-way send-only (e.g., CAN bus to certain CAN bus devices), or two-way, such as communication with other computer systems using a local area network or wide area digital network. As mentioned above, certain protocols and protocol stacks can be used on each of these networks and network interfaces.

[0061] The above-mentioned human interface devices, human-accessible storage devices and network interfaces may be attached to the kernel (240) of the computer system (200).

[0062] The core (240) may include at least one central processing unit (CPU) (241), a graphics processing unit (GPU) (242), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (243), hardware accelerators for specific tasks (244), a graphics adapter (250), and the like. These devices, along with read-only memory (ROM) (245), random access memory (246), internal mass storage (247) (such as an internal non-user accessible hard drive, SSD, etc.) can be connected via a system bus (248). In some computer systems, the system bus (248) can be accessed in the form of at least one physical plug to allow expansion by adding CPUs, GPUs, etc. Peripheral devices can be attached to the core's system bus (248) directly or via a peripheral bus (249). In an example, the screen (210) can be connected to the graphics adapter (250). Architectures for peripheral buses include PCI, USB, and the like.

[0063] The CPU (241), GPU (242), FPGA (243), and accelerator (244) can execute certain instructions, which, when combined, can constitute the aforementioned computer code. The computer code can be stored in ROM (245) or RAM (246). Transient data can also be stored in RAM (246), while permanent data can be stored, for example, in internal mass storage (247). Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with at least one of the CPU (241), GPU (242), mass storage (247), ROM (245), RAM (246), etc.

[0064] The computer readable medium may have computer code therein for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of this application, or they may be of a type well known and available to those skilled in the art of computer software.

[0065] As an example and not for limitation, a computer system having architecture (200), and in particular, a kernel (240), can provide functionality by virtue of at least one processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software contained in at least one tangible computer-readable medium. Such computer-readable medium can be a medium associated with a user-accessible mass storage as described above, as well as certain storage of the kernel (240) having non-volatile properties, such as a mass storage (247) or ROM (245) within the kernel. Software implementing various embodiments of the present application can be stored in such a device and executed by the kernel (240). Depending on specific needs, the computer-readable medium can include at least one memory device or chip. The software can cause the kernel (240) and in particular the processor therein (including a CPU, GPU, FPGA, etc.) to perform a specific method or a specific portion of a specific method described herein, including defining data structures stored in RAM (246) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system may be provided with functionality by logic hardwired or otherwise embodied in circuitry (e.g., accelerator (244)) that may operate in place of or in conjunction with software to perform a particular method or a particular portion of a particular method described herein. Where appropriate, reference to software may encompass logic hardware, and vice versa. Where appropriate, reference to a computer-readable medium may encompass circuitry (such as an integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or both. This application encompasses any suitable combination of hardware and software.

[0066] Figure 2The number and deployment of components shown are provided as examples. Figure 2 Device 200 may have more components, fewer components, different components, or components arranged differently than those shown. Furthermore, in addition or as an alternative, one group of components (e.g., at least one device) of device 200 may perform at least one function described as being performed by another group of devices of device 200.

[0067] Figure 3 is a block diagram of an example architecture 300 for performing video coding and decoding in various embodiments. In various embodiments, the architecture 300 can be a video coding for machines (VCM) architecture, or an architecture that is otherwise compatible with or configured to perform VCM coding. For example, the architecture 300 can be compatible with "Use cases and requirements for video coding for machines" (ISO / IEC JTC 1 / SC 29 / WG 2N18), "Draft evaluation framework for video coding for machines" (ISO / IEC JTC 1 / SC 29 / WG 2N19), and "Call for evidence for video coding for machines" (ISO / IEC JTC 1 / SC 29 / WG 2N20), the disclosures of which are incorporated herein by reference in their entirety.

[0068] In each embodiment, Figure 3 At least one of the components shown in may correspond to the above description of Figures 1 to 2 At least one of the components discussed (eg, at least one of the user device 110, the platform 120, the device 200, or any components included therein) or is implemented thereby.

[0069] from Figure 3 As can be seen in FIG, the architecture 300 may include a VCM encoder 310 and a VCM decoder 320. In some example embodiments, the VCM encoder may receive a sensor input 301, which may include, for example, at least one input image or input video. The sensor input 301 may be provided to a feature extraction module 311, which may extract features from the sensor input, and the extracted features may be converted by a feature conversion module 312 and encoded by a feature encoding module 313. In various embodiments, the term "encoding" may include or correspond to the term "compression", or may be used interchangeably with the term "compression". The architecture 300 may include an interface 302, which may allow the feature extraction module 311 to interface with a neural network (NN), which may assist in performing feature extraction.

[0070] The sensor input 301 may be provided to a video encoding module 314, which may generate an encoded video. In some example embodiments, after the features are extracted, converted, and encoded, the encoded features may be provided to the video encoding module 314, which may use the encoded features to assist in generating the encoded video. In various embodiments, the video encoding module 314 may output the encoded video as an encoded video stream, and the feature encoding module 313 may output the encoded features as an encoded feature stream. In various embodiments, the VCM encoder 310 may provide both the encoded video stream and the encoded feature stream to a stream multiplexer 315, which may generate an encoded stream by combining the encoded video stream and the encoded feature stream.

[0071] In various embodiments, the encoded bitstream can be received by a bitstream demultiplexer (demux), which can separate the encoded bitstream into an encoded video bitstream and an encoded feature bitstream. The encoded video bitstream and the encoded feature bitstream can be provided to the VCM decoder 320. The encoded feature bitstream can be provided to the feature decoding module 322, which can generate decoded features; and the encoded video bitstream can be provided to the video decoding module, which can generate decoded video. In various embodiments, the decoded features can also be provided to the video decoding module 323, which can use the decoded features to assist in generating the decoded video.

[0072] In various embodiments, the outputs of the video decoding module 323 and the feature decoding module 322 may be primarily used by machines, such as the machine vision module 332. In various embodiments, the outputs may also be used by humans, in Figure 3 3. A human vision module 331 is shown in the figure. A VCM system (e.g., architecture 300) from a client (e.g., from the VCM decoder 320 side) can perform video decoding to first obtain a video in the sample domain. Then, at least one machine task of understanding the video content can be performed, for example, by a machine vision module 332. In various embodiments, architecture 300 can include an interface 303 that can allow the machine vision module 332 to interface with a neural network, which can assist in performing the at least one machine task.

[0073] from Figure 3 It can be seen that in addition to the video encoding and decoding path (including the video encoding module 314 and the video decoding module 323), another path included in the architecture 300 can be a feature extraction, feature encoding and feature decoding path, which includes a feature extraction module 311, a feature conversion module 312, a feature encoding module 313 and a feature decoding module 322.

[0074] Various embodiments may relate to methods for enhancing decoded video for machine vision, human vision, or mixed human / machine vision. In various embodiments, each decoded image, such as that generated by the VCM decoder 320, may be enhanced for machine vision or human vision using an enhancement module and metadata sent from the encoder side. In various embodiments, these methods may be applied to any VCM codec. Although some embodiments may be described using broader terms such as "image / video" or more specific terms such as "image" and "video," it is understood that these descriptions are applicable to various embodiments.

[0075] In some exemplary embodiments, various encoding / decoding tools may be utilized during and after reconstruction of the input video to correct the video or change the purpose of the video. These tools may be referred to as encoding / decoding modules. From a decoding perspective, for example, various decoding modules may be executed by the decoder after the image / video has been reconstructed to correct the image / video or change the purpose of the image / video. For example, the reconstructed image / video may be processed by various decoding modules for machine vision in a VCM. These tools may be selectively called and executed by the decoder. Accordingly, these tools may be used by the encoder in the encoding process and decoding loop of the encoder. In the context of a VCM decoder, these modules may be as follows: Figure 4 The exemplary decoding modules are illustrated as a Region of Interest (RoI) module 410, a temporal resampling module 420 (e.g., a temporal upsampling module, a temporal interpolation module, a temporal extrapolation module, etc.), a spatial resampling module 430 (e.g., a spatial upsampling module), a post-filter module 440, a bit depth recovery module 450, a format adapter module 460, etc. These modules may be called in a specific order during or after reconstruction of the encoded video.

[0076] In some exemplary embodiments, the execution of these modules, whether on the encoder side or the decoder side, can follow a predefined execution order. For example, the execution order of these modules (if used) can be predefined as the RoI module, followed by the spatial resampling module (e.g., spatial upsampling module), followed by the temporal resampling module (e.g., temporal upsampling module), followed by the post-filter module, followed by the bit depth recovery module, and then the format adapter module. In another example, the execution order of these modules (if applied) can be predefined as the bit depth recovery module, followed by the RoI module, followed by the spatial resampling module, followed by the temporal resampling module, followed by the post-filter module, and then the format adapter module. In yet another example, the execution order of these modules (if applied) can be predefined as the RoI module, followed by the temporal resampling module, followed by the spatial resampling module, followed by the post-filter module, followed by the bit depth recovery module, and then the format adapter module. Any other predefined execution order of these modules can be predefined. Because the execution order of these modules is predefined, both the encoder and decoder know the execution order and no signaling in the bitstream is required. In one of these orders, the output of the previous module is input into the next module in a sequential manner.

[0077] The above-mentioned predefined execution order can be defined for the encoder. Thus, as an example, the encoder can execute the encoded versions of these modules in the predefined order, while the decoder can execute the decoded versions of these modules in the order opposite to the predefined order.

[0078] The above-mentioned predefined execution order can be defined for the decoder. In this way, the encoder can choose to execute the encoded versions of these modules in the order opposite to the predefined order, while the decoder can execute the decoded versions of these modules in the predefined order.

[0079] In some embodiments, a predefined execution order of these modules may be mandatory or suggested only for the decoder, while the encoder may continue to flexibly determine the execution order when performing encoding.

[0080] The following further exemplary embodiments provide for a flexible execution order of these modules on either or both the encoder and decoder sides. This flexible order execution of the encoding / decoding modules can be implemented at any encoding / decoding level, for example, at the frame level, sequence level, picture level, and any other suitable level. This flexible execution order may also be referred to as an adaptive execution order. This flexibility or adaptability can provide higher coding gain and less coding loss, as well as higher performance for machine vision in a VCM.

[0081] In some exemplary embodiments, a set of execution orders may be predefined, and one of them may be adaptively selected at various coding levels (e.g., frame level, slice level, picture level, etc.). The set of predefined execution orders may be known to both the encoder and the decoder and may be identified by an index in the set of execution orders.

[0082] For example, the set of predefined execution orders may include a first predefined order, in which the RoI module is executed first, followed by the spatial resampling module, followed by the temporal resampling module, followed by the post-filter module, followed by the bit depth recovery module, and then the format adapter module. The set of predefined execution orders may also include a second predefined order, in which the bit depth recovery module is executed first, followed by the RoI module, followed by the spatial resampling module, followed by the temporal resampling module, followed by the post-filter module, and then the format adapter module. The set of predefined execution orders may also include a third predefined order, in which the RoI module is executed first, followed by the temporal resampling module, followed by the spatial resampling module, followed by the post-filter module, followed by the bit depth recovery module, and then the format adapter module. The set of predefined execution orders may also include other execution orders. The plurality of modules may include other modules not explicitly described above. These predefined execution orders may be specified from an encoding perspective or from a decoding perspective.

[0083] For flexibility, an execution order can be adaptively selected from the set of predefined execution orders at a time. Alternatively, the execution order can be adaptively modified on a frame-by-frame, stripe-by-strip, etc. The selection or modification of the execution order can be adaptively varied on a frame-by-frame, sequence-by-sequence, stripe-by-strip, image-by-image, etc.

[0084] In some exemplary embodiments, the adaptability of the execution order may be based on one of at least one content characteristic associated with the video being encoded / decoded. The at least one content characteristic includes, but is not limited to, temporal dependency, content complexity, color complexity, motion complexity, texture characteristics, dynamic range variation, foreground-background segmentation, ratio of high-frequency regions, and ratio of RoI regions.

[0085] The specific execution order of the various modules described above may be determined by one of the content characteristics described above. Alternatively, a modification of the current execution order may be based on one of the content characteristics described above. For example, the execution order may be determined or modified on a frame-by-frame basis based on motion complexity. In particular, consider a video where the action moves quickly from one frame to the next. In this case, an adaptive change (e.g., of the execution order of the decoder) may involve increasing the priority of a temporal resampling module (e.g., a temporal interpolation module) or selecting an execution order that prioritizes temporal resampling modules. In other words, in the selected or modified execution order, the temporal resampling module is executed earlier.

[0086] For example, the decoder can analyze the two consecutive frames by comparing corresponding pixels or groups of pixels in the two consecutive frames. If the change in pixel values of a large number of pixels exceeds a certain threshold value, the decoder can determine that fast motion is detected and determine that the execution order of the earlier execution time resampling module can be used.

[0087] In another example, the decoder can quantify movement by calculating the average change in pixel values between frames. This average change in pixel values can be used by the decoder to detect fast movement. For example, when the average change in pixel values in a frame is higher than a predefined pixel value change threshold, the decoder can determine that fast movement has been detected. Alternatively, when the percentage of the average change between frames relative to the average pixel value is higher than a predefined percentage threshold, the decoder can determine that fast movement has been detected. A higher percentage or a larger average change generally indicates more significant movement, which means that the speed of movement is faster. When the decoder detects such fast movement, the execution order of the priority time resampling module (or the earlier executed time resampling module) can be used.

[0088] Alternatively, the specific execution order of the above modules can be adaptively determined or modified based on the content complexity. As an example, consider a complex urban scene with details, including moving vehicles, pedestrians, and constantly changing signs. In such a scene, the decoder can be configured to evaluate the spatial complexity by analyzing the changes in pixel values in one or more frames, detecting dense activities and complex details that characterize the scene. For example, the content complexity can be quantified by the change in pixel values across frames. In such high content complexity situations, the adaptive change of the execution order can involve prioritizing the spatial resampling module in the processing chain of the decoder. For example, when the change in pixel values across frames is equal to or above a predefined threshold, the spatial resampling module can be given priority over when the change in pixel values across frames is below the predefined threshold.

[0089] As an alternative, the specific execution order of the above-mentioned modules can be adaptively determined or modified based on the ratio of the high-frequency region. For example, the decoder can perform an analysis to identify regions within the frame that exhibit high spatial frequencies, which indicate fine details and distinct edges. The ratio of the high-frequency region can be quantified based on the proportion of the frame region containing these details compared to the overall size of the frame. If this ratio exceeds a predetermined ratio threshold, the decoder can adaptively prioritize the modules aimed at performing sharpening and detail enhancement, particularly for these high-frequency regions.

[0090] As an alternative, the specific execution order of the above-mentioned modules can be adaptively determined based on the RoI region. For example, the decoder can perform an analysis to determine the RoI region of the frame. If this ratio is higher than a predetermined threshold, the decoder can adaptively select or modify its execution order by prioritizing the RoI module among other modules.

[0091] In some of the exemplary embodiments provided above, a threshold for a specific content characteristic can be established or predefined, and depending on whether the value of the specific content characteristic exceeds or is lower than this threshold, the execution order can be adaptively selected or modified. However, in some other exemplary embodiments, multiple thresholds for a specific content characteristic can be established or predefined, thereby allowing for a more refined adjustment and greater flexibility of the execution order.

[0092] For example, in a scenario where the content characteristic is time-dependent, three different thresholds: threshold1, threshold2, and threshold3 can be established or predefined, where threshold1 < threshold2 < threshold3. This setting enables four different adaptive selection methods or four different adaptive modification methods of the execution order based on the relationship between the time-dependent representation value and these multiple thresholds.

[0093] In some of the above exemplary embodiments, the modification of the current execution order according to a specific content characteristic can involve advancing the corresponding module relative to other modules while keeping the order of other modules unchanged. The amount of advancement of the corresponding module (the number of positions advanced in the execution order) can be determined by the quantified value of the specific content characteristic.

[0094] Although the above examples are based on specific content characteristics to determine or modify the execution order of the modules, the basic principle also applies to the use of a combination of content characteristics. Only as an example, the execution order of the decoder modules can be adaptively determined or modified based on both time-dependence and content complexity.

[0095] In some exemplary embodiments, a neural network model can be used to adaptively predict the order in which modules are executed. The neural network can be pre-trained. For example, the neural network can take a set of derived video characteristics as input and generate a predicted execution order for the various modules described above on a frame-by-frame, sequence-by-sequence, or image-by-image basis.

[0096] In some exemplary embodiments, the module execution order may be determined by information related to at least one of the modules. For example, the execution order of a module may be adaptively determined by additional information about other modules. This additional information may include information about temporal resampling, spatial resampling, and bit depth truncation.

[0097] For example, the temporal resampling module may be associated with a temporal resampling rate (e.g., an upsampling rate). In one example, if the temporal resampling rate is 2, the execution order may be set to: temporal resampling module (e.g., a temporal upsampling module), spatial resampling module, RoI module, post-filter module, bit depth recovery module, format adapter module. Otherwise, if the temporal resampling rate is 4, the execution order may be set to: bit depth recovery module, spatial resampling module, RoI module, temporal resampling module, post-filter module, format adapter module.

[0098] For another example, a spatial resampling module can be associated with a spatial resampling rate. The module execution order can depend on the spatial resampling rate. For example, if the spatial resampling rate is 0.5, a first predefined module execution order can be used. If the spatial resampling rate is 0.75, a second predefined module execution order can be used. If the spatial resampling rate has other values, a third predefined module execution order can be used.

[0099] The above module information can be represented by a signal in the video bitstream for extraction by a decoder without actually decoding or reconstructing the video frame.

[0100] In some exemplary embodiments, a flexible or adaptive execution order of decoder modules can be facilitated by a signaling mechanism originating from the encoder side. Specifically, the encoder can perform the various analyses described above or other analyses on the input video content and adaptively determine the execution order of the decoding modules (e.g., frame by frame, strip by strip, sequence by sequence, image by image, etc.), and signal the decoder with this order or reordering in the bitstream. This approach ensures that the decoder's dynamic module ordering or reordering is directly informed by the encoder's analysis of the video content, thereby allowing a highly optimized decoding process customized to the specific characteristics of each video sequence or frame or strip or image. The signaling of the execution order can be implemented through high-level syntax embedded in the video stream (such as Supplemental Enhancement Information (SEI) messages, Video Supplemental Enhancement Information (VSEI) or other metadata carriers designed for this purpose). These messages can contain instructions to the decoder regarding the preferred or optimal order for module execution for the upcoming frame or sequence.

[0101] An example of a signaling syntax structure is shown below.

[0102] The following describes each of the above syntax elements:

[0103] adaptive_module_execution_order_flag indicates whether adaptive execution order of decoding modules is to be applied. This syntax element equal to 1 indicates that adaptive module execution order is enabled. This syntax element equal to 0 indicates that adaptive module execution order is disabled.

[0104] module_execution_order_idx is an n-bit unsigned integer whose value can range from 0 to n-1. Each index value indicates an execution order. For example, if the descriptor of module_execution_order_idx is u(2), then module_execution_order_idx can be equal to 0, 1, 2, or 3, indicating one of four different execution orders that can be predefined or signaled separately. Just as an example:

[0105] Idx 0 may indicate the following order: RoI module, spatial resampling module, temporal resampling module, post-filter module, bit depth recovery module, format adapter module.

[0106] Idx 1 may indicate the following order: spatial resampling module, RoI module, temporal resampling module, post-filter module, bit depth recovery module, format adapter module.

[0107] Idx 2 may indicate the following order: temporal resampling module, spatial resampling module, RoI module, post-filter module, bit depth recovery module, format adapter module.

[0108] Idx 3 may indicate the following order: bit depth recovery module, spatial resampling module, RoI module, temporal resampling module, post-filter module, format adapter module.

[0109] The above exemplary embodiments focus on the decoder's analysis of either or both signaling information and content characteristics / module information to adaptively determine the execution order of each decoding module. However, when an encoder uses encoding modules corresponding to each decoding module, the encoder can follow a similar analysis approach to determine the execution order of the encoding modules. In some implementations, the execution order of the encoding modules and the execution order of the decoding modules adopted by the encoder and decoder, respectively, can be symmetrical or mirror images of each other, meaning that the module executed earlier on the encoder side is processed last on the decoder side. In other words, the execution order of the decoding modules can be opposite to the execution order of the corresponding encoding modules. For example, the encoder can determine the order of the encoding modules and signal the symmetrical order of the decoding modules in the bitstream. For another example, both the encoder and decoder can follow the same analysis to determine the execution order of the encoding or decoding modules based on content characteristics and / or additional module information. The resulting execution order of the encoding modules and the execution order of the decoding modules can be symmetrical.

[0110] In some other exemplary embodiments, the above symmetry restriction can be lifted, allowing the execution order on the decoder side to be asymmetric or independent of the execution order on the encoder side. In such implementations, different order adaptation methods can be applied on the encoder side and the decoder side. Even if the execution order is signaled in the bitstream by the encoder, the decoder may not need to follow this signaled order and can still decide to perform analysis and adaptive execution order determination independently.

[0111] Figure 5 A flowchart of an example method (500) of an embodiment of the present application is shown. The method (500) starts at step (S501). In step (S510), an encoded bitstream of a video is received. In step (S520), a sequential order for executing multiple decoding modules is determined based on an indication extracted from the encoded bitstream. In step (S530), the encoded bitstream is decoded by executing the multiple decoding modules in the sequential order. The multiple decoding modules include at least one of the following: a region of interest (ROI) module, a temporal upsampling module, a spatial upsampling module, a post-filtering module, a bit depth recovery module, or a format adapter module. The method (500) stops at (S599).

[0112] Figure 6 A flowchart of an example method (600) of an embodiment of the present application is shown. The method (600) starts at step (S601). In step (S610), a sequential order for executing multiple encoding modules is adaptively determined. In step (S620), a video is encoded by executing multiple encoding modules in the sequential order to generate an encoded bitstream. The multiple encoding modules include at least one of the following: a region of interest (ROI) module, a temporal upsampling module, a spatial upsampling module, a post-filtering module, a bit depth recovery module, or a format adapter module. In step (S630), an indication is added to the encoded bitstream to indicate the sequential order to a decoder. The program (600) stops at (S699).

[0113] Methods (500) and (600) may be adjusted appropriately. At least one step in methods (500) and (600) may be modified and / or omitted. At least one additional step may be added. Any suitable implementation order may be used.

[0114] The techniques disclosed in this application can be used alone or in any combination in any order. In addition, each of the techniques (e.g., methods, embodiments), encoders, and decoders can be implemented by processing circuitry (e.g., at least one processor or at least one integrated circuit). In some examples, at least one processor executes a program stored in a non-volatile computer-readable medium.

[0115] Although several exemplary embodiments have been described herein, there are variations, permutations, and various equivalent alternative implementations that fall within the scope of the present invention. Therefore, it should be understood that those skilled in the art will be able to devise many systems and methods that, although not explicitly shown or described herein, embody the principles of the present invention and are therefore within the spirit and scope of the present invention.

Claims

1. A video decoding method, characterized in that: include: Receive the encoded video stream; determining a sequential order of executing a plurality of decoding modules based on an indication extracted from the encoded code stream; and Decoding the encoded code stream by executing the plurality of decoding modules according to the sequential order, The plurality of decoding modules include at least one of the following: a region of interest module, a temporal upsampling module, a spatial upsampling module, a post-filtering module, a bit depth recovery module, or a format adapter module.

2. The method according to claim 1, characterized in that The sequential order is selected from a set of predefined sequential orders based on the indication.

3. The method according to claim 2, characterized in that If the plurality of decoding modules are executed, a set of predefined sequential orders of the plurality of decoding modules includes at least one of the following: the region of interest module, followed by the spatial upsampling module, followed by the temporal upsampling module, followed by a post-filtering module, followed by the bit depth recovery module, and followed by the format adapter module; the bit depth recovery module, followed by the region of interest module, followed by the spatial upsampling module, followed by the temporal upsampling module, followed by the post filter module, and followed by the format adapter module; and The region of interest module, followed by the temporal upsampling module, followed by the spatial upsampling module, followed by the post filter module, followed by the bit depth recovery module, and followed by the format adapter module.

4. The method according to claim 1, wherein The sequential order is adaptively determined.

5. The method according to claim 4, characterized in that The indication is derived from at least one content characteristic of the video, the at least one content characteristic being extracted from the encoded bitstream; and The at least one content characteristic comprises at least one of: time dependency, content complexity, color complexity, motion complexity, texture characteristics, dynamic range variation, foreground-background segmentation, ratio of high frequency areas, ratio of regions of interest.

6. The method according to claim 5, characterized in that The indication is based on inter-frame content movement, and a higher degree of the inter-frame content movement indicates an earlier execution of the temporal upsampling module.

7. The method according to claim 6, characterized in that The inter-frame content movement is determined by: The number of pixels with frame-to-frame variation; the percentage of pixels with inter-frame value changes; and / or Average inter-frame pixel value change.

8. The method according to claim 5, characterized in that The indication is derived based on the content complexity, and a higher level of the content complexity indicates an earlier execution of the spatial upsampling module.

9. The method according to claim 5, characterized in that The indication is made based on the ratio of the high frequency region, and a higher level of the ratio of the high frequency region indicates an earlier execution of the spatial upsampling module.

10. The method according to claim 5, characterized in that The indication is made based on a ratio of the region of interest, and wherein a higher level of the ratio of the region of interest indicates an earlier execution of the region of interest module.

11. The method according to claim 5, characterized in that The sequential order is selected from a plurality of sequential orders based on a set of threshold levels of the at least one content characteristic.

12. The method according to claim 4, characterized in that The indication is derived by processing a set of video characteristics derived from the encoded bitstream using a pre-trained neural network.

13. The method according to claim 4, characterized in that The indication is derived from a set of additional parameters extracted from the encoded codestream associated with at least one decoding module of the plurality of decoding modules.

14. The method according to claim 13, characterized in that The set of additional parameters includes at least one of a temporal resampling rate or a spatial resampling rate.

15. The method according to claim 4, characterized in that Further including: Determining the sequential order is adaptively determined based on a flag in the encoded code stream; and The sequential order is determined based on a signaled index, the index being an index of the sequential order among a plurality of predefined sequential orders.

16. The method according to claim 1, characterized in that The sequential order is asymmetric to an encoding order of a plurality of encoding modules used to generate the encoded code stream, the plurality of encoding modules corresponding to the plurality of decoding modules.

17. A video encoding method, characterized in that: include: adaptively determining a sequential order for executing a plurality of encoding modules; Encoding the video by executing the plurality of encoding modules in the sequential order to generate an encoded bitstream; and adding an indication to the encoded bitstream to indicate the sequential order to a decoder, The plurality of encoding modules include at least one of the following: a region of interest module, a temporal upsampling module, a spatial upsampling module, a post-filtering module, a bit depth recovery module or a format adapter module.

18. The method according to claim 17, characterized in that The sequential order is determined based on at least one content characteristic of the video; and The at least one content characteristic comprises at least one of: time dependency, content complexity, color complexity, motion complexity, texture characteristics, dynamic range variation, foreground-background segmentation, ratio of high frequency areas, ratio of regions of interest.

19. The method according to claim 17, wherein The instructions include: a flag for indicating that the sequential order is adaptively determined; and An index of the sequential order among a plurality of predefined sequential orders.

20. A non-transitory computer-readable storage medium for storing an encoded video stream, characterized in that: The encoded code stream includes: A flag for indicating whether the decoder uses an adaptive sequential order to execute multiple decoding modules; and An indication is provided for, when the flag indicates to use the adaptive sequential order, executing the adaptive sequential order among a plurality of predefined sequential orders of the plurality of decoding modules.