Method, system, and program for video encoding using a hierarchical structure for neural network-based tools
By generating a virtual reference frame based on hierarchical levels for video coding, the method addresses the inefficiencies in traditional video coding standards, improving encoding efficiency and compression quality.
Patent Information
- Application Number
- JP2022559335
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-09-17
- Filing Date
- 2021-09-28
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2041-09-28
Smart Images

Figure 0007695042000010 
Figure 0007695042000011 
Figure 0007695042000012
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority from U.S. Provisional Patent Application No. 63 / 136,055, filed on January 11, 2021, and U.S. Patent Application No. 17 / 478,138, filed on September 17, 2021, the entire contents of which are incorporated herein by reference.
[0002] Field The present disclosure generally relates to the field of data processing, and more particularly to video encoding.
Background Art
[0003] Video encoding and decoding using inter - picture prediction with motion compensation have been known for decades. Uncompressed digital video can be composed of a series of pictures, each picture having, for example, spatial dimensions of 1920×1080 luminance samples and associated chrominance samples. A series of pictures can have a fixed or variable picture rate (also known informally as the frame rate), for example, 60 pictures per second or 60 Hz. Uncompressed video has significant bit - rate requirements. For example, 1080p60 4:2:0 video with 8 bits per sample (1920×1080 luminance sample resolution at a frame rate of 60 Hz) requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires a storage space of more than 600 GB.
[0004] Traditional video coding standards such as H.264 / Advanced Video Coding (H.264 / AVC), High-Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC) share a similar (recursive) block-based hybrid prediction / transformation framework, where individual coding tools such as intra / inter prediction, integer transformation, and context-adaptive entropy coding are crafted intensively to optimize the overall efficiency. Summary of the Invention Means for Solving the Problems
[0005] Embodiments relate to a method, a system, and a computer-readable medium for video coding. According to one aspect, a method for video coding is provided. The method may include receiving video data including a current picture. A virtual reference frame is generated for the current picture based on a hierarchical level associated with the current picture and the most recently decoded picture. The video data is decoded based on the generated reference frame.
[0006] According to another aspect, a computer system for video encoding is provided. The computer system includes one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage devices, and program instructions stored in at least one of the one or more storage devices for execution by at least one of the one or more processors via at least one of the one or more memories, whereby the computer system can execute a method. The method may include receiving video data including a current picture. A virtual reference frame is generated for the current picture based on a hierarchical level associated with the current picture and the most recently decoded picture. The video data is decoded based on the generated reference frame.
[0007] According to yet another aspect, a computer-readable medium for video encoding is provided. The computer-readable medium may include one or more computer-readable storage devices and program instructions stored in at least one of the one or more tangible storage devices, the program instructions being executable by a processor. The program instructions are executable by a processor to execute a method that may include receiving video data including a current picture. A virtual reference frame is generated for the current picture based on a hierarchical level associated with the current picture and the most recently decoded picture. The video data is decoded based on the generated reference frame.
Brief Description of the Drawings
[0008] These and other objects, features, and advantages will become apparent from the following detailed description of exemplary embodiments read in conjunction with the accompanying drawings. The drawings are for the purpose of facilitating understanding by those skilled in the art in connection with the detailed description and, for the sake of clarity, the various features of the drawings are not to scale.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
DETAILED DESCRIPTION OF THE INVENTION
[0009] Although detailed embodiments of the claimed structures and methods are disclosed herein, it can be understood that the disclosed embodiments merely exemplify the claimed structures and methods that can be embodied in various forms. However, these structures and methods can be embodied in many different forms and should not be construed as limited to the exemplary embodiments described herein. Rather, these exemplary embodiments are provided so that this disclosure will be thorough and complete and will fully convey the scope to those skilled in the art. Details of well-known features and techniques may be omitted in this document to avoid unnecessarily obscuring the presented embodiments.
[0010] Embodiments generally relate to the field of data processing, and more particularly, to video coding. The exemplary embodiments described below provide a system, method, and computer program for encoding and decoding video data based on a hierarchical temporal structure for loop filter / inter prediction in particular. Thus, some embodiments have the ability to improve the field of computing by allowing for improved efficiency in video coding.
[0011] As described above, video encoding and decoding using inter-picture prediction with motion compensation has been known for decades. Uncompressed digital video can be composed of a series of pictures, each picture having spatial dimensions, for example, of 1920×1080 luminance samples and associated chrominance samples. The series of pictures can have a fixed or variable picture rate (also informally known as the frame rate), for example, 60 pictures per second or 60 Hz. Uncompressed video has significant bitrate requirements. For example, 1080p60 4:2:0 video at 8 bits per sample (1920×1080 luminance sample resolution at a frame rate of 60 Hz) requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires storage space exceeding 600 GB. Traditional video coding standards such as H.264 / Advanced Video Coding (H.264 / AVC), High-Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC) share a similar (recursive) block-based hybrid prediction / transformation framework where individual coding tools such as intra / inter prediction, integer transformation, and context adaptive entropy coding are crafted intensively to optimize the overall efficiency.
[0012] One purpose of video encoding and decoding can be the reduction of redundancy in the input video signal by compression. Compression can in some cases help to reduce the aforementioned bandwidth or memory space requirements by more than an order of magnitude. Both reversible compression and irreversible compression, as well as combinations thereof, can be used. Reversible compression refers to a technique in which an exact copy of the original signal can be reconstructed from the compressed original signal. When using irreversible compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough to make the reconstructed signal useful for its intended purpose. In the case of video, irreversible compression is widely used. The amount of acceptable distortion depends on the application. For example, users of certain consumer streaming applications may tolerate higher distortion than users of television contribution applications. The achievable compression ratio can reflect the following: higher acceptable / tolerable distortion can result in a higher compression ratio.
[0013] The spatio-temporal pixel neighborhood is utilized for constructing a prediction signal to obtain corresponding residuals for subsequent transformation, quantization, and entropy encoding. On the other hand, the nature of a neural network (NN) is to extract different levels of spatio-temporal stimuli by analyzing spatio-temporal information from the receptive fields of neighboring pixels. The high degree of non-linearity and the ability to explore non-local spatio-temporal correlations offer promising opportunities to significantly improve compression quality.
[0014] However, one consideration in utilizing information from multiple neighboring video frames is the complex motion caused by camera motion and dynamic scenes. Conventional block-based motion vectors do not work well for non-translational motion. Learning-based optical flow can provide accurate motion information at the pixel level, but unfortunately, errors are prone to occur, especially along the boundaries of moving objects. In some hybrid inter-frame prediction, an NN-based model is applied to implicitly handle any complex motion in a data-driven manner.
[0015] Therefore, in order to obtain a better trade-off for performance and encoding runtime, when using an NN-based model as an LF, or for an inter-prediction tool, it may be advantageous to select different frames as reference frames for applying a loop filter (LF), or to generate intermediate frames.
[0016] Aspects are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer-readable media according to various embodiments. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0017] The exemplary embodiments described below provide a system, method, and computer program for encoding and decoding video data. Referring now to FIG. 1, there is shown a functional block diagram of a networked computer environment showing a video encoding system 100 (hereinafter, the "system") for encoding and decoding video data. It should be understood that FIG. 1 provides only an illustration of one implementation and does not imply any limitation as to the environments in which various embodiments may be implemented. Many modifications to the illustrated environment may be made based on design and implementation requirements.
[0018] System 100 may include a computer 102 and a server computer 114. The computer 102 can communicate with the server computer 114 via a communication network 110 (hereinafter referred to as the "network"). The computer 102 can include a processor 104 and a software program 108 stored in a data storage device 106, having an interface with a user and being enabled to communicate with the server computer 114. As will be described later with reference to FIG. 4, the computer 102 may include internal components 800A and external components 900A respectively, and the server computer 114 may include internal components 800B and external components 900B respectively. The computer 102 may be, for example, a mobile device, a phone, a personal digital assistant, a netbook, a laptop computer, a tablet computer, a desktop computer, or any type of computing device capable of executing a program, accessing a network, and accessing a database.
[0019] The server computer 114 may also operate in a cloud computing service model such as software as a service (SaaS), platform as a service (PaaS), or infrastructure as a service (IaaS), as will be described later with respect to FIGS. 5 and 6. The server computer 114 may also be located within a cloud computing deployment model such as a private cloud, a community cloud, a public cloud, or a hybrid cloud.
[0020] Server computer 114 that can be used to encode and decode video data is adapted to execute a video encoding program 116 (hereinafter referred to as "program") that can interact with database 112. The video encoding program method will be described in more detail below with respect to FIG. 3. In one embodiment, computer 102 may operate as an input device including a user interface, while program 116 may operate primarily on server computer 114. In an alternative embodiment, program 116 may operate primarily on one or more computers 102, while server computer 114 may be used for processing and storing data used by program 116. It should be noted that program 116 may be an independent program or may be integrated into a larger video encoding program.
[0021] However, it should be noted that the processing for program 116 can, in some cases, be shared between computer 102 and server computer 114 in any ratio. In another embodiment, program 116 may operate on two or more computers, server computers, or some combination of computers and server computers, for example, on multiple computers 102 that communicate with a single server computer 114 across network 110. In another embodiment, for example, program 116 may operate on multiple server computers 114 that communicate with multiple client computers across network 110. Alternatively, the program may also operate on a network server that communicates with servers and multiple client computers across the network.
[0022] Network 110 may include a wired connection, a wireless connection, a fiber optic connection, or some combination thereof. In general, network 110 can be any combination of connections and protocols that support communication between computer 102 and server computer 114. Network 110 can be, for example, a local area network (LAN), a wide area network (WAN) such as the Internet, a telecommunications network such as the public switched telephone network (PSTN), a wireless network, a public switched network, a satellite network, a cellular network (e.g., a fifth generation (5G) network, a long term evolution (LTE) network, a third generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a metropolitan area network (MAN), a private network, an ad hoc network, an intranet, a fiber optic-based network, etc., and / or combinations of these or other types of networks of various types.
[0023] The number and arrangement of the devices and networks shown in FIG. 1 are provided as an example. In practice, there may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks in a different arrangement than those shown in FIG. 1. Further, two or more of the devices shown in FIG. 1 may be implemented within a single device, or a single device shown in FIG. 1 may be implemented as a plurality of distributed devices. Additionally or alternatively, a set of the devices of system 100 (e.g., one or more devices) may perform one or more of the functions described as being performed by another set of the devices of system 100.
[0024] Referring now to FIG. 2, a hierarchical temporal structure 200 for loop filter / inter prediction is shown. The hierarchical temporal structure 200 may be used to apply an NN-based loop filter or inter-frame prediction in video encoding and decoding to determine the number / index of frames applied as reference frames for frame generation or improvement in video encoding for I / P / B frames, for using an NN model. A plurality of image frames x1, … x T Assume an input video x including (e.g., frames 1 to 16). In a first motion estimation step, the frame is divided into spatial blocks, and each block can be successively and iteratively divided into smaller blocks. The current frame x t and a set of previous reconstructed frames
Number
Number
Number
Number
Number
[0025] The quantization stage gives a quantized transform block. Both the motion vector m t and the quantized transform block are encoded into a bitstream by entropy coding, and this is sent to the decoder. Then, on the decoder side, the decoded block has inverse transformation and dequantization (typically through an inverse transformation such as IDCT with dequantized coefficients) applied to obtain a restored residual
Number
Number
Number
Number
[0026] In HEVC, VVC, or other video coding frameworks or standards, the decoded picture may be included in a reference picture list (RPL) and used as a reference picture for motion-compensated prediction, for coding subsequent pictures in encoding or decoding order, and for other parameter prediction. Alternatively, the decoded portion of the current picture may be used for intra prediction or intra-block copy for coding different regions or blocks of the current picture.
[0027] In one example, one or more virtual references may be generated and included in the RPL in both the encoder and decoder, or in the decoder only. The virtual reference picture may be generated by one or more processes including signal processing, spatial or temporal filtering, scaling, weighted averaging, up / downsampling, pooling, recursive processing in memory, linear system processing, non-linear system processing, neural network processing, deep learning-based processing, AI processing, pre-trained network processing, machine learning-based processing, online training network processing, or combinations thereof. For the process of generating the virtual reference, zero or more forward reference pictures preceding the current picture in both output / display order and encoding / decoding order, and zero or more backward reference pictures following the current picture in output / display order but preceding the current picture in encoding / decoding order are used as input data. The output of the process is a virtual picture / generated picture used as a new reference picture. Conventional motion compensation techniques may be applied when this new reference picture is selected to predict an encoded block in the current picture.
[0028] In one example, the NN-based method may be applied to loop filter design in combination with one or more of the additional components (e.g., DF, SAO, ALF, CCALF, etc.) described above at both the slice / CTU level on each frame, or to replace one or more of the additional components (e.g., DF, SAO, ALF, CCALF, etc.) described above. When applied, the reconstructed current picture is used as input data to the NN-based model to generate an NN-enhanced filtered picture. For each block or CTU, a decision can be made to select this NN-enhanced filtered picture as the post-filtering result or to use a conventional filtering method.
[0029] NN-based video coding tools may be applied to a picture if their hierarchical level IDs meet certain conditions. In one embodiment, the NN-based video coding tool may be applied to pictures having a temporal level ID below a given threshold. In other words, the NN-based video coding tool may not be applied to pictures having a temporal level ID greater than a given threshold. In another embodiment, the NN-based video coding tool may be applied to pictures having a temporal level ID greater than a given threshold. In other words, the NN-based video coding tool may not be applied to pictures having a temporal level ID below a given threshold. In another embodiment, the NN-based video coding tool may be applied only to pictures having a specific temporal level ID. In another embodiment, the NN-based video coding may be NN-based inter prediction or loop filtering, or both.
[0030] The hierarchical structure can be extended to a predefined prediction structure, and some of the pictures in the sequence can be used as references for other pictures, while some other pictures may not be used as references at all. In other cases, some of the pictures in the sequence are considered more important than others. They are encoded by setting a smaller QP. These pictures can be used as references more frequently than other pictures. In some cases, these pictures may be assigned with a certain hierarchical temporal level ID, and the NN-based video encoding tool may be applied to these pictures in the same sense as in the above embodiments, but may not be applied to the remaining pictures in the sequence.
[0031] The hierarchical structure for the NN-based encoding tool in video encoding may determine whether to apply the neural network-based encoding tool to a picture according to the hierarchical level to which the picture belongs. The NN-based encoding tool may include, but is not limited to, the NN-based virtual reference picture for inter prediction and the NN-based loop filtering. The following are some examples for further explaining the proposed method in more detail.
[0032] In one example, in the hierarchical temporal structure 200, when the current picture has a picture order count (POC) equal to 3, typically, decoded pictures having a POC equal to 0, 2, 4, or 8 may be stored in the decoded picture buffer, and some of them are included in the reference picture list (RPL) for decoding the current picture (POC = 3). For example, when a level 4 frame is selected as the frame for applying the NN-based coding tool, in most NN-based inter prediction models, to generate a virtual reference frame for the current picture (POC = 3), the nearest decoded picture having a POC equal to 2 or 4 can be used as input data for supplying to the virtual reference generation process using the NN-based model. The NN-based inter prediction model is applied 8 times to generate all virtual reference frames as reference pictures for each picture at level 4. In the same or another embodiment, when selecting temporal level 4 as the filter application frame, when using the NN-based loop filter for detail enhancement or post-filtering, the NN-based loop filter generation is launched 8 times, and at temporal level 4, from each picture, an NN-enhanced filtered picture is obtained. In the same or another embodiment, all pictures having a temporal level ID of 3 or less do not apply the NN-based video coding tool. In another example, all pictures having a temporal level ID of 3 or less apply the NN-based video coding tool, such as the NN-based loop filter or NN-based inter prediction.
[0033] Referring now to FIG. 3, an operation flowchart showing steps of a method 300 executed by a program for encoding video data based on a hierarchical temporal structure for loop filter / inter prediction is shown.
[0034] At 302, method 300 may include receiving video data including a current picture.
[0035] At 304, method 300 may include generating a virtual reference frame for the current picture based on a hierarchical level associated with the current picture and the most recently decoded picture.
[0036] At 306, method 300 may include decoding the video data based on the generated reference frame.
[0037] It can be understood that FIG. 3 only provides an illustration of one implementation and does not imply any limitation on how different embodiments can be implemented. Many modifications to the illustrated environment can be made based on design and implementation requirements.
[0038] FIG. 4 is a block diagram 400 of the internal and external components of the computer shown in FIG. 1 according to an exemplary embodiment. It should be understood that FIG. 4 only provides an illustration of one implementation and does not imply any limitation on the environment in which different embodiments can be implemented. Many modifications to the illustrated environment can be made based on design and implementation requirements.
[0039] Computer 102 (FIG. 1) and server computer 114 (FIG. 1) may each include a respective set of internal components 800A, B and external components 900A, B shown in FIG. 5. Each set of internal components 800 includes one or more processors 820 on one or more buses 826, one or more computer-readable RAMs 822, and one or more computer-readable ROMs 824, one or more operating systems 828, and one or more computer-readable tangible storage devices 830.
[0040] Processor 820 is implemented in hardware, firmware, or a combination of hardware and software. Processor 820 is a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or another type of processing component. In some implementations, Processor 820 includes one or more processors that can be programmed to perform functions. Bus 826 includes components that enable communication between internal components 800A, B.
[0041] One or more operating systems 828, software programs 108 (FIG. 1), and video coding programs 116 (FIG. 1) on server computer 114 (FIG. 1) are stored in one or more of respective computer-readable tangible storage devices 830 for execution by one or more of respective processors 820 via one or more of respective RAMs 822 (typically including cache memory). In the embodiment shown in FIG. 4, each of computer-readable tangible storage devices 830 is a magnetic disk storage device of an internal hard drive. Alternatively, each of computer-readable tangible storage devices 830 is a semiconductor storage device such as ROM 824, EPROM, flash memory, an optical disk, a magneto-optical disk, a solid state disk, a compact disk (CD), a digital versatile disk (DVD), a floppy disk, a cartridge, a magnetic tape, and / or other types of non-transitory computer-readable tangible storage devices capable of storing computer programs and digital information.
[0042] Each set of internal components 800A, B also includes an R / W drive or interface 832 for reading from and writing to one or more portable computer-readable tangible storage devices 936 such as CD-ROMs, DVDs, memory sticks, magnetic tapes, magnetic disks, optical disks, or semiconductor memory devices. Software programs such as software program 108 (FIG. 1) and video encoding program 116 (FIG. 1) are stored on one or more of the respective portable computer-readable tangible storage devices 936, read through the respective R / W drives or interfaces 832, and can be loaded onto the respective hard drives 830.
[0043] Each set of internal components 800A, B also includes a network adapter or interface 836 such as a TCP / IP adapter card, a wireless Wi-Fi interface card, or a 3G, 4G, or 5G wireless interface card, or other wired or wireless communication link. Software program 108 (FIG. 1) and video encoding program 116 (FIG. 1) on server computer 114 (FIG. 1) can be downloaded from an external computer to computer 102 (FIG. 1) and server computer 114 via a network (e.g., the Internet, a local area network, or other wide area network) and the respective network adapters or interfaces 836. From the network adapter or interface 836, software program 108 and video encoding program 116 on server computer 114 are loaded onto the respective hard drives 830. The network can include copper wire, fiber optic, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers.
[0044] Each set of external components 900A, B can include a computer display monitor 920, a keyboard 930, and a computer mouse 934. The external components 900A, B can also include a touch screen, a virtual keyboard, a touch pad, a pointing device, and other human interface devices. Each set of internal components 800A, B also includes a device driver 840 for interfacing with the computer display monitor 920, the keyboard 930, and the computer mouse 934. The device driver 840, the R / W drive or interface 832, and the network adapter or interface 836 include hardware and software (stored in the storage device 830 and / or the ROM 824).
[0045] This disclosure includes a detailed description of cloud computing, but it is to be understood in advance that the implementation of the teachings described herein is not limited to a cloud computing environment. Rather, some embodiments are implementable in connection with any other type of computing environment, whether currently known or later developed.
[0046] Cloud computing is a service delivery model that enables convenient on-demand network access to a shared pool of configurable computing resources (such as networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, services), which can be rapidly provisioned and released with minimal management effort or interaction with the service provider. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.
[0047] The characteristics are as follows: On-demand self-service: Cloud consumers can unilaterally provision computing capabilities such as server time and network storage automatically as needed, without the need for human interaction with the service provider. Broad network access: The capabilities are available over the network and accessed through standard mechanisms that promote use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs). Resource pooling: The provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, and various physical and virtual resources are dynamically assigned and reassigned according to demand. Consumers generally have a degree of location independence in that they have no control over or knowledge of the exact location of the provided resources, but it may be possible to specify location at a higher level of abstraction (e.g., country, state, data center). Rapid elasticity: Capabilities can be provisioned rapidly and elastically, in some cases automatically, to scale out quickly and scale in rapidly. To the consumer, the capabilities available for provisioning often appear to be unlimited and can be purchased in any quantity at any time. Measured service: The cloud system automatically controls and optimizes resource use by leveraging a metering capability at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource use can be monitored, controlled, and reported, providing transparency for both the provider and consumer of the utilized service.
[0048] The service model is as follows: Software as a Service (SaaS): The functionality provided to consumers is to utilize the provider's applications operating on cloud infrastructure. The applications can be accessed from various client devices through a client interface such as a web browser (e.g., web-based email). Consumers do not manage or control the underlying infrastructure, including the network, servers, operating systems, storage, or even individual application features. However, there may be exceptions for settings of application configurations specific to the user. Platform as a Service (PaaS): The functionality provided to consumers is to deploy applications created or obtained by consumers using programming languages and tools supported by the provider on cloud infrastructure. Consumers do not manage or control the underlying cloud infrastructure including the network, servers, operating systems, or storage, but have control over the deployed applications and potentially the configuration of the environment hosting the applications. Infrastructure as a Service (IaaS): The functionality provided to consumers is to provision processing, storage, network, and other basic computing resources, and consumers can deploy and run any software that may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but have control over the operating systems, storage, deployed applications, and potentially limited control over selected network components (e.g., host firewalls).
[0049] The deployment models are as follows: Private Cloud: The cloud infrastructure is operated solely for a single organization. It may be managed by that organization or a third party and can exist either on - premise or off - premise. Community Cloud: The cloud infrastructure is shared by several organizations and supports a specific community with shared concerns (e.g., mission, security requirements, policies, and compliance considerations). It may be managed by those organizations or a third party and can exist either on - premise or off - premise. Public Cloud: The cloud infrastructure is owned by an organization selling cloud services and made available to the general public or large industry groups. Hybrid Cloud: The cloud infrastructure remains a distinct entity but is composed of two or more clouds (private, community, public) that are bound together by standardized or proprietary technologies (such as cloud bursting for load distribution between clouds) that enable data and application portability.
[0050] Cloud computing environments are service - oriented, focusing on statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is an infrastructure consisting of a network of interconnected nodes.
[0051] Referring to FIG. 5, an exemplary cloud computing environment 500 is shown. As illustrated, cloud computing environment 500 includes one or more cloud computing nodes 10, and local computing devices used by cloud consumers, such as a portable digital assistant (PDA) or mobile phone 54A, desktop computer 54B, laptop computer 54C, and / or automotive computer system 54N, can communicate with the one or more cloud computing nodes 10. The cloud computing nodes 10 can communicate with each other. They may be grouped physically or virtually into one or more networks such as the private, community, public, or hybrid clouds as described above, or combinations thereof (not shown). Thereby, cloud computing environment 500 can provide infrastructure, platform, and / or software as a service such that cloud consumers do not need to maintain resources on local computing devices. The types of computing devices 54A-N shown in FIG. 5 are merely intended to be exemplary, and it is understood that cloud computing node 10 and cloud computing environment 500 can communicate with any type of computerized device through any type of network and / or network addressable connection (e.g., using a web browser).
[0052] Referring to FIG. 6, a set of functional abstraction layers 600 provided by cloud computing environment 500 (FIG. 5) is shown. It should be understood in advance that the components, layers, and functions shown in FIG. 6 are merely intended to be exemplary and embodiments are not limited thereto. As illustrated, the following layers and corresponding functions are provided.
[0053] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include: mainframe 61, RISC (Reduced Instruction Set Computer) architecture-based server 62, server 63, blade server 64, storage device 65, and network and networking components 66. In some embodiments, software components include network application server software 67 and database software 68.
[0054] The virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities may be provided: virtual server 71, virtual storage device 72, virtual network 73 including a virtual private network, virtual applications and operating systems 74, and virtual client 75.
[0055] In one example, the management layer 80 may provide the functions described below. Resource provisioning 81 provides for the dynamic procurement of computing resources and other resources utilized to execute tasks within a cloud computing environment. Metering and valuation 82 provides for cost tracking when resources are utilized within a cloud computing environment and for billing or invoicing for the consumption of these resources. In one example, these resources can include application software licenses. Security provides for the authentication of cloud consumers and tasks and for the protection of data and other resources. The user portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 84 provides for the allocation and management of cloud computing resources such that the required service levels are met. Planning and fulfillment of service level agreements (SLAs) 85 provides for the advance arrangement and procurement of cloud computing resources for which future requirements are predicted according to the SLA.
[0056] The workload layer 90 provides, for that purpose, examples of functions where a cloud computing environment can be utilized. Examples of workloads and functions that can be provided from this layer include mapping and navigation 91, software development and life cycle management 92, virtual classroom education delivery 93, data analysis processing 94, transaction processing 95, and video encoding 96. Video encoding 96 can use a hierarchical temporal structure for loop filter / inter prediction.
[0057] Some embodiments can relate to systems, methods, and / or computer-readable media at any possible level of technological detail integration. The computer-readable media can include a computer-readable non-transitory storage medium having thereon computer-readable program instructions for causing a processor to execute operations.
[0058] A computer-readable storage medium may be a tangible device that holds and stores instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVDs), memory sticks, floppy disks, punch cards, or mechanically encoded devices such as raised structures within grooves having recorded instructions, and any suitable combination thereof. As used herein, a computer-readable storage medium should not be construed to be a transient signal per se, such as, for example, radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., optical pulses passing through an optical fiber cable), or electrical signals transmitted through a wire.
[0059] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded from an external computer or an external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each computing / processing device.
[0060] The computer-readable program code / instructions for performing the calculations may be in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in an object-oriented programming language such as Smalltalk, C++, and a procedural programming language such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), can execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to perform aspects or operations and personalize the electronic circuit.
[0061] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / steps specified in the flowchart and / or block(s) of the block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium having instructions stored therein includes a manufacture comprising instructions for implementing aspects of the functions / steps specified in the flowchart and / or block(s) of the block diagram.
[0062] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / steps specified in the flowchart and / or block(s) of the block diagram.
[0063] Flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in the flowchart or block diagram can represent a module, segment, or portion of instructions that includes one or more executable instructions for implementing the specified logical function(s). The method, computer system, and computer-readable media may include additional blocks, fewer blocks, different blocks, or differently arranged blocks compared to those shown in the drawings. In some alternative implementations, the functions described in the blocks may be performed out of the order described in the drawings. For example, two blocks shown sequentially may in fact be executed simultaneously or substantially simultaneously, or the blocks may sometimes be executed in the reverse order depending on the functions involved. Also, note that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a special-purpose hardware-based system that performs the specified function or acts, or combinations of special-purpose hardware and computer instructions.
[0064] It will be apparent that the systems and / or methods described herein can be implemented in various forms of hardware, firmware, or a combination of hardware and software. The actual specialized control hardware or software code used to implement these systems and / or methods is not limiting of the implementation. Thus, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, and it is understood that software and hardware can be designed based on the description herein to implement the systems and / or methods.
[0065] No element, step, or instruction used herein should be construed as essential or required unless expressly stated. Also, as used herein, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." Additionally, as used herein, the term "set" is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and may be used interchangeably with "one or more." When only one item is intended, the term "one" or similar language is used. Also, as used herein, the terms "have," "having," "having," and the like are intended to be open-ended terms. Additionally, the phrase "based on" is intended to mean "based at least in part on" unless expressly stated otherwise.
[0066] The description of various aspects and embodiments has been presented for illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Even if combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features may be combined in ways not specifically recited in the claims and / or disclosed in the specification. Although each dependent claim listed below may directly refer to only one claim, the disclosure of possible implementations includes each dependent claim combined with every other claim in the claims. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terms used in this specification are selected to best explain the principles of the embodiments, practical applications or technical improvements to the technology found in the market, or to enable other skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for video coding to be executed by a processor, comprising: receiving video data including a current picture; generating a virtual reference frame for the current picture using a neural network based on a hierarchical level related to the current picture and the closest decoded picture; decoding the video data based on a reference picture list including the generated virtual reference frame, wherein the neural network is applied to pictures coded using a quantization parameter (QP) smaller than a given threshold and not applied to other pictures, a method.
2. The method according to claim 1, wherein the virtual reference frame corresponds to one or more of an I-frame, a P-frame, and a B-frame.
3. The method according to claim 1 or 2, wherein the decoded video data is included in the reference picture list.
4. The method according to claim 3, wherein subsequent frames from the video data are decoded based on motion-compensated prediction, intra prediction, or intra block copy based on the reference picture list.
5. The method according to claim 4, wherein the virtual reference frame is generated based on one or more of signal processing, spatial or temporal filtering, scaling, weighted averaging, up / downsampling, pooling, recursive processing in memory, linear system processing, non-linear system processing, neural network processing, deep learning-based processing, AI processing, processing of a pre-trained network, machine learning-based processing, and online training network processing.
6. A computer system for video coding, the computer system comprising: One or more computer-readable non-transitory storage media configured to store computer program code; One or more computer processors configured to access the computer program code and operate as instructed by the computer program code, the computer program code causing the one or more computer processors to execute the method according to any one of claims 1 to 5; A computer system. Claim 7 A computer program for causing one or more computer processors to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method and apparatus using virtual reference pictures
JP2009543521A
Method and apparatus for inter-picture prediction using hypothetical reference pictures for video coding
JP2023504418A
Inter-prediction method and apparatus using reference frame generated based on deep learning
US20190306526A1
Method and device for encoding or decoding image
US20200120340A1