Method, computer system, and computer program for coding image data
By selecting a subset of hybrid transform kernels for video data decoding, the method addresses the computational and bitrate overhead issues in AV1 video coding, enhancing efficiency and performance.
Patent Information
- Application Number
- JP2024025597
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-09
- Filing Date
- 2024-02-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-02-02
AI Technical Summary
The existing AV1 video coding format incurs additional computational complexity and bitrate overhead due to the need to select and signal specific hybrid transform kernels for each residual coding block.
The proposed method involves identifying a set of hybrid transform kernels corresponding to the video data and selecting a subset of these kernels, either explicitly or implicitly, to decode the video data, thereby reducing computational complexity and bitrate overhead.
This approach improves coding efficiency by reducing computational complexity and bitrate overhead while maintaining or enhancing coding performance, thus optimizing video data processing in the AV1 format.
Smart Images

Figure 0007683868000002 
Figure 0007683868000003 
Figure 0007683868000004
Abstract
Description
Technical Field
[0001] This application claims priority from U.S. Provisional Patent Application No. 63 / 032,216, filed May 29, 2020, and U.S. Patent Application No. 17 / 066,791, filed Oct. 9, 2020, the entireties of which are hereby incorporated by reference.
[0002] This disclosure relates generally to the field of data processing, and more specifically to video encoding and decoding.
Background Art
[0003] AOMedia Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. It was developed as a successor to VP9 by the Alliance for Open Media (AOMedia), a consortium established in 2015 that includes semiconductor companies, video-on-demand providers, video content producers, software development companies, and web browser vendors. Many of the components of the AV1 project are based on the previous research efforts of the alliance members. Individual contributors had started experimental technology platforms years earlier, with Xiph / Mozilla's Daala already releasing code in 2010, Google's experimental VP9 evolution project VP10 announced on September 12, 2014, and Cisco's Thor announced on August 11, 2015. Based on the VP9 codebase, AV1 incorporates additional technologies, some of which were developed in these experimental formats. The first version 0.1.0 of the AV1 reference codec was released on April 7, 2016. The alliance released the AV1 bitstream specification, along with a reference software-based encoder and decoder, on March 28, 2018. On June 25, 2018, a verified version 1.0.0 of the specification was released. On January 8, 2019, a verified version 1.0.0 with the Errata 1 specification was released. The AV1 bitstream specification includes a reference video codec. Summary of the Invention
[0004] Embodiments relate to a method, a system, and a computer-readable medium for coding video data. According to one aspect, a method for coding video data is provided. The method may include receiving video data. A set of hybrid transform kernels corresponding to the video data is identified. A subset of the set of hybrid transform kernels is selected, either explicitly or implicitly, from the set of hybrid transform kernels. The video data is decoded based on the selected subset of hybrid transform kernels.
[0005] According to another aspect, a computer system for coding video data is provided. The computer system can include one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage devices, and program instructions stored in at least one of the one or more storage devices for execution by at least one of the one or more processors via at least one of the one or more memories, whereby the computer system can execute the method. The method may include receiving video data. A set of hybrid transform kernels corresponding to the video data is identified. A subset of the set of hybrid transform kernels is selected, either explicitly or implicitly, from the set of hybrid transform kernels. The video data is decoded based on the selected subset of hybrid transform kernels.
[0006] According to another aspect, a computer-readable medium for coding video data is provided. The computer-readable medium may include one or more computer-readable storage devices and program instructions executable by a processor stored in at least one of the one or more tangible storage devices. The program instructions are executable by the processor to perform a method that may include receiving video data accordingly. A set of hybrid transform kernels corresponding to the video data is identified. A subset of the hybrid transform kernels is selected, either explicitly or implicitly, from the set of hybrid transform kernels. The video data is decoded based on the selected subset of hybrid transform kernels.
Brief Description of the Drawings
[0007] These and other objects, features and advantages will become apparent from the following detailed description of exemplary embodiments read in conjunction with the accompanying drawings. These illustrations are for clarity in facilitating the understanding of those skilled in the art together with the detailed description, and the various features of the drawings are not to scale.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
[0008] Detailed embodiments of the claimed structures and methods are described herein, but it will be understood that the disclosed embodiments are merely illustrative of the claimed structures and methods that may be embodied in various forms. These structures and methods can be embodied in many different forms and should not be construed as limited to the exemplary embodiments described herein. Rather, these exemplary embodiments are provided so that this disclosure will be thorough and complete and will fully convey the scope to those skilled in the art. In the description, details of well-known mechanisms and techniques may be omitted so as not to unnecessarily obscure the presented embodiments.
[0009] Embodiments generally relate to the field of data processing, and more specifically to video encoding and decoding. The exemplary embodiments described below provide, among other things, a system, method, and computer program for encoding and decoding video data based on selecting a hybrid transform kernel, either implicitly or explicitly. Thus, some embodiments have the ability to improve the field of computing by enabling enhanced coding efficiency through the use of a hybrid transform kernel that may be implied by a computer from video data.
[0010] As described above, AOMedia Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. It was developed as a successor to VP9 by the Alliance for Open Media (AOMedia), a consortium established in 2015 that includes semiconductor companies, video-on-demand providers, video content producers, software development companies, and web browser vendors. Many of the components of the AV1 project are based on the previous research efforts of alliance members. Individual contributors had started experimental technology platforms years earlier, with Xiph / Mozilla's Daala already releasing code in 2010, Google's experimental VP9 evolution project VP10 announced on September 12, 2014, and Cisco's Thor announced on August 11, 2015. Based on the VP9 codebase, AV1 incorporates additional technologies, some of which were developed in these experimental formats. The first version 0.1.0 of the AV1 reference codec was released on April 7, 2016. The alliance released the AV1 bitstream specification on March 28, 2018, along with a reference software-based encoder and decoder. On June 25, 2018, the verified version 1.0.0 of the specification was released. On January 8, 2019, the verified version 1.0.0 with the Errata 1 specification was released. The AV1 bitstream specification includes a reference video codec.
[0011] Unlike VP9 where each coding block has only one type of transformation, AV1 enables each transformation block to independently select its own transformation kernel. AV1 utilizes a set of hybrid transformation kernels to code the intra prediction residual. Hybrid transformation kernels generally refer to 2D separable transformation kernels that are extended to combinations of various 1D kernels, such as DCT, ADST, inverse ADST (FLIPADDST), and identity transformation (IDTX). The set of hybrid transformation kernels and their availability for luma intra prediction residuals depend on the size of the residual block. For chroma intra prediction residuals, the transformation type selection is done implicitly depending on the intra prediction mode. However, with the introduction of LGT (and its inverse version) and KLT in the AV2 development process, the set of hybrid transformation kernels available for coding luma and chroma intra prediction residuals has been extended. Selecting specific hybrid transformation types from this extended set and signaling them in the bitstream for each residual coding block incurs additional computational complexity and bitrate overhead. Therefore, it may be advantageous to use LGT and KLT that are intra-mode dependent and residual block size dependent to utilize the directionality of changes in the magnitude of the residuals in order to reduce both the computational complexity and the bitrate overhead while improving coding performance.
[0012] Here, aspects will be described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer-readable media according to various embodiments. It is understood that each block of the flowcharts and / or block diagrams, as well as combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0013] Next, refer to FIG. 1, which is a functional block diagram of a networked computer environment showing a video coding system 100 (hereinafter, "the system") for encoding and decoding video data based on selecting an implicit or explicit hybrid transform kernel. It should be understood that FIG. 1 merely provides an example of one implementation and does not imply any limitation regarding an environment in which multiple different embodiments may be implemented. Numerous changes to the illustrated environment can be made based on design and implementation requirements.
[0014] System 100 may include a computer 102 and a server computer 114. Computer 102 may communicate with server computer 114 via a communication network 110 (hereinafter, "the network"). Computer 102 can include a processor 104 and a software program 108 stored in a data storage device 106, and is enabled to interface with a user and communicate with server computer 114. As will be described later with reference to FIG. 4, computer 102 can include internal components 800A and external components 900A respectively, and server computer 114 can include internal components 800B and external components 900B respectively. Computer 102 can be, for example, a mobile device, a phone, a personal digital assistant, a netbook, a laptop computer, a tablet computer, a desktop computer, or any type of computing device capable of running a program, accessing a network, and accessing a database.
[0015] Server computer 114 may also operate in a cloud computing service model, such as, for example, software - as - a - service (SaaS), platform - as - a - service (PaaS), or infrastructure - as - a - service (IaaS), as will be described later with respect to FIGS. 5 and 6. Server computer 114 may also be placed within a cloud computing deployment model, such as, for example, a private cloud, a community cloud, a public cloud, or a hybrid cloud.
[0016] Server computer 114, which can be used to code video data based on implicitly or explicitly selecting a hybrid conversion kernel, is enabled to execute a video coding program 116 (hereinafter referred to as "the program") that can interact with database 112. The video coding program method will be described in more detail later with respect to FIG. 3. In one embodiment, computer 102 can operate as an input device including a user interface, while the program 116 can run mainly on server computer 114. In an alternative embodiment, the program 116 may run mainly on one or more computers 102, and server computer 114 may be used for processing and storing data used by the program 116. Note that the program 116 may be a stand - alone program or may be integrated into a larger video coding program.
[0017] However, it should be noted that, in some examples, the processing for program 116 may be shared between computer 102 and server computer 114 in any ratio. In another embodiment, program 116 may operate on two or more computers, server computers, or some combination of a computer and a server computer, for example, multiple computers 102 that communicate with a single server computer 114 across network 110. In another embodiment, for example, program 116 may operate on multiple server computers 114 that communicate with multiple client computers across network 110. Alternatively, the program may operate on a network server that communicates with a server and multiple client computers across the network.
[0018] Network 110 may include a wired connection, a wireless connection, an optical fiber connection, or some combination thereof. Generally, network 110 can be any combination of connections and protocols that support communication between computer 102 and server computer 114. Network 110 can be, for example, a local area network (LAN), a wide area network (WAN) such as the Internet, a telecommunications network such as the public switched telephone network (PSTN), a wireless network, a public switched network, a satellite network, a cellular network (e.g., a fifth generation (5G) network, a long term evolution (LTE) network, a third generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a metropolitan area network (MAN), a private network, an ad hoc network, an intranet, an optical fiber-based network, or the like, and / or a combination of these or other types of networks.
[0019] The number and configuration of the devices and networks shown in FIG. 1 are provided as an example. In practice, there may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks configured differently than those shown in FIG. 1. Further, two or more of the devices shown in FIG. 1 may be implemented within a single device, or alternatively, a single device shown in FIG. 1 may be implemented as a plurality of distributed devices. Additionally, or alternatively, a set of devices (e.g., one or more devices) of system 100 may perform one or more functions described as being performed by another set of devices of system 100.
[0020] Next, referring to FIG. 2, an exemplary line graph transform (LGT) 200 is shown. A graph can be a general mathematical structure composed of a set of vertices and edges used to model affinity relationships between objects of interest. In practice, a weighted graph (for which a set of weights is assigned to edges and optionally vertices) can provide a sparse representation for robust modeling of signals / data. The LGT can improve coding efficiency by providing better adaptation to various block statistics. A separable LGT can be designed and optimized by learning a line graph from data to model the inherent row and column statistics of the block residual signal, and a related generalized graph Laplacian (GGL) matrix is used to derive the LGT.
[0021] For example, given a weighted graph G(W,V), the GGL matrix can be defined as L c = D - W + V. Here, W can be an adjacency matrix composed of non - negative edge weights wc, D can be a diagonal degree matrix, and V can be a diagonal matrix showing weighted self - loops v c1 , v c2 . The matrix L c is:
Number
[0022] And the LGT can be derived by the eigenvalue decomposition of GGL Lc:UΦU T where the columns of the orthogonal matrix U are the basis vectors of the LGT and Φ is a diagonal eigenvalue matrix. In fact, the discrete cosine transform (DCT) and discrete sine transform (DST), including DCT-2, DCT-8, and DST-7, are LGTs derived from a specific form of GGL. DCT-2 is derived by setting v c1 = 0. DST-7 is derived by setting v c1 = w c . DCT-8 is derived by setting v c2 = w c . DST-4 is derived by setting v c1 = 2w c . DCT-4 is derived by setting v c2 = 2w c .
[0023] The LGT is implemented using matrix multiplication for transform sizes 4, 8, and 16. The 4-point LGT core is derived by setting v c at L c1 = 2w c , which means it is DST-4. The 8-point LGT core is derived by setting v c at L c1 = 1.5w c , and the 16-point LGT core is derived by setting v c at L c1 = w c , which means it is DST-7.
[0024] An extended set of hybrid transform kernels can be referred to as set A. Set A comprehensively includes all combinations such as the Discrete Cosine Transform (DCT), Identity Transform (IDTX, which skips transform coding in a specific direction), Asymmetric Discrete Sine Transform (ADST), Flipped Asymmetric Discrete Sine Transform (FLIPADST, which applies ADST in reverse order), Line Graph Transform (LGT), Flipped Line Graph Transform (FLIPLGT), Karhunen - Loeve Transform (KLT), etc. A subset of the elements of A that can be a reduced set of transform types can be referred to as x, and thus, x ∈ A. The subset x can include one or more transform types (e.g., DCT, ADST, LGT, KLT) and / or one or more combinations of vertical and horizontal transform types (e.g., DCT_DCT, LGT_LGT, DCT_LGT, LGT_DCT).
[0025] According to one or more embodiments, an implicit method can be used to select the elements of x such that a hybrid transform type can be selected based on the coded information available to both the encoder and the decoder. Thus, it can be said that no additional signaling is required to specify the transform type in the decoder. In one embodiment, this selection can be made depending on the intra prediction mode and / or the block size. In one embodiment, one or more of eight nominal modes, five non - angular smoothing modes, and angular delta values (e.g., from - 3 to + 3) can be considered during the selection process. In one embodiment, for the directional intra prediction mode, only the nominal mode may be used to select the transform type (i.e., directional intra prediction modes that have different angular delta values but share the same nominal mode can apply the same implicit transform type).
[0026] In one embodiment, the same hybrid conversion type is selected in the recursive filtering mode and the DC mode. In one embodiment, the recursive filtering mode and the SMOOTH mode select the same hybrid conversion type. In one embodiment, the SMOOTH, SMOOTH_H, and SMOOTH_V modes select the same hybrid conversion type. In one embodiment, the SMOOTH, SMOOTH_H, SMOOTH_V, and Paeth prediction modes select the same hybrid conversion type. In one embodiment, the recursive filtering mode, SMOOTH, and Paeth prediction mode select the same hybrid conversion type. In one embodiment, the vertical mode (V_PRED) and the SMOOTH_V prediction mode select the same hybrid conversion type. In one embodiment, the horizontal mode (H_PRED) and the SMOOTH_H prediction mode select the same hybrid conversion type. In one embodiment, the CfL mode and the DC mode select the same hybrid conversion type. In one embodiment, the CfL mode and the SMOOTH mode select the same hybrid conversion type. In one embodiment, the CfL mode and the Paeth mode select the same hybrid conversion type.
[0027] In one embodiment, depending on one or more of eight nominal modes, five non-angular smoothing modes, and an angular delta value (e.g., from -3 to +3), and a block size, different self-loop weights (v c1 v c2 ) of the LGT can be used. In one embodiment, depending on one or more of eight nominal modes, five non-angular smoothing, and an angular delta value (e.g., from -3 to +3), and a block size, different statistical characteristics of the KLT can be used. In one embodiment, for the same intra prediction mode that can be enabled for both the luma and chroma components, the implicit hybrid conversion selection can be the same.
[0028] According to one or more embodiments, an explicit method for selecting elements of x may be proposed such that the selection needs to be specified by a syntax signaled within a bitstream (i.e., the encoder needs to explicitly select and signal the conversion type at the block level). The block level may include a super block level, a coding block level, a prediction block level, or a conversion block level. In one embodiment, an explicit conversion method (at least two conversion type candidates) can be applied to all intra prediction modes, but the number of hybrid conversion candidates can be different for different intra prediction modes. In another embodiment, in some intra prediction modes, either an implicit or an explicit conversion method can be used, while other intra prediction modes apply an implicit conversion method (only one available conversion type). In one embodiment, when the explicit conversion method involves the use of LGT, an identifier of the self-loop weight defining the LGT candidate can be signaled in the bitstream at the block level, and the identifier can be either an index of the associated self-loop rate value or the self-loop rate value. In one embodiment, when the explicit conversion method involves the use of KLT, an identifier of the KLT kernel can be signaled in the bitstream at the block level, and the identifier can be either an index of the KLT or a KLT matrix element value.
[0029] The switch between the explicit method and the implicit method can be indicated either at a high-level syntax or at the block level. If the selection can be indicated at the HLS, it may include a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a slice header. If the switch can be indicated at the block level, it may include a super block level, a coding block level, a prediction block level, and / or a conversion block level.
[0030] Next, referring to FIG. 3, an operation flowchart showing the steps of a method 300 for coding video data is shown. In some implementations, one or more of the process blocks in FIG. 3 may be executed by computer 102 (FIG. 1) and server computer 114 (FIG. 1). In some implementations, one or more of the process blocks in FIG. 3 may be executed by another device or group of devices separate from or including computer 102 and server computer 114.
[0031] At 302, method 300 includes receiving video data.
[0032] At 304, method 300 includes identifying a set of hybrid transform kernels corresponding to the video data.
[0033] At 306, method 300 includes selecting a subset of the hybrid transform kernels, either explicitly or implicitly, from the set of hybrid transform kernels.
[0034] At 308, method 300 includes decoding the video data based on the selected subset of hybrid transform kernels.
[0035] It should be understood that FIG. 3 merely provides an illustration of one implementation and does not imply any limitation as to how different embodiments may be implemented. Numerous changes to the illustrated environment may be made based on design and implementation requirements.
[0036] FIG. 4 is a block diagram 400 of the internal and external components of the computer shown in FIG. 1 according to an exemplary embodiment. It should be understood that FIG. 4 merely provides an illustration of one implementation and does not imply any limitation as to the environment in which different embodiments may be implemented. Numerous changes to the illustrated environment may be made based on design and implementation requirements.
[0037] Computer 102 (FIG. 1) and server computer 114 (FIG. 1) may include respective sets 800A, B of internal components and sets 900A, B of external components shown in FIG. 4. Each of the sets 800 of internal components includes one or more processors 820, one or more computer-readable RAMs 822, and one or more computer-readable ROMs 824 on one or more buses 826, one or more operating systems 828, and one or more computer-readable tangible storage devices 830.
[0038] Processor 820 is implemented in hardware, firmware, or a combination of hardware and software. Processor 820 is a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or another type of processing component. In some implementations, processor 820 includes one or more processors that can be programmed to execute functions. Bus 826 includes components that enable communication between internal components 800A, B.
[0039] One or more operating systems 828, software programs 108 (FIG. 1), and video coding programs 116 (FIG. 1) on the server computer 114 (FIG. 1) are stored in one or more of respective computer-readable tangible storage devices 830 for execution by one or more of respective processors 820 via one or more of respective RAMs 822 (typically including cache memory). In the embodiment shown in FIG. 4, each of the computer-readable tangible storage devices 830 is a magnetic disk storage device of an internal hard drive. Alternatively, each of the computer-readable tangible storage devices 830 is a semiconductor storage device such as, for example, ROM 824, EPROM, flash memory, optical disk, magneto-optical disk, solid state disk, compact disk (CD), digital versatile disk (DVD), floppy disk (registered trademark), cartridge, magnetic tape, and / or another type of non-transitory computer-readable tangible storage device capable of storing computer programs and digital information.
[0040] Each set 800A, B of internal components also includes an R / W drive or interface 832 for reading from and writing to one or more portable computer-readable tangible storage devices 936 such as, for example, CD-ROM, DVD, memory stick, magnetic tape, magnetic disk, optical disk, or semiconductor storage device. Software programs such as, for example, software program 108 (FIG. 1) and video coding program 116 (FIG. 1) are stored in one or more of respective portable computer-readable tangible storage devices 936, read out via respective R / W drives or interfaces 832, and can be loaded into respective hard drives 830.
[0041] Each set 800A, B of internal components also includes a network adapter or interface 836, such as a TCP / IP adapter card, a wireless Wi-Fi interface card, or a 3G, 4G, or 5G wireless interface card, or other wired or wireless communication links. The software program 108 (FIG. 1) and the video coding program 116 (FIG. 1) on the server computer 114 (FIG. 1) can be downloaded from an external computer to the computer 102 (FIG. 1) and the server computer 114 via a network (e.g., the Internet, a local area network, or other wide area network) and their respective network adapters or interfaces 836. From the network adapter or interface 836, the software program 108 and the video coding program 116 on the server computer 114 are loaded onto their respective hard drives 830. The network can have copper wires, optical fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers.
[0042] Each of the sets 900A, B of external components can include a computer display monitor 920, a keyboard 930, and a computer mouse 934. The external components 900A, B can also include a touch screen, a virtual keyboard, a touch pad, a pointing device, and other human interface devices. Each of the sets 800A, B of internal components also includes a device driver 840 for interfacing with the computer display monitor 920, the keyboard 930, and the computer mouse 934. The device driver 840, the R / W drive or interface 832, and the network adapter or interface 836 have hardware and software (stored in the storage device 830 and / or the ROM 824).
[0043] It should be understood in advance that this disclosure includes a detailed description of cloud computing, but the implementation of the teachings described herein is not limited to a cloud computing environment. Rather, some embodiments can be implemented with any other type of computing environment, whether currently known or later developed.
[0044] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (such as networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services), which can be rapidly provisioned and released with minimal management effort or interaction with the service provider. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.
[0045] The characteristics are as follows: On-demand self-service: Cloud consumers can automatically provision computing capabilities, such as server time and network storage, as needed, without the need for human interaction with the service provider. Broad network access: The capabilities are available over the network and accessed through standard mechanisms that promote use by heterogeneous thin or thick client platforms (such as mobile phones, laptops, and PDAs). Resource pooling: The provider's computing resources are pooled to serve multiple users using a multi-tenant model, where different physical and virtual resources are dynamically assigned and reassigned according to demand. Users can generally specify their location at a higher level of abstraction (e.g., country, state, or data center), giving a sense of location independence in that they have no control or knowledge of the exact location of the resources provided; Rapid adaptability: Functions can be made quickly and elastically available, in some cases automatically, to scale out rapidly and then be released quickly to scale in. To the user, the capabilities available for provisioning often appear limitless, and any amount can be purchased at any time; Measured service: The cloud system automatically controls and optimizes resource utilization by leveraging metering capabilities at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). It can monitor, control, and report resource usage, providing transparency to both the provider and the user of the services being utilized.
[0046] The service model is as follows: Software as a Service (SaaS): The functionality provided to the user is to use the provider's applications running on the cloud infrastructure. The applications are accessible from various client devices via a thin-client interface such as a web browser (e.g., web-based email). The user does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application functionality, except for limited user-specific application configuration settings; Platform as a Service (PaaS): The functionality provided to users is to deploy applications created or obtained by users, which are created using programming languages and tools supported by the provider, onto the cloud infrastructure. Users have control over the deployed applications and, optionally, the application hosting environment settings, but do not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage; Infrastructure as a Service (IaaS): The functionality provided to users is to enable users to deploy and run any software that may include operating systems and applications on processing resources, storage resources, network resources, and other basic computing resources. Users have control over the operating system, storage, deployed applications, and, optionally, limited control over selecting network components (e.g., host firewalls), but do not manage or control the underlying cloud infrastructure.
[0047] The deployment models are as follows: Private cloud: The cloud infrastructure is operated solely for an organization. This can be managed by the organization or a third party and can be located on - premise or off - premise; Community cloud: The cloud infrastructure is shared by several organizations to support a specific community with common concerns (e.g., mission, security requirements, policies, and compliance considerations). This can be managed by those organizations or a third party and can be located on - premise or off - premise; Public cloud: The cloud infrastructure is made available to the general public or large industry groups and is owned by an organization that sells cloud services; Hybrid Cloud: A cloud infrastructure that combines two or more clouds (private, community, or public), where the two or more clouds remain distinct entities but are joined together by standardized or proprietary technologies that enable data and application portability (such as cloud bursting for load balancing between clouds).
[0048] A cloud computing environment is service-oriented, focusing on statelessness, loose coupling, modularity, and semantic interoperability. At the center of cloud computing is an infrastructure with a network of interconnected nodes.
[0049] Referring to FIG. 5, an exemplary cloud computing environment 500 is shown. As illustrated, cloud computing environment 500 has one or more cloud computing nodes 10, which are used to communicate with local computing devices used by cloud consumers such as, for example, a personal digital assistant (PDA) or cellular phone 54A, desktop computer 54B, laptop computer 54C, and / or automotive computer system 54N. The cloud computing nodes 10 can communicate with one another. They may be physically or virtually grouped (not shown) in one or more networks such as, for example, the private, community, public, or hybrid clouds described above herein, or a combination thereof. This enables the cloud computing environment 500 to provide infrastructure, platforms, and / or software as services such that a cloud consumer need not maintain resources on a local computing device. It should be understood that the types of computing devices 54A-54N shown in FIG. 4 are merely illustrative, and that the cloud computing nodes 10 and the cloud computing environment 500 can communicate with any type of computerized device on any type of network and / or network addressable connection (e.g., using a web browser).
[0050] Referring to FIG. 6, a set 600 of functional abstractions provided by cloud computing environment 500 (FIG. 5) is shown. It should be understood in advance that the components, layers, and functions shown in FIG. 6 are merely illustrative and are not intended to limit the embodiments. As illustrated, the following layers and corresponding functions are provided.
[0051] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframe 61, servers 62 based on RISC (Reduced Instruction Set Computer) architecture, server 63, blade server 64, storage device 65, and network and networking components 66. In some embodiments, software components include network application server software 67 and database software 68.
[0052] The virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual server 71, virtual storage 72, virtual network 73 including a virtual private network, virtual applications and operating systems 74, and virtual client 75.
[0053] In one example, the management layer 80 can provide the functions described below. Resource provisioning 81 provides for the dynamic procurement of computing resources and other resources used to execute tasks within a cloud computing environment. Metering and pricing 82 provides for cost tracking when resources are utilized within a cloud computing environment and for billing or invoicing for the consumption of these resources. In one example, these resources may have application software licenses. Security provides for the authentication of cloud users and tasks and for the protection of data and other resources. User portal 83 provides access to the cloud computing environment for users and system administrators. Service level management 84 provides for the allocation and management of cloud computing resources such that the required service levels are met. Service level agreement (SLA) formulation and fulfillment 85 provides for the pre - preparation and procurement of cloud computing resources for which future demands are predicted, in accordance with the SLA.
[0054] The workload layer 90 provides examples of functions for which a cloud computing environment can be utilized. Examples of workloads and functions that can be provided from this layer include mapping and navigation 91, software development and lifecycle management 92, virtual classroom lesson delivery 93, data analysis processing 94, transaction processing 95, and video coding 96. Video coding 96 can encode and decode video data based on implicitly or explicitly selecting a hybrid conversion kernel.
[0055] Some embodiments may relate to systems, methods, and / or computer-readable media at any possible technical detail level of integration. The computer-readable media may include a computer-readable non-transitory storage medium (or media) having computer-readable program instructions for causing a processor to execute operations.
[0056] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVDs), memory sticks, floppy disks (registered trademark), mechanically encoded devices such as punch cards or raised structures in grooves that record instructions, and any suitable combination thereof. A computer-readable storage medium, as used herein, is not to be construed as being a transient signal per se, such as, for example, radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., optical pulses passing through an optical fiber cable), or electrical signals transmitted through a wire.
[0057] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded via a network, such as, for example, the Internet, a local area network, a wide area network, and / or a wireless network, from an external computer or an external storage device. The network can include a copper transmission cable, an optical transmission fiber, wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium within the respective computing / processing device.
[0058] The computer-readable program code / instructions for performing the operations can be in any combination of one or more programming languages, including, for example, assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in an object-oriented programming language such as Smalltalk, C++, or the like, and a procedural programming language such as the "C" programming language or a similar programming language. The computer-readable program instructions can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit for performing the aspects or operations.
[0059] These computer-readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which are executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium storing the instructions comprises an article of manufacture including instructions for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0060] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0061] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in the flowchart or block diagram may represent one or more executable instructions of a module, segment, or portion of code that implements a particular (one or more) logical function. The method, computer system, and computer-readable media may include additional blocks, fewer blocks, different blocks, or blocks configured differently than those shown in the figures. In some alternative implementations, the functions recited in the blocks may be performed out of the order recited in the figures. For example, two blocks shown in succession may actually be executed concurrently or substantially concurrently, or may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or combinations of dedicated hardware and computer instructions.
[0062] It should be understood that the systems and / or methods described herein can be implemented in various forms of hardware, firmware, or combinations of hardware and software. The actual specialized control hardware or software code used to implement these systems and / or methods is not limiting. Accordingly, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, and it should be understood that software and hardware can be designed based on the description herein to implement the systems and / or methods.
[0063] None of the elements, acts, or instructions used herein are expressly so recited Unless otherwise indicated, it should not be construed as important or essential. Also, as used herein, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." Further, as used herein, the term "set" is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and may be used interchangeably with "one or more." When only one item is intended, the term "one" or similar words are used. Also, as used herein, the terms "comprising," "having," "including," or the like are intended to be open-ended terms. Further, the phrase "based on" is intended to mean "at least partially based on" unless expressly stated otherwise.
[0064] These descriptions of various aspects and embodiments are presented for purposes of illustration and are not intended to be exhaustive or limited to the disclosed embodiments. Even if combinations of features are recited in the claims and / or disclosed in the specification, such combinations are not intended to limit the disclosed implementations. In fact, many of these features may be combined in ways not specifically recited in the claims and / or specifically disclosed in the specification. Each of the dependent claims listed below may directly depend on only one claim, but the disclosed implementations may include each dependent claim in combination with all other claims in the claim set. Numerous changes and modifications will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terms used herein are chosen to best explain the principles of the embodiments, the practical application, or a technical improvement found in the marketplace, or to enable one skilled in the art to understand the embodiments disclosed herein.
Claims
1. 1. A processor-implemented method for video decoding, comprising: receiving video data; identifying a plurality of transform kernels corresponding to the video data, the plurality of transform kernels having a plurality of pairs of transform types including a vertical transform type and a horizontal transform type, at least one of which is a line graph transform (LGT); selecting one or more transformation kernels from the plurality of transformation kernels; decoding the video data based on the one or more selected transformation kernels; having Selecting the one or more transformation kernels comprises: In the recursive filtering mode and the DC mode, the same transformation type is selected; In the recursive filtering mode and the SMOOTH mode, the same transformation type is selected; In the SMOOTH mode, the SMOOTH_H mode, and the SMOOTH_V mode, the same conversion type is selected; In the SMOOTH mode, the SMOOTH_H mode, the SMOOTH_V mode, and the Paeth mode, the same conversion type is selected; In the recursive filtering mode, the SMOOTH mode, and the Paeth mode, the same transformation type is selected; In V_PRED mode and SMOOTH_V mode, the same conversion type is selected; In the H_PRED and SMOOTH_H modes, the same conversion type is selected; In chroma-from-luma and DC modes, the same conversion type must be selected, or In chroma-from-luma mode and SMOOTH mode, select the same conversion type. having at least one of: method.
2. The method of claim 1 , wherein the one or more transform kernels are selected based on coded information available to both an encoder and a decoder.
3. The method of claim 1 , wherein the one or more transform kernels are selected based on both an intra-prediction mode and a block size associated with the video data.
4. The method of claim 1 , wherein the one or more transform kernels are explicitly specified by syntax elements signaled in a bitstream associated with the video data.
5. The method of claim 1 , in which an explicit transformation scheme is applied for all intra-prediction modes.
6. The method of claim 5 , wherein different intra-prediction modes have different numbers of transform candidates.
7. The method of claim 1 , in which an explicit transformation scheme is used for a subset of intra-prediction modes.
8. The method of claim 1 , wherein the one or more transform kernels are switched between explicit and implicit based on signaling in a high-level syntax or at a block level.
9. one or more memories storing a computer program; one or more processors; having The computer program causes the one or more processors to carry out a method according to any one of claims 1 to 8. Computer system.
10. A computer program causing a computer to carry out the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Implementation design for hybrid transform coding scheme
US20150145874A1
Method and apparatus for processing video signal using graph-based transform
US20180041760A1
Method and device for processing a video signal by using an adaptive separable graph-based transform
US20180146195A1
Method and device for processing video signal by using separable graph-based transform
US20180213233A1
Transform Kernel Selection and Entropy Coding
US20180249179A1