Video encoding method and device, video decoding method and device, computer system and computer readable medium
By identifying the luma block context at predefined locations in AV1 video encoding, the signaling notification of the chroma intra-prediction mode is improved, solving the problem of low signaling efficiency in the chroma intra-prediction mode and improving the efficiency of video encoding and decoding.
Patent Information
- Application Number
- CN202511161933.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-02
- Filing Date
- 2020-11-09
- Publication Date
- 2025-11-14
AI Technical Summary
In existing AV1 video encoding and decoding technologies, the signaling notification efficiency of the chroma intra-frame prediction mode is not high, resulting in insufficient video encoding and decoding efficiency.
The signaling notification process is improved by identifying entropy coding context for intra-chroma prediction modes by recognizing co-located luminance blocks at one or more predefined locations.
It improves the efficiency of video encoding and decoding by optimizing the signaling notification of the chroma intra-frame prediction mode, thereby enhancing the performance of encoding and decoding.
Smart Images

Figure CN120956901A_ABST
Abstract
Description
[0001] This application is a divisional application, with original application number 2020800632912 and original application date of November 9, 2020. The entire contents of the original application are incorporated herein by reference. Cross-reference to related applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 026,495, filed May 18, 2020, and U.S. Patent Application No. 17 / 061,854, filed October 2, 2020, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This invention relates to the field of data processing, and more specifically to video encoding and / or decoding. Specifically, this invention relates to a video decoding method and apparatus, a computer system, and a computer-readable medium. Background Technology
[0004] AOMedia Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. AOMedia Video 1 was developed by the Alliance for Open Media (AOMedia) as a successor to VP9. The Alliance, founded in 2015, includes semiconductor companies, video-on-demand providers, video content producers, software development companies, and web browser vendors. Many components of the AV1 project are derived from previous research by Alliance members. Individual contributors launched experimental technology platforms several years ago: Xiph / Mozilla's Daala released its code in 2010, Google's experimental VP9 evolution project VP10 was announced on September 12, 2014, and Cisco's Thor released it on August 11, 2015. Building upon the VP9 codebase, AV1 incorporates other technologies, some of which were developed using these experimental formats. The first version 0.1.0 of the AV1 reference codec was released on April 7, 2016. The consortium announced the release of the AV1 bitstream specification, along with a software-based reference encoder and decoder, on March 28, 2018. Verification version 1.0.0 of the specification was released on June 25, 2018. Verification version 1.0.0 with errata table 1 was released on January 8, 2019. The AV1 bitstream specification includes a reference video codec. In existing AV1 technology, video encoding and decoding efficiency is not high. Summary of the Invention
[0005] Implementations relate to methods, systems, and computer-readable media for encoding and / or decoding video data. According to one aspect, a method for encoding and / or decoding video data is provided. The method may include receiving video data comprising chroma and luma components; identifying one or more contexts for entropy coding of intra-chroma prediction modes based on luma blocks co-located at one or more predefined locations; and decoding the video data based on the identified context.
[0006] According to another aspect, a computer system for encoding and / or decoding video data is provided. The computer system may include one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage devices, and program instructions stored on at least one of the one or more storage devices, the program instructions being executable by at least one of the one or more processors via at least one of the one or more memories, thereby enabling the computer system to perform a method. The method may include receiving video data comprising chroma components and luma components; identifying one or more contexts for entropy coding of intra-chroma prediction modes based on luma blocks co-located at one or more predefined locations; and decoding the video data based on the identified context.
[0007] According to another aspect, an apparatus for encoding and / or decoding video data is provided. The apparatus includes a receiving module for receiving video data comprising chroma and luma components; an identification module for identifying one or more contexts for entropy coding of intra-chroma prediction modes based on luma blocks co-located at one or more predefined locations; and a decoding module for decoding the video data based on the identified contexts.
[0008] According to another aspect, a computer-readable medium is provided for encoding and / or decoding video data. The computer-readable medium may include one or more computer-readable storage devices and program instructions stored on at least one of the one or more tangible storage devices, the program instructions being executable by a processor. The program instructions, executable by a processor, are used to perform a method that may accordingly include receiving video data comprising chroma and luma components; identifying one or more contexts for entropy coding of intra-chroma prediction modes based on luma blocks co-located at one or more predefined locations; and decoding the video data based on the identified context.
[0009] This invention provides a method, apparatus, computer system, and computer-readable medium for encoding and / or decoding video data. The method includes receiving video data comprising chroma and luma components; identifying one or more contexts for entropy coding of chroma intra-prediction modes based on luma blocks co-located at one or more predefined locations; and decoding the video data based on the identified contexts. By identifying one or more contexts for entropy coding of chroma intra-prediction modes based on luma blocks co-located at one or more predefined locations, signaling notification for chroma intra-prediction modes is improved, thereby increasing the efficiency of video encoding and decoding. Attached Figure Description
[0010] These and other objects, features, and advantages will become apparent from the following detailed description of illustrative embodiments, which should be read in conjunction with the accompanying drawings. The various features in the drawings are not to scale because they are illustrated for clarity in order to facilitate understanding by those skilled in the art in conjunction with the detailed description. In the drawings:
[0011] Figure 1 A networked computer environment according to at least one embodiment is illustrated;
[0012] Figure 2 It is a diagram of the coding tree structure of the luminance and chrominance components of video data according to at least one embodiment.
[0013] Figure 3 This is an operation flowchart illustrating the steps performed by a program that encodes video data according to at least one embodiment;
[0014] Figure 4 It is based on at least one embodiment. Figure 1 A block diagram depicting the internal and external components of a computer and server;
[0015] Figure 5 It includes, according to at least one embodiment. Figure 1 A block diagram illustrating a cloud computing environment for a computer system; and
[0016] Figure 6 It is based on at least one embodiment. Figure 5 A block diagram illustrating the functional layers of an illustrative cloud computing environment. Detailed Implementation
[0017] Specific embodiments of the claimed structures and methods are disclosed herein; however, it is to be understood that the disclosed embodiments are merely illustrative of the claimed structures and methods that can be implemented in various forms. These structures and methods can be implemented in many different forms and should not be construed as limited to the exemplary embodiments set forth herein. Rather, these exemplary embodiments are provided so that this disclosure will be thorough and complete and will fully convey the scope to those skilled in the art. Details of well-known features and techniques may be omitted in the specification to avoid unnecessarily obscuring the presented embodiments.
[0018] The implementations generally relate to the field of data processing, and more specifically to video encoding and decoding. The exemplary implementations described below provide systems, methods, and computer programs for encoding and / or decoding video data, particularly based on context associated with luma blocks co-located at one or more predefined locations. Thus, some implementations have the ability to improve computational efficiency by using improved signaling for intra-chroma prediction modes.
[0019] As previously described, AOMedia Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. AOMedia Video 1 was developed by the Open Media Consortium (AOMedia) as the successor to VP9. This consortium, founded in 2015, includes semiconductor companies, video-on-demand providers, video content producers, software development companies, and web browser vendors. Many components of the AV1 project are derived from previous research work by consortium members. Individual contributors launched experimental technology platforms several years ago: Xiph / Mozilla's Daala released its code in 2010, Google's experimental VP9 evolution project VP10 was announced on September 12, 2014, and Cisco's Thor released it on August 11, 2015. Building upon the VP9 codebase, AV1 incorporates other technologies, some of which were developed using these experimental formats. The first version 0.1.0 of the AV1 reference codec was released on April 7, 2016. The consortium announced the release of the AV1 bitstream specification, along with software-based reference encoders and decoders, on March 28, 2018. Verification version 1.0.0 of the specification was released on June 25, 2018. Verification version 1.0.0 with errata table 1 was released on January 8, 2019. The AV1 bitstream specification includes a reference video codec.
[0020] In AV1, semi-decoupled partitioning (SDP) can be used. However, in SDP, the luma and chroma blocks within a superblock can have different partitions, and the area of a chroma block may cover multiple luma coding blocks. Therefore, always using the top-left corner of the chroma block to locate the corresponding luma mode may not be optimal. Furthermore, when the luma and chroma blocks within a superblock have different partitions, the CfL mode is more likely to be selected as the optimal mode, but this characteristic is not well utilized in chroma mode signaling methods. Additionally, when signaling the chroma intra-prediction mode, all luma modes within the current superblock are available. However, this is not used to design better codewords for signaling the chroma intra-prediction mode. Therefore, to improve the signaling notification of the chroma intra-prediction mode, it may be advantageous to identify one or more contexts for entropy coding of the chroma intra-prediction mode based on luma blocks co-located at one or more predefined locations.
[0021] This document describes aspects with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer-readable media according to various embodiments. It will be understood that each block in the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0022] Now refer to Figure 1 This diagram illustrates a networked computer environment for a video coding system 100 (hereinafter referred to as the "System") that encodes and / or decodes video data based on a coding tree structure type. It should be understood that... Figure 1 This is merely an illustration of one implementation method and does not imply any limitation on the environments in which different implementation methods can be implemented. Many modifications can be made to the described environment based on design and implementation requirements.
[0023] System 100 may include computer 102 and server computer 114. Computer 102 may communicate with server computer 114 via communication network 110 (hereinafter referred to as "network"). Computer 102 may include processor 104 and software program 108, which is stored on data storage device 106 and is capable of connecting to a user interface and communicating with server computer 114. (Referencing below...) Figure 4 The computer 102 discussed may include internal components 800A and external components 900A, and the server computer 114 may include internal components 800B and external components 900B. The computer 102 may be, for example, a mobile device, telephone, personal digital assistant, netbook, laptop computer, tablet computer, desktop computer, or any type of computing device capable of running programs, accessing networks, and accessing databases.
[0024] As follows about Figure 6 The server computer 114 discussed can also run in a cloud computing service model, such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS). The server computer 114 can also reside in a cloud computing deployment model, such as a private cloud, community cloud, public cloud, or hybrid cloud.
[0025] A server computer 114, which can be used to encode video data, is capable of running a video encoding program 116 (hereinafter referred to as the "program") that can interact with a database 112. See below for reference. Figure 3 The video encoding program method will be described in more detail below. In one embodiment, computer 102 may operate as an input device including a user interface, while program 116 may run primarily on server computer 114. In an alternative embodiment, program 116 may run primarily on one or more computers 102, while server computer 114 may be used to process and store data used by program 116. It should be noted that program 116 may be a standalone program or may be integrated into a larger video encoding program.
[0026] However, it should be noted that in some instances, the processing of program 116 can be shared between computer 102 and server computer 114 at any ratio. In another embodiment, program 116 can run on more than one computer, server computer, or some combination of computers and server computers; for example, multiple computers 102 communicate with a single server computer 114 via network 110. In another embodiment, for example, program 116 can run on multiple server computers 114 that communicate with multiple client computers via network 110. Alternatively, the program can run on a web server that communicates with the server and multiple client computers via a network.
[0027] Network 110 may include wired connections, wireless connections, fiber optic connections, or combinations thereof. Typically, network 110 may be any combination of connections and protocols supporting communication between computer 102 and server computer 114. Network 110 may include various types of networks, such as local area networks (LANs), wide area networks (WANs) such as the Internet, telecommunications networks such as the Public Switched Telephone Network (PSTN), wireless networks, public switched networks, satellite networks, cellular networks (e.g., fifth-generation (5G) networks, long-term evolution (LTE) networks, third-generation (3G) networks, code division multiple access (CDMA) networks, etc.), public land mobile networks (PLMNs), metropolitan area networks (MANs), private networks, self-organizing networks, intranets, fiber-optic-based networks, etc., and / or combinations of these or other types of networks.
[0028] Figure 1 The number and arrangement of devices and networks shown are merely examples. In reality, additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or networks may exist. Figure 1 The devices and / or networks shown are arranged differently. Furthermore, Figure 1 The two or more devices shown can be implemented within a single device, or Figure 1 The single device shown can be implemented as multiple distributed devices. Alternatively or alternatively, a group of devices in system 100 (e.g., one or more devices) can perform one or more functions described as being performed by another group of devices in system 100.
[0029] Now refer to Figure 2 A block diagram 200 depicts an exemplary coding tree structure for video data. The coding tree structure may include a luminance component 202 and a chrominance component 204.
[0030] A semi-decoupled partitioning (SDP) scheme (i.e., semi-separated tree (SST) or flexible block partitioning of chroma components) can be used. According to SDP, the luma component 202 and chroma component 204 in a superblock (SB) can have the same or different block partitions, depending on the luma coding block size or the luma tree depth. When the luma block area size is greater than a threshold T1 or the luma block coding tree partition depth is less than or equal to a threshold T2, the chroma block can use the same coding tree structure as the luma. Otherwise, when the block area size is less than or equal to T1 or the luma partition depth is greater than T2, the corresponding chroma block can have a different coding block partition than the luma component; this can be called flexible block partitioning of the chroma component. T1 can be a positive integer, such as 128 or 256. T2 can be a positive integer, such as 1 or 2.
[0031] According to one or more implementations, when SDP can be applied and a chroma coding block can be associated with multiple luma coding blocks, the context for entropy coding of the chroma intra-prediction mode can depend on the corresponding luma blocks located at one or more predefined positions. In one implementation, one or more predefined positions may include the center position and / or the top-left corner of the current chroma block. In one implementation, the center position of the current chroma block may be the top-left corner of the current chroma block.
[0032] In one implementation, when the center and top-left corner of the current chroma block can both be used to locate the corresponding luminance block, the two luminance modes can be quantized before performing context selection. In one implementation, the luminance mode can be quantized into two values before performing the context selection process; the luminance mode can be a directional mode or a non-directional mode.
[0033] In another implementation, the non-directional pattern can be quantized into a single value, and the directional pattern can be quantized into a smaller set based on its angle. In one example, the directional pattern can be quantized into four values: 0 indicates that the angle of the current pattern can be equal to or less than 90 degrees, 1 indicates that the angle of the current pattern can be between 90 and 135 degrees, 2 indicates that the angle of the current pattern can be between 135 and 180 degrees, and 3 indicates that the angle of the current pattern can be greater than 180 degrees.
[0034] In one implementation, the context can be derived as an intra-frame prediction mode that can be used to predict the majority of samples in a co-located luma block.
[0035] In one implementation, multiple sample locations can be predefined, and intra-frame prediction modes associated with these locations can be identified for predicting co-localized luma blocks. Context values can then be derived using one of these identified prediction modes. In one example, the prediction mode most frequently used among the identified prediction modes can be used to derive the context values. In a second example, the predefined sample locations include four corner samples and a center / middle sample. In a third example, the predefined sample locations include four corner samples. In a fourth example, the predefined sample locations include two selected locations among the four corner samples and a center / middle sample. In a fifth example, the predefined sample locations include three selected locations among the four corner samples and a center / middle sample.
[0036] In one implementation, if co-located luma blocks at one or more predefined locations cannot be predicted using an intra-prediction mode, but the current chroma-coded block can be predicted using an intra-prediction mode, then the prediction modes of the co-located luma blocks can be mapped to one or more predefined intra-prediction modes. In one example, when co-located luma blocks can be encoded using IBC or Palette modes, a default intra-prediction mode can be used to derive the context for entropy coding of the chroma intra-prediction mode. The default intra-prediction mode includes, but is not limited to, DC, SMOOTH, SMOOTH-H, SMOOTH-V, or Paeth prediction modes.
[0037] According to one or more implementations, when signaling an intra-chroma prediction mode, a flag, namely the CfL flag, can first be signaled to indicate whether the current chroma mode can be a CfL mode. In one implementation, the intra-chroma prediction modes of adjacent blocks can be used to derive the context for signaling the CfL flag. In one example, a first context can be used when neither of the adjacent chroma modes is a CfL mode. Otherwise, a second context can be used. In another example, a first context can be used when neither of the adjacent chroma modes is a CfL mode. Otherwise, a second context can be used when one of the adjacent chroma modes is a CfL mode. Otherwise, a third context can be used.
[0038] In one implementation, the corresponding luminance mode can be used to derive the context for signaling the CfL flag. In one implementation, the coordinates of the corresponding luminance mode can be located at the center and top-left corner of the current chroma block. In another implementation, a first context can be used when the corresponding luminance mode is an directional mode. Otherwise, a second context can be used. In another implementation, three contexts can be used when two corresponding luminance modes can be applied to the context for determining the CfL flag. A first context can be used when both corresponding luminance modes are non-directional modes. Otherwise, a second context can be used when one of the corresponding luminance modes is a non-directional mode. Otherwise, a third context can be used.
[0039] According to one or more implementations, in order to signal chroma intra-frame prediction modes, a list can be constructed that includes previously encoded luma modes within the current picture / slice / tile. For the current chroma block, only the N luma modes with the highest occurrence rate can be allowed and signaled, where N can be a positive integer. In one implementation, N can be a power of 2. In one implementation, only the previously encoded luma modes within the current superblock row can be used. In one implementation, when SDP can be enabled, only the previously encoded luma modes within the current superblock can be used.
[0040] According to one or more implementations, all nominal intra-prediction angles are allowed for the luma-coded block, and all nominal intra-prediction angles may also be allowed and signaled for the chroma-coded block, while only a subset of incremental angles relative to the nominal angle may be allowed and signaled for the chroma-coded block. In one implementation, all non-directional modes, such as DC, SMOOTH, SMOOTH-H, and SMOOTH-V modes, may be allowed and signaled for the chroma-coded block. In one implementation, only incremental angles for co-located luma intra-prediction modes may be allowed and signaled for the chroma-coded block. In one implementation, the nominal angles along with non-directional modes may be signaled first. Then, if the current mode is a directional mode and equal to the co-located luma nominal mode, a second flag may be signaled to indicate the index of the incremental angle relative to the nominal angle. In another implementation, all allowed intra-prediction modes for the chroma-coded block may be signaled together.
[0041] Now refer to Figure 3 This describes an operational flowchart illustrating the steps of a method 300 for encoding and / or decoding video data. In some implementations, Figure 3 One or more processing blocks can be processed by computer 102 ( Figure 1 ) and server computer 114 ( Figure 1 ) Execution. In some implementations, Figure 3 One or more processing blocks may be executed by another device or group of devices that are separate from or include computer 102 and server computer 114.
[0042] At 302, method 300 includes receiving video data including chroma and luminance components.
[0043] At 304, method 300 includes identifying one or more contexts for entropy coding of intra-chroma prediction modes based on luma blocks co-located at one or more predefined locations.
[0044] At 306, method 300 includes decoding video data based on the identified context.
[0045] Understandable. Figure 3 This provides only an illustration of one implementation method and does not imply any limitations on how different implementation methods can be implemented. Many modifications can be made to the described environment based on design and implementation requirements.
[0046] Figure 4 According to the illustrative implementation method Figure 1 Block diagram 400 depicts the internal and external components of a computer. It should be understood that... Figure 4 This is merely an illustration of one implementation method and does not imply any limitation on the environments in which different implementation methods can be implemented. Many modifications can be made to the described environment based on design and implementation requirements.
[0047] Computer 102 ( Figure 1 ) and server computer 114 ( Figure 1 ) can include Figure 4 The corresponding sets of internal components 800A, 800B and external components 900A, 900B shown. Each set of internal components 800 includes one or more processors 820 on one or more buses 826, one or more computer-readable RAMs 822 and one or more computer-readable ROMs 824, one or more operating systems 828, and one or more computer-readable tangible storage devices 830.
[0048] Processor 820 is implemented in hardware, firmware, or a combination of hardware and software. Processor 820 is a central processing unit (CPU), graphics processing unit (GPU), accelerated processing unit (APU), microprocessor, microcontroller, digital signal processor (DSP), field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), or another type of processing unit. In some implementations, processor 820 includes one or more processors that can be programmed to perform functions. Bus 826 includes components that allow communication between internal components 800A and 800B.
[0049] One or more operating systems 828, software programs 108 ( Figure 1 ) and server computer 114 ( Figure 1 The video encoding program 116 on ) Figure 1 The data is stored on one or more of the corresponding computer-readable tangible storage devices 830 for execution by one or more of the corresponding processors 820 via one or more of the corresponding RAMs 822 (which typically include cache memory). Figure 4In the embodiments shown, each of the computer-readable tangible storage devices 830 is a disk storage device of an internal hard disk drive. Alternatively, each of the computer-readable tangible storage devices 830 is a semiconductor storage device, such as a ROM 824, EPROM, flash memory, optical disc, magneto-optical disc, solid-state disk, compact disc (CD), digital versatile disc (DVD), floppy disk, cassette tape, magnetic tape, and / or another type of non-transitory computer-readable tangible storage device that can store computer programs and digital information.
[0050] Each set of internal components 800A and 800B also includes an R / W drive or interface 832 for reading and writing to one or more portable computer-readable tangible storage devices 936, such as CD-ROMs, DVDs, memory sticks, magnetic tapes, disks, optical discs, or semiconductor storage devices. Software programs such as 108 ( Figure 1 ) and video encoding program 116 ( Figure 1 The software program can be stored on one or more of the corresponding portable computer-readable tangible storage devices 936, read via the corresponding R / W drive or interface 832 and loaded into the corresponding hard disk drive 830.
[0051] Each set of internal components 800A and 800B also includes a network adapter or interface 836, such as a TCP / IP adapter card; a wireless Wi-Fi interface card; or a 3G, 4G, or 5G wireless interface card or other wired or wireless communication link. Software program 108 ( Figure 1 ) and server computer 114 ( Figure 1 The video encoding program 116 on ) Figure 1 ) can be downloaded from an external computer to computer 102 via a network (e.g., the Internet, a local area network, or other, a wide area network) and a corresponding network adapter or interface 836. Figure 1 The network includes a network adapter or interface 836 and a server computer 114. Software program 108 and video encoding program 116 on server computer 114 are loaded from the network adapter or interface 836 into the corresponding hard disk drive 830. The network may include copper wire, fiber optic, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers.
[0052] Each set of external components 900A and 900B may include a computer monitor 920, a keyboard 930, and a computer mouse 934. External components 900A and 900B may also include a touchscreen, a virtual keyboard, a touchpad, a pointing device, and other human-machine interface devices. Each set of internal components 800A and 800B also includes a device driver 840 that interfaces with the computer monitor 920, keyboard 930, and computer mouse 934. Device driver 840, R / W driver or interface 832, and network adapter or interface 836 include hardware and software (stored in storage device 830 and / or ROM 824).
[0053] It should be understood in advance that although this disclosure includes a detailed description of cloud computing, the implementation of the teachings described herein is not limited to cloud computing environments. Rather, some implementations can be combined with any other type of computing environment now known or developed hereafter.
[0054] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with service providers. This cloud model may include at least five features, at least three service models, and at least four deployment models.
[0055] The characteristics are as follows:
[0056] On-demand self-service: Cloud consumers can unilaterally and automatically supply computing power, such as server time and network storage, as needed, without requiring human interaction with the service provider.
[0057] Extensive network access: Capabilities are available via the network and are accessed through standard mechanisms that facilitate use by heterogeneous thin-client or thick-client platforms, such as mobile phones, laptops, and PDAs.
[0058] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically allocated and reallocated based on demand. There is a sense of location agnosticness because consumers typically do not control or know the exact location of the resources provided, but can specify the location at a higher level of abstraction (e.g., country, state, or data center).
[0059] Rapid and elastic: Capacity can be supplied quickly and elastically (in some cases automatically) to scale outwards rapidly and to scale inwards rapidly. For consumers, the capacity available for supply often appears unlimited and can be purchased in any quantity at any time.
[0060] Measurement services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to service types (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency for both the providers and consumers of the services being utilized.
[0061] The service model is as follows:
[0062] Software as a Service (SaaS): This provides consumers with the ability to use the provider's applications running on cloud infrastructure. Applications can be accessed from various client devices via thin client interfaces such as web browsers (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating system, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.
[0063] Platform as a Service (PaaS): This provides consumers with the ability to deploy consumer-created or acquired applications, built using programming languages and tools supported by the provider, onto cloud infrastructure. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but they have control over the deployed applications and the configuration of possible application hosting environments.
[0064] Infrastructure as a Service (IaaS): This provides consumers with the capability to offer processing, storage, networking, and other basic computing resources that enable consumers to deploy and run arbitrary software, which may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but they do have control over the operating system, storage, deployed applications, and possibly limited control over the selection of network components (e.g., host firewalls).
[0065] The deployment model is as follows:
[0066] Private cloud: Cloud infrastructure for organization operations only. It can be managed by the organization or a third party and can exist on-site or off-site.
[0067] Community cloud: A cloud infrastructure shared by several organizations and supporting a specific community with shared concerns (e.g., tasks, security requirements, policies, and compliance considerations). It can be managed by an organization or a third party and can exist on-site or off-site.
[0068] Public cloud: Cloud infrastructure available to the general public or large industry groups and owned by the organization that sells cloud services.
[0069] Hybrid cloud: A cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain a single entity but are bound together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursts for load balancing between clouds).
[0070] Cloud computing environments are service-oriented, emphasizing statelessness, loose coupling, modularity, and semantic interoperability. At the heart of cloud computing is the infrastructure of a network of interconnected nodes.
[0071] Reference Figure 5 The illustration depicts a cloud computing environment 500. As shown, the cloud computing environment 500 includes one or more cloud computing nodes 10, with local computing devices used by cloud consumers, such as personal digital assistants (PDAs) or cellular phones 54A, desktop computers 54B, laptop computers 54C, and / or automotive computer systems 54N, capable of communicating with the one or more cloud computing nodes 10. The cloud computing nodes 10 can communicate with each other. They can be physically or virtually grouped in one or more networks, such as private clouds, community clouds, public clouds, or hybrid clouds, or combinations thereof, as described above (not shown). This allows the cloud computing environment 600 to provide infrastructure, platform, and / or software as a service, without requiring cloud consumers to maintain resources on their local computing devices. It should be understood that... Figure 5 The types of computing devices 54A to 54N shown are intended to be illustrative only, and cloud computing node 10 and cloud computing environment 500 can communicate with any type of computerized device via any type of network and / or network-addressable connection (e.g., using a web browser).
[0072] Reference Figure 6 This demonstrates the 500-fold cloud computing environment ( Figure 5 This provides a set of functional abstraction layers, 600. It should be understood beforehand that... Figure 6 The components, layers, and functions shown are intended to be illustrative only, and the implementation is not limited thereto. As depicted, the following layers and corresponding functions are provided:
[0073] The hardware and software layer 60 includes hardware components and software components. Examples of hardware components include: a mainframe 61; a server 62 based on a RISC (Reduced Instruction Set Computer) architecture; a server 63; a blade server 64; a storage device 65; and a network and network components 66. In some implementations, software components include network application server software 67 and database software 68.
[0074] The virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual server 71; virtual storage 72; virtual network including virtual private network 73; virtual application and operating system 74; and virtual client 75.
[0075] In one example, management layer 80 may provide the following functionalities: Resource Provisioning 81 provides dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and Pricing 82 provides cost tracking for utilizing resources within the cloud computing environment, as well as billing or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Protection provides authentication for cloud consumers and tasks, and protection for data and other resources. User Portal 83 provides access to the cloud computing environment for consumers and system administrators. Service Level Management 84 provides cloud resource allocation and management to meet the required service level. Service Level Agreement (SLA) Planning and Fulfillment 85 provides pre-scheduling and procurement of cloud resources and anticipates future demand for cloud resources according to the SLA.
[0076] Workload layer 90 provides examples of functionalities that can be used in cloud computing environments. Examples of workloads and functionalities that can be provided from this layer include: mapping and navigation 91; software development and lifecycle management 92; virtual classroom delivery 93; data analytics and processing 94; transaction processing 95; and video encoding 96. Video encoding 96 can encode and / or decode video data based on context associated with luma blocks co-located at one or more predefined locations.
[0077] Some implementations may relate to systems, methods, and / or computer-readable media at any possible level of integration technical detail. A computer-readable medium may include a computer-readable non-transitory storage medium (or multiple media) having computer-readable program instructions on it for causing a processor to perform operations.
[0078] Computer-readable storage media can be tangible devices capable of retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to: electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable optical disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanical encoding devices such as punched cards or raised structures in recesses on which instructions are recorded, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as being a transient signal, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through fiber optic cables), or electrical signals transmitted through wires.
[0079] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device, or downloaded via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network) to an external computer or external storage device. The network may include copper cables, optical fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the suitable computing / processing device.
[0080] Computer-readable program code / instructions used to perform operations can be: assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of execution entirely on a remote computer or server, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the use of the Internet provided by an Internet service provider). In some implementations, electronic circuit systems, including, for example, programmable logic circuit systems, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), can personalize the electronic circuit system by executing computer-readable program instructions using state information of computer-readable program instructions in order to perform aspects or operations.
[0081] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other device to function in a particular manner, such that the computer-readable storage medium storing the instructions includes an article of writing comprising instructions for implementing aspects of the functions / actions specified in one or more blocks of a flowchart and / or block diagram.
[0082] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device such that a series of operational steps to be performed on a computer, other programmable apparatus or other device can produce a computer-implemented process, and that the instructions to be performed on a computer, other programmable apparatus or other device implement the functions / actions specified in one or more blocks of a flowchart and / or block diagram.
[0083] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing a specific logical function. Methods, computer systems, and computer-readable media may include additional blocks, fewer blocks, different blocks, or blocks arranged differently compared to those depicted in the figures. In some alternative implementations, the functions indicated in a block may not occur in the order indicated in the figures. For example, two blocks shown consecutively may actually be executed simultaneously or substantially simultaneously, or blocks may sometimes be executed in reverse order depending on the functions involved. It will also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a hardware-based dedicated system that performs a specific function or action or executes a combination of dedicated hardware and computer instructions.
[0084] It will be apparent that the systems and / or methods described herein can be implemented in various forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the implementation method. Therefore, while the operation and behavior of the systems and / or methods are described herein without reference to specific software code, it should be understood that software and hardware can be designed to implement the systems and / or methods based on the descriptions herein.
[0085] Unless explicitly stated otherwise, no element, action, or instruction used herein should be construed as critical or necessary. Furthermore, as used herein, the articles “a” and “an” are intended to include one or more items and may be used interchangeably with “one or more.” Additionally, as used herein, the term “set” is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and may be used interchangeably with “one or more.” The term “one” or similar language is used where only one item is intended. Furthermore, as used herein, the terms “have,” “possess,” “contain,” etc., are intended to be open-ended terms. Furthermore, unless explicitly stated otherwise, the phrase “based on” is intended to mean “at least partially based on.”
[0086] Descriptions of various aspects and implementations have been presented for illustrative purposes, but these descriptions are not intended to be exhaustive or limited to the disclosed implementations. Although combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features can be combined in ways not specifically recited in the claims and / or not disclosed in the specification. Although each dependent claim listed below may directly depend on only one claim, the disclosure of possible implementations includes each dependent claim combined with each other claim in the claim set. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described implementations. The terminology used herein has been chosen to best illustrate the principles of the implementations, their practical application, or technical improvements relative to technologies found in the market, or to enable others skilled in the art to understand the implementations disclosed herein.
Claims
1. A video encoding method, characterized in that, The method includes: Receives video data including chroma and luminance components; Set up one or more contexts for entropy coding of chroma intra-frame prediction modes, the one or more contexts corresponding to one or more co-located luma blocks at one or more predefined locations; Based on one or more contexts, a first flag is signaled to indicate that the intra-prediction chroma mode associated with the current chroma block is a chroma mode derived from luma, wherein when the first flag is a first value, it indicates that the one or more co-located luma blocks are encoded using the directional mode, and when the first flag is a second value, it indicates that the one or more co-located luma blocks are not encoded using the directional mode; and The video data is encoded based on one or more contexts that are set.
2. The method according to claim 1, characterized in that, The one or more predefined positions include one or more of the middle position and the top left corner of the current chroma block.
3. The method according to claim 2, characterized in that, The middle position of the current chroma block corresponds to the upper left corner of the chroma block.
4. The method according to claim 1, characterized in that, The method further includes: Identify one or more predefined sample locations; and Identify one or more intra-frame prediction patterns associated with the one or more predefined sample locations for predicting the one or more co-localized brightness blocks. The one or more contexts set are also based on prediction modes in one or more identified intra-frame prediction modes.
5. The method according to claim 1, characterized in that, The one or more contexts set are also based on the one or more chroma intra-frame prediction modes, including the chroma mode from the luminance.
6. The method according to claim 4, characterized in that, The one or more predefined sample locations include one or more corner samples and center samples; or The predefined sample locations include four corner samples.
7. The method according to claim 1, characterized in that, The one or more contexts set are also based on one or more intra-luminance prediction modes associated with the one or more co-located luminance blocks.
8. The method according to claim 7, characterized in that, When it is determined that the one or more intra-frame prediction modes of luminance include one or more directional modes, the one or more contexts set include a first context; or When it is determined that at least one of the one or more intra-frame prediction modes for luminance includes a directional mode, the set one or more contexts include a second context; or When it is determined that none of the one or more luminance intra-frame prediction modes is a directional mode, the one or more contexts set include a third context.
9. The method according to claim 5, characterized in that, When at least one of the one or more chroma intra-frame prediction modes is determined to be the chroma mode derived from luminance, the one or more contexts set include a second context.
10. The method according to claim 1, characterized in that, For chroma-coded blocks, all nominal intra-frame prediction angles are allowed and signaled to be permitted for luminance-coded blocks.
11. The method according to claim 10, characterized in that, For chroma intra-coded blocks, only a subset of incremental angles for the nominal intra-predicted angle are permitted and signaled.
12. The method according to claim 11, characterized in that, For chroma intra-frame coded blocks, allow and notify all non-directional modes via signaling; or For the chroma coding block, only incremental angles corresponding to the co-located intra-frame prediction mode of lumen are permitted and signaled.
13. The method according to claim 12, characterized in that, The nominal intra-frame prediction angle, along with the non-directional mode, is notified via signaling; and Based on the chroma intra-prediction mode associated with the current chroma block being the orientation mode and equal to the co-located luminance nominal mode, a second flag is signaled to indicate the index of the incremental angle for the nominal intra-prediction angle.
14. The method according to claim 12, characterized in that, All allowed intra-prediction modes for a chroma-coded block are notified together via signaling.
15. A video decoding method, characterized in that, The method includes: Receive a video bitstream comprising multiple blocks and a first flag of the current chroma block; the multiple blocks include the current chroma block. Based on the number of neighboring blocks using the chroma CfL mode from the luma of the current block, an entropy decoding context is selected for entropy decoding of one or more parameters of the current chroma block, including: When the number of neighboring blocks using the CfL mode is less than a threshold number, the first context is selected as the entropy decoding context. When the number of neighboring blocks using the CfL mode is greater than the threshold number, a second context is selected as the entropy decoding context, and the second context is different from the first context; The first flag is entropy decoded from the video bitstream using the entropy decoding context, wherein the first flag indicates whether the CfL mode is enabled for the current chroma block; and Decode the current chroma block based on the value of the first flag.
16. A video encoding method, characterized in that, The method includes: Receive video data comprising multiple video blocks, wherein the multiple video blocks include the current chroma block; Based on the number of neighboring blocks using the chroma CfL mode from the luma of the current block, an entropy coding context is selected for entropy coding one or more parameters of the current chroma block, including: When the number of neighboring blocks using the CfL mode is less than a threshold number, the first context is selected as the entropy coding context. When the number of neighboring blocks using the CfL mode is greater than the threshold number, a second context is selected as the entropy coding context, and the second context is different from the first context; The first flag is entropy-encoded using the entropy coding context, wherein the first flag indicates whether the CfL mode is enabled for the current chroma block; and The entropy-encoded first flag is communicated via signaling in the video bitstream.
17. A method for processing video data, characterized in that, The method includes: Obtain the source video sequence, which includes multiple video blocks; The conversion is performed between the source video sequence and the video bitstream. The video bitstream includes: Multiple encoded blocks corresponding to the plurality of video blocks; and A first flag indicates whether a chroma CfL mode from luma is enabled for a first chroma block among the plurality of encoded blocks, wherein the first flag uses context for entropy coding, the context being selected based on the number of neighboring blocks of the first chroma block that use the chroma CfL mode from luma, including: When the number of neighboring blocks using the CfL mode is less than a threshold number, the first flag is entropy-encoded using the first context; and When the number of neighboring blocks using the CfL mode is greater than the threshold number. The first flag is entropy encoded using a second context, which is different from the first context.
18. A method for storing or transmitting video bitstreams, characterized in that, The video bitstream is generated by the method according to any one of claims 1 to 14 and 16, or the video bitstream is decoded based on the method of claim 15.
19. A computer system for decoding video data, characterized in that, The computer system includes: One or more computer-readable non-transitory storage media are configured to store computer program code; and One or more computer processors are configured to access the computer program code and operate as instructed by the computer program code to perform the method of any one of claims 1 to 18.
20. A non-transitory computer-readable medium, characterized in that, The non-transitory computer-readable medium stores a video bitstream generated by the method according to any one of claims 1 to 14 and 16.