Method, System, Program for Signal Transmission of Symbolic Tree Unit Size
By using separate flags to signal the coding tree unit size in video encoding, the inefficiencies in existing standards are addressed, resulting in improved bit savings and reduced memory usage.
Patent Information
- Application Number
- JP2024095604
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-17
- Filing Date
- 2024-06-13
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-09-22
AI Technical Summary
The existing video encoding standards, such as HEVC, face inefficiencies in coding tree unit size signaling, which can result in wasted bits and increased memory usage.
The proposed solution involves using separate flags to signal the coding tree unit size instead of fixed-length coding, allowing for more efficient bit allocation and reduced memory usage.
This approach enables better bit savings and reduced memory usage by allowing flexible signaling of the coding tree unit size, thereby improving the efficiency of video encoding and decoding processes.
Smart Images

Figure 0007684008000001 
Figure 0007684008000002 
Figure 0007684008000003
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority from U.S. Provisional Patent Application No. 62 / 905,339, filed on September 24, 2019, and U.S. Patent Application No. 17 / 024,246, filed on September 17, 2020. The entireties of those are incorporated herein by reference.
[0002] Field The present disclosure generally relates to the field of data processing, and more particularly to video encoding and decoding.
Background Art
[0003] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). In 2015, both standardization organizations jointly formed the JVET (Joint Video Exploration Team) to explore the possibility of developing the next video coding standard beyond HEVC, and in October 2017, they announced a "Call for Proposals (CfP) on Video Compression with Beyond-HEVC Capabilities". By February 15, 2018, a total of 22 CfP responses regarding standard dynamic range (SDR), 12 CfP responses regarding high dynamic range (HDR), and 12 CfP responses regarding the 360 video category were submitted respectively. In April 2018, all received CfP responses were evaluated at the 122nd MPEG / 10th JVET meeting. As a result of this meeting, JVET officially started the standardization process for the next-generation video coding beyond HEVC. The new standard was named Versatile Video Coding (VVC), and JVET was renamed the Joint Video Expert Team. The current version of the VTM (VVC Test Model) is VTM6.
Summary of the Invention
[0004] Embodiments relate to a method, system, and computer-readable medium for encoding video data. According to one aspect, a method for encoding video data is provided. The method may include receiving video data having a certain coding tree unit size. The coding tree unit size associated with the video data is signaled by setting two or more flags. The video data is encoded / decoded based on the flags corresponding to the signaled coding tree unit size.
[0005] According to another aspect, a computer system for encoding video data is provided. The computer system may include one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage devices, and program instructions. The program instructions are stored in at least one of the one or more storage devices for execution by at least one of the one or more processors via at least one of the one or more memories, whereby the computer system can execute the method. The method may include receiving video data having a coding tree unit size. The coding tree unit size associated with the video data is signaled by setting two or more flags. The video data is encoded / decoded based on the flags corresponding to the signaled coding tree unit size.
[0006] According to yet another aspect, a computer-readable medium for encoding video data is provided. The computer-readable medium may include one or more computer-readable storage devices and program instructions stored in at least one of the one or more tangible storage devices, the program instructions being executable by a processor. The program instructions are executable by a processor to perform a method, which may suitably include receiving video data having a coding tree unit size. The coding tree unit size associated with the video data is signaled by setting two or more flags. The video data is encoded / decoded based on the flags corresponding to the signaled coding tree unit size.
Brief Description of the Drawings
[0007] These and other objects, features and advantages will become apparent from the following detailed description of exemplary embodiments read in conjunction with the accompanying drawings. The drawings are for the purpose of facilitating understanding by those skilled in the art in conjunction with the detailed description and, for the sake of clarity, the various features of the drawings are not to scale.
[0008]
Figure 1
[0009]
Figure 2
[0010]
Figure 3A
Figure 3B
Figure 3C
Figure 3D
[0011]
Figure 4
[0012]
Figure 5
[0013]
Figure 6
[0014]
Figure 7
Best Mode for Carrying Out the Invention
[0015] Detailed embodiments of the structures and methods according to the claims are disclosed herein, but it can be understood that the disclosed embodiments merely exemplify the structures and methods according to the claims that can be embodied in various forms. However, these structures and methods can be embodied in many different forms and should not be construed as being limited to the exemplary embodiments described herein. Rather, these exemplary embodiments are provided so that this disclosure is complete and full and conveys the scope fully to those skilled in the art. In the description, details of well-known features and techniques may be omitted in order to avoid obscuring the presented embodiments unnecessarily.
[0016] Embodiments generally relate to the field of data processing, and more particularly, to video encoding and decoding. The exemplary embodiments described below provide, among other things, a system, method, and computer program for encoding video data using a separate flag to replace the coded tree unit size syntax. Thus, some embodiments have the ability to improve the field of computing by saving bits by signaling the coded tree unit size through a flag and thereby enabling less memory usage.
[0017] As described above, ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). In 2015, both standard organizations formed the Joint Video Exploration Team (JVET) jointly to explore the possibility of developing the next video coding standard beyond HEVC, and in October 2017, they announced the "Call for Proposals (CfP) on Video Compression with Beyond-HEVC capabilities". By February 15, 2018, a total of 22 CfP responses regarding standard dynamic range (SDR), 12 CfP responses regarding high dynamic range (HDR), and 12 CfP responses regarding the 360 video category were submitted respectively. In April 2018, all received CfP responses were evaluated at the 122nd MPEG / 10th JVET meeting. As a result of this meeting, JVET officially started the standardization process for the next-generation video coding beyond HEVC. The new standard was named Versatile Video Coding (VVC), and JVET was renamed the Joint Video Expert Team. The current version of the VVC Test Model (VTM) is VTM6.
[0018] In HEVC, the coding tree unit is divided into coding units by using a quadtree structure called a coding tree in order to adapt to various local features. The decision of whether to code the picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction is made at the level of the coding unit. Each coding unit can be further divided into one, two, or four prediction units according to the prediction unit division type. Within one prediction unit, the same prediction process is applied, and the relevant information is transmitted to the decoder on a prediction unit basis. After obtaining the residual block by applying the prediction process based on the prediction unit division type, the coding unit can be divided into transform units according to another quadtree structure such as the coding tree for the coding unit. One of the important features of the HEVC structure is having multiple division concepts including coding units, prediction units, and transform units. However, describing the syntax log2_ctu_size_minus5 using fixed-length coding u(2) may waste one bit. There may be only three numbers to be coded, namely 0, 1, and 2, and if the number to be coded is 0 or 1, u(2) may waste one bit. Therefore, in order to save bits in the sequence parameter set, it may be advantageous to replace the original coding tree unit size syntax with a plurality of separate flags.
[0019] Aspects are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer-readable media according to various embodiments. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0020] Referring now to FIG. 1, a functional block diagram of a network-connected computer environment is shown for a video encoding system 100 (hereinafter, "system") for encoding video data using a separate flag that replaces the coded tree unit size syntax. FIG. 1 provides only an illustration of one implementation and should be understood not to imply any limitation regarding the environments in which various embodiments may be implemented. Many modifications to the illustrated environment may be made based on design and implementation requirements.
[0021] System 100 can include a computer 102 and a server computer 114. Computer 102 may communicate with server computer 114 via a communication network 110 (hereinafter, "network"). Computer 102 may include a processor 104 and a software program 108 stored in data storage device 106 that interfaces with a user and is enabled to communicate with server computer 114. As will be described later with reference to FIG. 5, computer 102 may include internal components 800A and external components 900A, respectively, and server computer 114 may include internal components 800B and external components 900B, respectively. Computer 102 may be, for example, a mobile device, a telephone, a personal digital assistant, a netbook, a laptop computer, a tablet computer, a desktop computer, or any type of computing device capable of executing a program, accessing a network, and accessing a database.
[0022] Server computer 114 may also operate in a cloud computing service model such as software as a service (SaaS), platform as a service (PaaS), or infrastructure as a service (IaaS), as will be described later with respect to FIGS. 6 and 7. Server computer 114 may also be located within a cloud computing deployment model such as a private cloud, community cloud, public cloud, or hybrid cloud.
[0023] Server computer 114, which can be used to encode video data, is enabled to execute a video encoding program 116 (hereinafter, “program”) that can interact with database 112. The video encoding program method will be described in more detail below with respect to FIG. 4. In one embodiment, computer 102 may operate as an input device including a user interface, while program 116 may be executed primarily on server computer 114. In an alternative embodiment, program 116 may be executed primarily on one or more computers 102, while server computer 114 may be used for processing and storing data used by program 116. It should be noted that program 116 may be a stand-alone program or may be integrated into a larger video encoding program.
[0024] However, it should be noted that the processing for program 116 can, in some cases, be shared between computer 102 and server computer 114 in any ratio. In another embodiment, program 116 can operate on two or more computers, server computers, or some combination of computers and server computers, for example, multiple computers 102 that communicate with a single server computer 114 across network 110. In another embodiment, for example, program 116 can operate on multiple server computers 114 that communicate with multiple client computers across network 110. Alternatively, the program may operate on a network server that communicates with servers and multiple client computers across the network.
[0025] Network 110 may include a wired connection, a wireless connection, a fiber optic connection, or some combination thereof. Generally, network 110 can be any combination of connections and protocols that support communication between computer 102 and server computer 114. Network 110 can be, for example, a local area network (LAN), a wide area network (WAN) such as the Internet, a telecommunications network such as a public switched telephone network (PSTN), a wireless network, a public switched network, a satellite network, a cellular network (e.g., a fifth generation (5G) network, a long term evolution (LTE) network, a third generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a metropolitan area network (MAN), a private network, an ad hoc network, an intranet, a fiber optic-based network, etc., and / or combinations of these or other types of networks.
[0026] The number and arrangement of the devices and networks shown in FIG. 1 are given as an example. In practice, there may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks with different arrangements than those shown in FIG. 1. Further, two or more of the devices shown in FIG. 1 may be implemented within a single device, or the single device shown in FIG. 1 may be implemented as a plurality of distributed devices. Additionally or alternatively, a set of devices (e.g., one or more devices) of system 100 may perform one or more functions described as being performed by another set of devices of system 100.
[0027] Referring now to FIG. 2, an exemplary QTBT block structure 200 is shown. The QTBT block structure 200 may include block splitting by using QTBT. A corresponding tree representation can also be drawn. Solid lines may indicate quadtree splitting, and dotted lines may indicate binary tree splitting. At each split (i.e., non-leaf) node of the binary tree, one flag may be signaled to indicate which split type (i.e., horizontal or vertical) can be used, where 0 may indicate a horizontal split and 1 may indicate a vertical split. For quadtree splitting, since the quadtree splitting always splits the block in both the horizontal and vertical directions to generate four sub-blocks of equal size, it may not be necessary to indicate the split type.
[0028] QTBT can remove the concept of multiple split types. For example, QTBT can remove the separation of the concepts of coding unit, prediction unit, and transform unit, and support greater flexibility for the split shape of the coding unit. In the QTBT block structure 200, the coding unit can have either a square or rectangular shape. The coding tree unit (CTU) may be split by a quadtree structure. The quadtree leaf node may be further split by a binary tree structure. There may be two split types for binary tree splitting: symmetric horizontal splitting and symmetric vertical splitting.
[0029] A binary tree leaf node may be called a coding unit (CU), and its segmentation can be used for prediction and transformation processing without further splitting. This means that the coding unit, prediction unit, and transformation unit may have the same block size in the QTBT coding block structure 200. A coding unit may contain coding blocks (CBs) of different color components (e.g., one coding unit can contain one luma CB and two chroma CBs for P slices and B slices in 4:2:0 chroma format), or it may contain a single-component CB (e.g., one coding unit can contain one luma CB or two chroma CBs for I slices).
[0030] For the QTBT splitting method, the following parameters may be defined: The coding tree unit size may be the root node size of a quadtree, which is the same concept as in HEVC; MinQTSize may be the minimum allowable quadtree leaf node size; MaxBTSize may be the maximum allowable binary tree root node size; MaxBTDepth may be the maximum allowable binary tree depth; MinBTSize may be the minimum allowable binary tree leaf node size.
[0031] In an example of the QTBT splitting structure 200, the coding tree unit size may be set as 128×128 luma samples having two corresponding 64×64 blocks of chroma samples. MinQTSize may be set to 16×16. MaxBTSize may be set to 64×64. MinBTSize (for both width and height) may be set to 4×4. MaxBTDepth may be set to 4. The quadtree splitting may first be applied to the coding tree unit to generate quadtree leaf nodes. The quadtree leaf nodes can have sizes ranging from 16×16 (i.e., MinQTSize) to 128×128 (i.e., the coding tree unit size). When the leaf quadtree node is 128×128, it cannot be further split by a binary tree because its size may exceed MaxBTSize (i.e., 64×64). Otherwise, the leaf quadtree node may be further split by a binary tree. Thus, the quadtree leaf node may also be the root node for the binary tree and have a binary tree depth of 0.
[0032] When the binary tree depth reaches MaxBTDepth (i.e., 4), further splitting may not be considered. When the binary tree node has a width equal to MinBTSize (i.e., 4), further horizontal splitting may not be considered. Similarly, when the binary tree node has a height equal to MinBTSize, further vertical splitting may not be considered. The leaf nodes of the binary tree are further processed by prediction and transformation processing without further splitting. In an example, the maximum coding tree unit size may be 256×256 luma samples.
[0033] Furthermore, the QTBT scheme may support the flexibility for luma and chroma to have separate QTBT structures. Currently, for P and B slices, the luma and chroma CTBs in one coding tree unit may share the same QTBT structure. However, for I slices, the luma CTB may be divided into CUs by the QTBT structure, and the chroma CTB may be divided into chroma coding units by another QTBT structure. This means that a CU in an I slice may contain a coding block of the luma component or coding blocks of two chroma components, while a coding unit in a P or B slice consists of coding blocks of all three color components.
[0034] Here, referring to A - D in FIG. 3, exemplary syntax elements 300A - 300D are shown. The syntax elements 300A - 300D may be used to signal the coding tree unit size so as to save bits.
[0035] According to one or more embodiments, two of the three flags, namely, use_32_ctu_size_flag, use_64_ctu_size_flag, and use_128_ctu_size_flag, may be used to signal the coding tree unit size. In certain embodiments, use_32_ctu_size_flag may be signaled first. If use_32_ctu_size_flag is equal to 1, signaling of the coding tree unit size may be complete. Otherwise, use_64_ctu_size_flag may be signaled. In certain embodiments, use_64_ctu_size_flag may be signaled first. If use_64_ctu_size_flag is equal to 1, signaling of the coding tree unit size may be complete. Otherwise, use_32_ctu_size_flag may be signaled. In certain embodiments, use_32_ctu_size_flag may be signaled first. If use_32_ctu_size_flag is equal to 1, signaling of the coding tree unit size may be complete. Otherwise, use_128_ctu_size_flag may be signaled. In certain embodiments, use_128_ctu_size_flag may be signaled first. If use_128_ctu_size_flag is equal to 1, signaling of the coding tree unit size may be complete. Otherwise, use_32_ctu_size_flag may be signaled. In certain embodiments, use_64_ctu_size_flag may be signaled first. If use_64_ctu_size_flag is equal to 1, signaling of the coding tree unit size may be complete. Otherwise, use_128_ctu_size_flag may be signaled. In certain embodiments, use_128_ctu_size_flag may be signaled first. If use_128_ctu_size_flag is equal to 1, signaling of the coding tree unit size may be complete. Otherwise, use_64_ctu_size_flag may be signaled.
[0036] Referring now to A - B of FIG. 3, according to one or more embodiments, individual flags within the sequence parameter set indicate whether the smallest coding tree unit size can be applied (use_smallest_ctu_size_flag) and whether the largest coding tree unit size can be applied (use_largest_ctu_size_flag). In one embodiment, a sequence parameter set flag indicating whether the smallest coding tree unit size can be applied may be signaled first, and if the smallest coding tree unit size is not applied, another sequence parameter set flag indicating whether the largest coding tree unit size can be applied may be signaled. In another embodiment, a sequence parameter set flag indicating whether the largest coding tree unit size can be applied may be signaled first, and if the largest coding tree unit size is not applied, another sequence parameter set flag indicating whether the smallest coding tree unit size can be applied may be signaled.
[0037] If use_smallest_ctu_size_flag is equal to 1, it may be specified that the luma coding tree block size of each coding tree unit may be equal to 32×32, and if use_smallest_ctu_size_flag is equal to 0, it may be specified that use_largest_ctu_size_flag may exist.
[0038] If use_largest_ctu_size_flag is equal to 1, it may be specified that the luma coding tree block size of each coding tree unit may be equal to 128×128, and if use_largest_ctu_size_flag is equal to 0, it may be specified that the luma coding tree block size of each coding tree unit may be equal to 64×64.
[0039] Adding 2 to log2_min_luma_coding_block_size_minus2 may specify the minimum luma coding block size.
[0040] The variables CtbLog2SizeY, CtbSizeY, MinCbLog2SizeY, MinCbSizeY, MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSizeY, MaxTbSizeY, PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC, and PicHeightInSamplesC may be derived as follows: if(use_smallest_ctu_size_flag) CtbLog2SizeY = 0 else if(use_largest_ctu_size_flag) CtbLog2SizeY = 2 else CtbLog2SizeY = 1 CtbSizeY = 1 << CtbLog2SizeY
[0041] Referring now to FIG. 3C, according to one or more embodiments, sps_max_luma_transform_size_64_flag may be signaled only if the coding tree unit size may be 64×64 or greater. In one embodiment, use_32_ctu_size_flag may be signaled first, and if it is equal to 1, sps_max_luma_transform_size_64_flag may not be signaled. In one embodiment, use_64_ctu_size_flag may be signaled first, and if it is equal to 0, then use_32_ctu_size_flag may be signaled and may be equal to 1, and sps_max_luma_transform_size_64_flag may not be signaled. In one embodiment, use_128_ctu_size_flag may be signaled first, and if it is equal to 0, then use_32_ctu_size_flag may be signaled and may be equal to 1, and sps_max_luma_transform_size_64_flag may not be signaled.
[0042] That use_32_ctu_size_flag is equal to 1 may specify that the luma coding tree block size of each coding tree unit may be equal to 32×32, and that use_32_ctu_size_flag is equal to 0 may specify that use_128_ctu_size_flag may exist. That use_128_ctu_size_flag is equal to 1 may specify that the luma coding tree block size of each coding tree unit may be 128×128, and that use_128_ctu_size_flag is equal to 0 may specify that the luma coding tree block size of each coding tree unit may be equal to 64×64.
[0043] Adding 2 to log2_min_luma_coding_block_size_minus2 may specify the minimum luma coding block size.
[0044] The sps_max_luma_transform_size_64_flag being equal to 1 may specify that the maximum transform size in luma samples may be equal to 64, and the sps_max_luma_transform_size_64_flag being equal to 0 may specify that the maximum transform size in luma samples may be equal to 32. If it does not exist, the value of the sps_max_luma_transform_size_64_flag may be presumed to be equal to 0.
[0045] The variables CtbLog2SizeY, CtbSizeY, MinCbLog2SizeY, MinCbSizeY, MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSizeY, MaxTbSizeY, PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC, and PicHeightInSamplesC may be derived as follows: if(use_32_ctu_size_flag) CtbLog2SizeY = 0 elseif(use_128_ctu_size_flag) CtbLog2SizeY = 2 else CtbLog2SizeY = 1 CtbSizeY = 1 << CtbLog2SizeY
[0046] Referring to D in FIG. 3 here, according to one or more embodiments, the sps_max_luma_transform_size_64_flag may be signaled only if the coding tree unit size is not the smallest coding tree unit size.
[0047] That use_smallest_ctu_size_flag is equal to 1 may specify that the luma coding tree block size of each coding tree unit may be equal to 32×32, and that use_smallest_ctu_size_flag is equal to 0 may specify that use_largest_ctu_size_flag may exist.
[0048] That use_largest_ctu_size_flag is equal to 1 may specify that the luma coding tree block size of each coding tree unit may be equal to 128×128, and that use_largest_ctu_size_flag is equal to 0 may specify that the luma coding tree block size of each coding tree unit may be equal to 64×64.
[0049] Adding 2 to log2_min_luma_coding_block_size_minus2 may specify the minimum luma coding block size.
[0050] That sps_max_luma_transform_size_64_flag is equal to 1 may specify that the maximum transform size in luma samples may be equal to 64, and that sps_max_luma_transform_size_64_flag is equal to 0 may specify that the maximum transform size in luma samples may be equal to 32. If it does not exist, the value of sps_max_luma_transform_size_64_flag may be presumed to be equal to 0.
[0051] The variables CtbLog2SizeY, CtbSizeY, MinCbLog2SizeY, MinCbSizeY, MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSizeY, MaxTbSizeY, PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC, and PicHeightInSamplesC may be derived as follows: if(use_smallest_ctu_size_flag) CtbLog2SizeY = 0 else if(use_largest_ctu_size_flag) CtbLog2SizeY = 2 else CtbLog2SizeY = 1 CtbSizeY = 1 << CtbLog2SizeY
[0052] Referring now to FIG. 4, an operational flowchart showing the steps of a method 400 for encoding video data is shown. In some implementations, one or more of the process blocks of FIG. 4 may be performed by computer 102 (FIG. 1) and server computer 114 (FIG. 1). In some implementations, one or more of the process blocks of FIG. 4 may be performed by another device or group of devices separate from, or including, computer 102 and server computer 114.
[0053] At 402, method 400 includes receiving video data having a certain coding tree unit size.
[0054] At 404, method 400 includes signaling the coding tree unit size associated with the video data by setting two or more flags.
[0055] In 406, method 400 includes encoding video data based on a flag corresponding to the signaled coded tree unit size.
[0056] FIG. 4 only provides an illustration of one implementation and should be understood not to imply any limitation as to how various embodiments may be implemented. Many modifications to the illustrated environment may be made based on design and implementation requirements.
[0057] FIG. 5 is a block diagram 500 of the internal and external components of the computer shown in FIG. 1 according to an exemplary embodiment. FIG. 5 only gives a diagram of one implementation and should be understood not to imply any limitation regarding the environments in which different embodiments may be implemented. Many modifications to the illustrated environment may be made based on design and implementation requirements.
[0058] Computer 102 (FIG. 1) and server computer 114 (FIG. 1) may each include a respective set of internal components 800A, B and external components 900A, B shown in FIG. 4. Each set of internal components 800 includes one or more processors 820, one or more computer-readable RAMs 822 and one or more computer-readable ROMs 824 on one or more buses 826, one or more operating systems 828, and one or more tangible computer-readable storage devices 830.
[0059] Processor 820 is implemented in hardware, firmware, or a combination of hardware and software. Processor 820 is a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or another type of processing component. In some implementations, Processor 820 includes one or more processors that can be programmed to perform functions. Bus 826 includes components that enable communication between internal components 800A, B.
[0060] One or more operating systems 828, software program 108 (FIG. 1), and video encoding program 116 (FIG. 1) on server computer 114 (FIG. 1) are stored in one or more of respective computer-readable tangible storage devices 830 for execution by one or more of respective processors 820 via one or more of respective RAMs 822. In the embodiment shown in FIG. 5, each of computer-readable tangible storage devices 830 is a magnetic disk storage device of an internal hard drive. Alternatively, each of computer-readable tangible storage devices 830 is a semiconductor storage device, such as ROM 824, EPROM, flash memory, optical disk, magneto-optical disk, solid state disk, compact disk (CD), digital versatile disk (DVD), floppy disk, cartridge, magnetic tape, and / or other types of non-transitory computer-readable tangible storage devices capable of storing computer programs and digital information.
[0061] Each set of internal components 800A, B also includes an R / W drive or interface 832 for reading and writing to one or more portable computer-readable tangible storage devices 936 such as CD-ROMs, DVDs, memory sticks, magnetic tapes, magnetic disks, optical disks, or semiconductor memory devices. Software programs such as software program 108 (FIG. 1) and video encoding program 116 (FIG. 1) are stored on one or more of the respective portable computer-readable tangible storage devices 936, read via the respective R / W drive or interface 832, and can be loaded onto the respective hard drives 830.
[0062] Each set of internal components 800A, B also includes a network adapter or interface 836 such as a TCP / IP adapter card, a wireless Wi-Fi interface card, or a 3G, 4G, or 5G wireless interface card, or other wired or wireless communication link. The software program 108 (FIG. 1) and the video encoding program 116 (FIG. 1) on the server computer 114 can be downloaded from an external computer to the computer 102 (FIG. 1) and the server computer 114 via a network (e.g., the Internet, a local area network, or other wide area network) and the respective network adapter or interface 836. From the network adapter or interface 836, the software program 108 and the video encoding program 116 on the server computer 114 are loaded onto the respective hard drives 830. The network can include copper wire, fiber optic, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers.
[0063] Each of the sets of external components 900A, B can include a computer display monitor 920, a keyboard 930, and a computer mouse 934. The external components 900A, B can also include a touch screen, a virtual keyboard, a touch pad, a pointing device, and other human interface devices. Each of the sets of internal components 800A, B also includes a device driver 840 for interfacing with the computer display monitor 920, the keyboard 930, and the computer mouse 934. The device driver 840, the R / W drive or interface 832, and the network adapter or interface 836 include hardware and software (stored in the storage device 830 and / or the ROM 824).
[0064] The present disclosure includes a detailed description regarding cloud computing, but it is to be understood in advance that the implementation of the teachings described herein is not limited to a cloud computing environment. Rather, some embodiments can be implemented in connection with any other type of computing environment, whether currently known or later developed.
[0065] Cloud computing is a service delivery model that enables convenient on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, services) that can be rapidly provisioned and released with minimal management effort or interaction with a service provider. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.
[0066] The characteristics are as follows: On - demand self - service: Cloud consumers can automatically and unilaterally provision computing capabilities such as server time and network storage as needed, without the need for human interaction with the service provider. Broad network access: The capabilities are available over a network and accessed through standard mechanisms. This facilitates use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs). Resource pooling: The provider's computing resources are pooled to serve multiple consumers using a multi - tenant model, and various physical and virtual resources are dynamically assigned and re - assigned according to demand. Consumers generally have no control or knowledge over the exact location of the provided resources, but have a sense of location independence in that they may be able to specify the location at a higher level of abstraction (e.g., country, state, data center). Rapid elasticity: The capabilities are provisioned rapidly and elastically, in some cases automatically, to scale out quickly and released quickly to scale in. To the consumer, the capabilities available for provisioning often appear to be unlimited, and they can purchase any amount at any time. Measured service: The cloud system automatically controls and optimizes resource use by leveraging a metering function at an appropriate level of abstraction for the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage is monitored, controlled, reported, and provides transparency for both the provider and consumer of the service being utilized.
[0067] The service model is as follows: Software as a Service (SaaS): The function provided to consumers is to use the provider's applications running on cloud infrastructure. The applications are accessible from various client devices through a client interface such as a web browser (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, which includes the network, servers, operating systems, storage, or even individual application functions. However, limited, user-specific application configuration settings may be an exception. Platform as a Service (PaaS): The function provided to consumers is to deploy the applications created or obtained by consumers, which are created using the programming languages and tools supported by the provider, onto the cloud infrastructure. Consumers do not manage or control the underlying cloud infrastructure including the network, servers, operating systems, and storage, but have control over the deployed applications and possibly the configuration of the application hosting environment. Infrastructure as a Service (IaaS): The function provided to consumers is to provision processing, storage, network, and other basic computing resources. Here, consumers can deploy and run any software that can include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but have control over the operating system, storage, the deployed applications, and limited control over the selected network components (e.g., host firewall).
[0068] The deployment models are as follows: Private cloud: The cloud infrastructure is operated solely for an organization. This may be managed by the organization or a third party and can be located either on - premise or off - premise. Community cloud: The cloud infrastructure is shared by several organizations and supports a specific community with shared concerns (e.g., mission, security requirements, policies, and compliance considerations). This may be managed by the organization or a third party and can be located either on - premise or off - premise. Public cloud: The cloud infrastructure is made available to the general public or a large industry group and is owned by an organization that sells cloud services. Hybrid cloud: The cloud infrastructure is a combination of two or more clouds (private, community, or public), which remain separate entities but are linked by standardized or proprietary technologies (e.g., cloud bursting for load balancing between clouds) that enable data and application portability.
[0069] Cloud computing environments are service - oriented, focusing on statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is an infrastructure consisting of a network of interconnected nodes.
[0070] Referring to FIG. 6, an exemplary cloud computing environment 600 is shown. As shown in the figure, the cloud computing environment 600 includes one or more cloud computing nodes 10, and local computing devices used by cloud consumers, such as a personal digital assistant (PDA) or mobile phone 54A, desktop computer 54B, laptop computer 54C, and / or automotive computer system 54N, can communicate with them. The cloud computing nodes 10 can communicate with each other. They may be grouped (not shown) physically or virtually into one or more networks such as the private, community, public, or hybrid clouds as described above or combinations thereof. Thereby, the cloud computing environment 600 can provide infrastructure, platform, and / or software as a service such that cloud consumers do not need to maintain resources on local computing devices. The types of computing devices 54A - N shown in FIG. 5 are merely intended to be exemplary, and it is understood that the cloud computing nodes 10 and the cloud computing environment 600 can communicate with any type of computerized device through any type of network and / or network addressable connection (e.g., using a web browser).
[0071] Referring to FIG. 7, a set of functional abstraction layers 700 provided by the cloud computing environment 600 (FIG. 6) is shown. It should be understood that the components, layers, and functions shown in FIG. 7 are merely intended to be exemplary, and embodiments are not limited thereto. As shown in the figure, the following layers and corresponding functions are provided.
[0072] The hardware and software layer 60 includes hardware and software components. Examples of hardware components are: mainframe 61; RISC (Reduced Instruction Set Computer) architecture-based server 62; server 63; blade server 64; storage device 65; and network and network components 66. In some embodiments, the software components include network application server software 67 and database software 68.
[0073] From the abstraction layer provided by the virtualization layer 70, the following examples of virtual entities may be provided: virtual server 71; virtual memory 72; virtual network 73 including a virtual private network; virtual applications and operating systems 74; and virtual client 75.
[0074] In one example, the management layer 80 can provide the functions described below. Resource provisioning 81 provides for the dynamic procurement of computing resources and other resources utilized to execute tasks within a cloud computing environment. Metering and Pricing 82 provides for cost tracking when resources are utilized within a cloud computing environment and for billing or issuing an itemized statement for the consumption of these resources. In one example, these resources may include application software licenses. Security provides for the authentication of cloud consumers and tasks and for the protection of data and other resources. The user portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 84 provides for the allocation and management of cloud computing resources such that the required service levels are met. Service Level Agreement (SLA) formulation and fulfillment 85 provides for the advance arrangement and procurement of cloud computing resources where future requirements are predicted according to the SLA.
[0075] The workload layer 90 provides examples of functions for which a cloud computing environment can be utilized. Examples of workloads and functions that can be provided from this layer include: mapping and navigation 91; software development and life cycle management 92; virtual classroom education delivery 93; data analysis processing 94; transaction processing 95; and video encoding 96. Video encoding 96 may encode video data using a separate flag that replaces the coded tree unit size syntax.
[0076] Some embodiments can relate to systems, methods, and / or computer-readable media at any possible technical detail level of integration. A computer-readable medium can include a computer-readable non-transitory storage medium (or media) having thereon computer-readable program instructions for causing a processor to execute operations.
[0077] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disks (DVD), memory sticks, floppy disks, punch cards, or mechanically encoded devices such as a raised structure in a groove having instructions recorded thereon, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as being a transient signal per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., optical pulses passing through an optical fiber cable), or electrical signals transmitted through a wire.
[0078] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded from an external computer or an external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each respective computing / processing device.
[0079] The computer-readable program code / instructions for performing the calculations may be in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in an object-oriented programming language such as Smalltalk, C++, and a procedural programming language such as the "C" programming language or a similar programming language. The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on a remote computer or server. In this last scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit to perform aspects or operations.
[0080] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to generate a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / steps specified in the block(s) of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to function in a particular manner, thereby, the computer-readable storage medium storing the instructions contains a manufacture including instructions for implementing aspects of the functions / steps specified in the block(s) of a flowchart and / or block diagram.
[0081] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process. Thereby, the instructions executed on the computer, other programmable apparatus, or other devices implement the functions / steps specified in the block(s) of a flowchart and / or block diagram.
[0082] Flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of one or more executable instructions for implementing the specified logical function(s). The methods, computer systems, and computer-readable media may include additional blocks, fewer blocks, different blocks, or differently arranged blocks than those shown in the drawings. In some alternative implementations, the functions recited in the blocks may occur out of the order recited in the drawings. For example, two blocks shown in succession may, in fact, be executed simultaneously or substantially simultaneously, or the blocks may be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or combinations of special-purpose hardware and computer instructions.
[0083] It will be apparent that the systems and / or methods described herein can be implemented in various forms of hardware, firmware, or a combination of hardware and software. The actual specialized control hardware or software code used to implement these systems and / or methods is not limiting of the implementation. Thus, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, and it is understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0084] No element, step, or instruction used in this specification should be construed as decisive or essential unless explicitly stated. Also, as used in this specification, the articles "a" and "an" are intended to include one or more items and can be used interchangeably with "one or more". Further, as used in this specification, the term "set" is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and can be used interchangeably with "one or more". When only one item is intended, the term "one" or similar language is used. Also, as used in this specification, terms such as "have", "possess", "having", etc. are intended to be open-ended terms. Further, the phrase "based on" is intended to mean "at least in part based on" unless otherwise specified.
[0085] Descriptions of various aspects and embodiments have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the disclosed embodiments. Even if combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features may be combined in ways not specifically recited in the claims and / or disclosed in the specification. Each of the dependent claims listed below may directly depend only on one claim, but the disclosure of possible implementations includes combinations of each dependent claim with all other claims in the claims. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terms used in this specification are chosen in order to best explain the principles of the embodiments, the practical application, or technical improvements found in the marketplace, or to enable other skilled artisans to understand the embodiments disclosed in this specification.
Claims
1. A video encoding method performed by an encoder, comprising: generating encoded video data; and transmitting the encoded video data; the encoded video data includes a flag corresponding to a coding tree unit size that is signaled; Generating the encoded video data includes: receiving video data having a coding tree unit size; signaling the coding tree unit size associated with the video data by setting an adaptive one or two flags; and encoding the video data based on the flag corresponding to a signaled coding tree unit size. method.
2. The method described in claim 1, wherein the adaptive one or two flags include two of a 32 pixel coding tree unit size flag, a 64 pixel coding tree unit size flag, and a 128 pixel coding tree unit size flag.
3. The method described in claim 2, wherein the 32 pixel coding tree unit size flag is signaled first, and the 64 pixel coding tree unit size flag is signaled based on the 32 pixel coding tree unit size flag not being equal to 1.
4. The method described in claim 2, wherein the 64 pixel coding tree unit size flag is signaled first, and the 32 pixel coding tree unit size flag is signaled based on the 64 pixel coding tree unit size flag not being equal to 1.
5. The method described in claim 2, wherein the 32 pixel coding tree unit size flag is signaled first, and the 128 pixel coding tree unit size flag is signaled based on the 32 pixel coding tree unit size flag not being equal to 1.
6. The method described in claim 2, wherein the 128 pixel coding tree unit size flag is signaled first, and the 32 pixel coding tree unit size flag is signaled based on the 128 pixel coding tree unit size flag not being equal to 1.
7. The method described in claim 2, wherein the 64 pixel coding tree unit size flag is signaled first, and the 128 pixel coding tree unit size flag is signaled based on the 64 pixel coding tree unit size flag not being equal to 1.
8. The method described in claim 2, wherein the 128 pixel coding tree unit size flag is signaled first, and the 64 pixel coding tree unit size flag is signaled based on the 128 pixel coding tree unit size flag not being equal to 1.
9. The method of claim 2, wherein a 64 pixel maximum luma transform size flag is signaled based on the coding tree unit size associated with the video data being 64 pixels by 64 pixels or greater.
10. The method of claim 9, wherein the 64 pixel maximum luma transform size flag is not signaled based on the 32 pixel coding tree unit size flag being initially signaled and equal to 1.
11. The method described in claim 9, wherein based on the 64 pixel coding tree unit size flag being initially signaled and equal to 0, the 32 pixel coding tree unit size flag is signaled and set equal to 1, and the 64 pixel maximum luma transform size flag is not signaled.
12. The method described in claim 9, wherein based on the 128 pixel coding tree unit size flag being initially signaled and equal to 0, the 32 pixel coding tree unit size flag is signaled and set equal to 1, and the 64 pixel maximum luma transform size flag is not signaled.
13. The method described in claim 1, wherein a first flag and a second flag of the adaptive one or two flags correspond to a sequence parameter set associated with the video data, the first flag indicating whether a minimum coding tree unit size is applied, and the second flag indicating whether a maximum coding tree unit size is applied.
14. The method of claim 13, wherein the first flag is signaled first and the second flag is signaled based on the minimum coding tree unit size not being applied.
15. The method of claim 13, wherein a second flag is signaled first and the first flag is signaled based on the maximum coding tree unit size not being applied.
16. A method according to any one of claims 13 to 15, wherein a 64 pixel maximum luma transform size flag is signalled on the basis that the minimum coding tree unit size is not applied.
17. A video encoding method performed by an encoder, comprising: generating encoded video data; and transmitting the encoded video data; the encoded video data includes a flag signaled in a sequence parameter set or a slice header corresponding to a coding tree unit size; Generating the encoded video data includes: receiving video data having a coding tree unit size; signaling a flag corresponding to the coding tree unit size in a sequence parameter set or a slice header; The signaling step includes: signaling a first flag of the flags, the first flag corresponding to a first coding tree unit size; The video encoding method includes: if the first flag is one, encoding the video data based on the first coding tree unit size without signaling a second one of the flags; signaling a second one of the flags if the first flag is 0, and encoding the video data based on the signaled second flag; if the second flag is 1, the video data is encoded based on a second coding tree unit size, and if the second flag is 0, the video data is encoded based on a third coding tree unit size. method.
18. The method of claim 17, wherein the flags include two of a 32 pixel coding tree unit size flag, a 64 pixel coding tree unit size flag, and a 128 pixel coding tree unit size flag.
19. The method described in claim 18, wherein the 32 pixel coding tree unit size flag is signaled first, and the 64 pixel coding tree unit size flag is signaled based on the 32 pixel coding tree unit size flag not being equal to 1.
20. The method described in claim 18, wherein the 64 pixel coding tree unit size flag is signaled first, and the 32 pixel coding tree unit size flag is signaled based on the 64 pixel coding tree unit size flag not being equal to 1.
21. The method described in claim 18, wherein the 32 pixel coding tree unit size flag is signaled first, and the 128 pixel coding tree unit size flag is signaled based on the 32 pixel coding tree unit size flag not being equal to 1.
22. The method described in claim 18, wherein the 128 pixel coding tree unit size flag is signaled first, and the 32 pixel coding tree unit size flag is signaled based on the 128 pixel coding tree unit size flag not being equal to 1.
23. The method described in claim 18, wherein the 64 pixel coding tree unit size flag is signaled first, and the 128 pixel coding tree unit size flag is signaled based on the 64 pixel coding tree unit size flag not being equal to 1.
24. The method described in claim 18, wherein the 128 pixel coding tree unit size flag is signaled first, and the 64 pixel coding tree unit size flag is signaled based on the 128 pixel coding tree unit size flag not being equal to 1.
25. The method of claim 18, wherein a 64 pixel maximum luma transform size flag is signaled based on the coding tree unit size associated with the video data being 64 pixels by 64 pixels or greater.
26. The method described in claim 25, wherein the 64 pixel maximum luma transform size flag is not signaled based on the 32 pixel coding tree unit size flag being initially signaled and equal to 1.
27. The method described in claim 25, wherein based on the 64 pixel coding tree unit size flag being initially signaled and equal to 0, the 32 pixel coding tree unit size flag is signaled and set equal to 1, and the 64 pixel maximum luma transform size flag is not signaled.
28. The method described in claim 25, wherein based on the 128 pixel coding tree unit size flag being initially signaled and equal to 0, the 32 pixel coding tree unit size flag is signaled and set equal to 1, and the 64 pixel maximum luma transform size flag is not signaled.
29. The method of claim 17, wherein the first flag indicates whether a minimum coding tree unit size is applied, and the second flag indicates whether a maximum coding tree unit size is applied.
30. The method of claim 29, wherein the first flag is signaled first and the second flag is signaled based on the minimum coding tree unit size not being applied.
31. The method described in claim 17, wherein the second flag indicates whether a minimum coding tree unit size is applied and the first flag indicates whether a maximum coding tree unit size is applied.
32. The method of claim 29, wherein a 64 pixel maximum luma transform size flag is signaled based on the minimum coding tree unit size not being applied.
Citation Information
Patent Citations
Methods and apparatuses for encoding / decoding high resolution images
US20150208091A1
Method of Video Coding Using Separate Coding Tree for Luma and Chroma
US20180288446A1