Shape Adaptive Discrete Cosine Transform with Geometric Partitioning Having Region Number Adaptivity
By adopting shape adaptive discrete cosine transform (SA-DCT) and geometric division mode in video compression technology, the transformation type is selected based on the prediction error in the block, which solves the problems of low encoding efficiency and waste of computing resources in the prior art, and achieves more efficient video encoding and decoding.
Patent Information
- Application Number
- CN202080022269.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-01-28
- Filing Date
- 2020-01-28
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-01-28
AI Technical Summary
When the existing video compression technology processes non-rectangular blocks, conventional blocked discrete cosine transformation (B-DCT) is inefficient and wastes computing resources, making it difficult to effectively represent pixel information in non-rectangular areas.
Shape adaptive discrete cosine transform (SA-DCT) combined with geometric division mode is used to select the transformation type according to the prediction error within the block, select B-DCT or SA-DCT to more effectively represent the residual information, and reduce the consumption of computing resources.
Through the combination of SA-DCT and geometric division mode, the complexity and processing performance of video encoding and decoding are improved, the transmission bit rate of the bit stream is reduced, and the consumption of computing resources is reduced.
Smart Images

Figure CN113597757B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Application Serial No. 62 / 797,799, filed on January 28, 2019, entitled “SHAPE ADAPTIVE DISCRETE COSINE TRANSFORM FORGEOMETRIC PARTITIONING WITH AN ADAPTIVE NUMBER OF REGIONS,” the entire contents of which are hereby incorporated by reference into this application. Technical Field
[0003] The present invention generally relates to the field of video compression. In particular, the present invention relates to shape adaptive discrete cosine transform for geometric partitioning with region number adaptation. Background Art
[0004] A video codec may include electronic circuits or software that compresses or decompresses digital video. The video codec can convert uncompressed video to a compressed format and vice versa. In the field of video compression, a device that compresses video (and / or performs some of its functions) is often called an encoder, while a device that decompresses video (and / or performs some of its functions) is called a decoder.
[0005] The compressed data may be in a format that conforms to standard video compression specifications. Compression may also be lossy in that some information from the original video is lost in the compressed video. This may result in the decompressed video not having enough information to accurately reconstruct the original video, making it lower quality than the original uncompressed video.
[0006] Video quality has a complex relationship with the amount of data used to characterize the video (e.g., determined by the bit rate), the complexity of the encoding and decoding algorithms, susceptibility to data loss and errors, ease of editing, random access, end-to-end latency (e.g., delay), etc. Summary of the invention
[0007] In one aspect, a decoder includes: a circuit configured to receive a bitstream; determine a first region, a second region, and a third region of a current block based on a geometric partitioning pattern; and use an inverse discrete cosine transform on each of the first region, the second region, and the third region to decode the current block.
[0008] On the other hand, the decoder includes: a circuit configured to receive a bitstream; determine a first region, a second region, and a third region of a current block based on a geometric partitioning pattern; determine a coding transformation type based on a signal contained in the bitstream to decode each of the first region, the second region, and / or the third region, the coding transformation type characterizing at least a discrete cosine inverse transform and a shape adaptive discrete cosine inverse transform of the block; and decode the current block, decoding the current block includes inversely transforming each of the first region, the second region, and / or the third region using the determined transformation type.
[0009] On the other hand, a method includes: receiving a bitstream by a decoder; determining a first region, a second region, and a third region of a current block according to a geometric partitioning pattern; determining a coding transformation type according to a signal contained in the bitstream to decode the first region, the second region, and / or the third region, the coding transformation type at least characterizing an inverse discrete cosine transform and a shape adaptive inverse discrete cosine transform of the block; and decoding the current block, decoding the current block includes inversely transforming each of the first region, the second region, and / or the third region using the determined transformation type.
[0010] One or more variations of the subject matter described herein are described in detail in the following drawings and description. Other features and advantages of the subject matter described herein will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] For the purpose of illustrating the present invention, aspects of one or more embodiments of the present invention are shown in the accompanying drawings. However, it should be understood that the present invention is not limited to the precise configuration and apparatus shown in the accompanying drawings, in which:
[0012] Figure 1 is a diagram of an example of an exponentially partitioned residual block (eg, a current block), wherein there are three segments with different prediction errors;
[0013] Figure 2 is a system block diagram of an exemplary video encoder capable of shape adaptive discrete cosine transformation (SA-DCT), wherein the SA-DCT is used for geometric partitioning with adaptive number of regions, which can improve the complexity and processing performance of video encoding and decoding;
[0014] Figure 3 is a process flow chart illustrating an exemplary process of encoding a video using SA-DCT for geometric partitioning with adaptive number of regions;
[0015] Figure 4is a system block diagram illustrating an exemplary decoder capable of decoding a bitstream using SA-DCT for geometric partitioning with adaptation of the number of regions;
[0016] Figure 5 is a process flow diagram illustrating an exemplary process of decoding a bitstream using SA-DCT for geometric partitioning with region number adaptation; and
[0017] Figure 6 is a block diagram of a computing system that can be used to implement any one or more of the methods of the present disclosure, and any one or more portions thereof.
[0018] The accompanying drawings are not necessarily drawn to scale, but are shown in dotted lines, schematic diagrams and partial views. In some cases, details that are not necessary for understanding the embodiments or that make other details difficult to understand are omitted. In the various drawings, the same reference numerals represent the same components. DETAILED DESCRIPTION
[0019] Embodiments in the present disclosure relate to encoding and decoding blocks in geometric partitions, where not all blocks are necessarily rectangular. Embodiments may include and / or be configured to perform encoding and / or decoding using discrete cosine transform (DCT) and / or inverse DCT. In some embodiments of the present disclosure, DCT is selected based on the information content in the geometric partitioned blocks. In some existing video encoding and decoding methods, all blocks are rectangular, and conventional block discrete cosine transform (Block DCT, B-DCT) is used to encode residual blocks for the entire rectangular block. However, in blocks where the geometric partitioning can be divided into multiple non-rectangular regions, the use of conventional B-DCT will inefficiently represent the potential pixel information of some blocks and may require unnecessary computing resources to perform. In some embodiments of the current inventive subject matter, when using a geometric partitioning mode, the encoder may use a shape adaptive discrete cosine transform (SA-DCT) instead of B-DCT or in addition to B-DCT. In some embodiments, the encoder may select between B-DCT and SA-DCT based on the prediction error level of each region of the block (e.g., geometric partition block); the selection may be signaled in the bitstream for decoding. The non-rectangular region may be encoded and / or decoded using either B-DCT or SA-DCT and the selection may be signaled; since the residual may be more efficiently represented, the transmission bit rate in the bitstream may be reduced, thereby reducing the computational resources required to perform the processing. The current inventive subject matter may be applicable to relatively large blocks, e.g., blocks of size 128×128 or 64×64. In some embodiments, the geometric partitioning may include partitioning the current block into an adaptive number of regions, e.g., three or more regions of the current specified block; the type of DCT transform (e.g., B-DCT or SA-DCT) for each region may be signaled.
[0020] In an embodiment, the B-DCT may be a DCT performed on a NxN block of values using an NxN reversible matrix, where the NxN blocks of values are, for example, but not limited to, chrominance and / or luminance values of a corresponding NxN pixel array. For example, in a non-limiting example, if the NxN matrix X is transformed, the "DCT-I" transform may calculate each element of the transform matrix as follows:
[0021]
[0022] Where k = 0, ..., N-1. In another non-limiting example, the "DCT-II" transform can calculate the transformed matrix as follows:
[0023]
[0024] Wherein, k=0, ..., N-1. In an illustrative example, if the pixel block size of the partition is 4×4, the generalized discrete cosine transform matrix may include a generalized discrete cosine transform matrix II of the following form:
[0025]
[0026] where a is 1 / 2 and b is And c is
[0027] In some embodiments, an integer approximation algorithm of the transformation matrix may be used, which may be implemented with efficient hardware and software. For example, in the case where the pixel block size of the partition is 4×4, the generalized discrete cosine transform matrix may include a generalized discrete cosine transform II matrix of the following form:
[0028]
[0029] The inverse B-DCT transform can be calculated using the same NxN transform matrix through a second matrix multiplication; the output result can be normalized to restore the original value. For example, the inverse DCT-I transform can be multiplied by 2 / (N-1) for normalization.
[0030] The SA-DCT may be performed on a non-rectangular pixel array. In an embodiment, the SA-DCT may be calculated by performing a one-dimensional version of a DCT (e.g., DCT-I, DCT-II) or similar transform on a vertical column vector representing pixel values in a shape of interest, and then grouping the resulting values into horizontal vectors and performing a one-dimensional DCT again; the second DCT may fully transform the pixel values. The SA-DCT variables may be further scaled and / or normalized by coefficients to correct for weighted average defects and / or non-standard orthogonal defects caused by the above-mentioned transformation, quantization of the above-mentioned transformation output, and / or inverse transformation of the transformation output and / or quantized transformation output. Prior to performing the above-mentioned SA-DCT process, the average single value of the subject image area may be subtracted from each pixel value or a scaled version thereof, and may require one or other of the scaling processes to be applied before and / or after the transformation, quantization, and / or inverse transformation to further correct. A person skilled in the art will recognize, after reading the entire contents of this disclosure, that various alternative or additional variations of the SA-DCT process may be applied as described above.
[0031] Motion compensation may include methods for predicting a video frame or a portion thereof given a previous frame and / or future frame by calculating the motion of a camera and / or objects in the video including the current frame, previous frame and / or future frame, and / or represented by the current frame, previous frame and / or future frame. Motion compensation may be used for encoding and decoding of video data for video compression, for example, encoding and decoding using the Moving Picture Experts Group-2 (MPEG-2) (also known as Advanced Video Coding (AVC)) standard. Motion compensation may describe a picture based on a transformation from a reference picture to a current picture. The reference picture may be a previous or future picture in time when compared to the current picture. Compression efficiency may be improved when an image may be accurately synthesized from previously transmitted and / or stored images.
[0032] Block partitioning as used in the present disclosure refers to a method of finding similar motion areas in video coding. Some form of block partitioning can be found in video codec standards, including MPEG-2, H.264 (also known as AVC or MPEG-4 Part 10), and H.265 (also known as High Efficiency Video Coding (HEVC)). In an exemplary block partitioning method, non-overlapping blocks of a video frame can be divided into rectangular sub-blocks to find block partitions containing similar motion pixels. The method is effective when all pixels of the block partition have similar motion. The motion of the pixels in the block can be determined relative to the previously encoded frame.
[0033] Shape adaptive DCT and / or B-DCT can be effectively used for geometric partitioning with adaptation of the number of regions. Figure 1is a diagram of a non-limiting example of a residual block (eg, current block) 100 of size 64×64 or 128×128 by geometric partitioning, wherein there are three segments S0, S1, and S2 with different prediction errors. Figure 1 Three segments are shown for example purposes, but more or fewer segments may be used alternatively or additionally. The current block may be geometrically divided according to two line segments (P1P2 and P3P4) into three regions S0, S1, and S2. In this example, the prediction error of S0 is relatively high, while the prediction errors of S1 and S2 are relatively low. For segment S0 (also referred to as a region), the encoder may select and use B-DCT for residual encoding. For segments S1 and S2 with small prediction errors, the encoder may select and use SA-DCT. The residual coding transform may be selected based on the prediction error (e.g., residual size). Since the SA-DCT algorithm is relatively simpler in complexity and does not require many calculations like B-DCT, encoding the lower prediction error residual using SA-DCT may improve the complexity and processing performance of video encoding and decoding.
[0034] Therefore, still refer to Figure 1 , for segments with low prediction error, SA-DCT may be signaled as an additional transform selection to the full block DCT. The error considered low or high is a parameter that can be set on the encoder and may vary from application to application. The selection of the transform type may be signaled in the bitstream. In the decoder, the bitstream is parsed and for a given current block, the residual may be decoded using the transform type signaled in the bitstream. In some embodiments, multiple coefficients associated with the transform may alternatively or additionally be signaled in the bitstream.
[0035] Specifically, and with reference to Figure 1 , geometric partitioning with adaptive number of regions may include video encoding and decoding techniques, wherein a rectangular block is further divided into two or more regions that may be non-rectangular. For example, Figure 1A non-limiting example of pixel-level geometric partitioning with region number adaptation is shown in FIG. An exemplary rectangular block 100 (which has a width of M pixels and a height of N pixels, expressed as M×N pixels) can be divided into three regions (S0, S1, and S2) along line segments P1P2 and P3P4. When the motion of pixels in S0 is similar, a single motion vector can be used to describe the motion of all pixels in the region. This motion vector can be used to compress region S0. Similarly, when the motion of pixels in region S1 is similar, the motion of pixels in region S1 can be described by an associated motion vector. Similarly, when the motion of pixels in region S2 is similar, the motion of pixels in region S2 can be described by an associated motion vector. Such geometric partitioning can be signaled to a receiver (e.g., a decoder) in a video bitstream by encoding positions P1, P2, P3, P4 and / or representations of these positions, for example, including but not limited to using coordinates (e.g., polar coordinates, Cartesian coordinates, etc.) and indices, into a predefined template, or encoding other partitioning features.
[0036] Still refer to Figure 1 , when encoding video data using pixel-level geometric partitioning, the line segment P1P2 (or more specifically, the point P1 and the point P2) can be determined. In order to determine the line segment P1P2 (or more specifically, the point P1 and the point P2) that can best divide the block when using pixel-level geometric partitioning, possible combinations of the points P1 and P2 can be made according to M and N (block width and block height). For a block of size M×N, there are (M-1)×(N-1)×3 possible partitioning methods. Therefore, identifying the correct partitioning method becomes a costly computational task that requires evaluating motion estimates for all possible partitions. Compared to encoding using rectangular partitioning (e.g., no pixel-level geometric partitioning), this increases the time required to encode the video and / or requires increased processing power. The area that constitutes the best or correct partition can be determined based on a metric and can vary with implementation.
[0037] In some embodiments, and still with reference to Figure 1 , iterative partitioning can be performed. In iterative partitioning, a first partition forming two regions (e.g., the determination line P1P2 and the associated region) can be determined, and then one of these regions can be further divided. For example, performing a combination Figure 1 The described partitioning may divide the block into two regions. One of these regions may be further partitioned (e.g., to form new regions S1 and S2). The process may continue to perform block-level geometric partitioning until a stopping criterion is reached.
[0038] Figure 22 is a system block diagram showing an exemplary video encoder capable of performing SA-DCT and / or B-DCT for geometric partitioning with adaptive number of regions, which can improve the complexity and processing performance of video encoding and decoding. The exemplary video encoder 200 receives an input video 205, which can be initially partitioned or divided according to a processing scheme, such as a tree-structured macroblock partitioning scheme (e.g., quad-tree plus binary tree). An example of a tree-structured macroblock partitioning scheme may include partitioning a picture frame into large block elements called coding tree units (CTUs). In some embodiments, each CTU may be further divided into multiple sub-blocks called coding units (CUs) one or more times. The final result of this partitioning may include a group of sub-blocks called predictive units (PUs). Transform units (TUs) may also be used. Such a partitioning scheme may include performing geometric partitioning with adaptive number of regions according to some aspects of the current inventive subject matter.
[0039] Continue to refer to Figure 2 , the exemplary video encoder 200 includes an intra prediction processor 215, a motion estimation / compensation processor 220 (also referred to as an inter prediction processor) capable of supporting geometric partitioning with region number adaptation, a transform / quantization processor 225, an inverse quantization / inverse transform processor 230, a loop filter 235, a decoded picture buffer 240, and an entropy coding processor 245. In some embodiments, the motion estimation / compensation processor 220 may perform geometric partitioning. Bitstream parameters that signal the geometric partitioning mode may be input to the entropy coding processor 245 for inclusion in the output bitstream 250.
[0040] In operation, and continue to refer to Figure 2 For each frame block of the input video 205, it may be determined whether the block is processed by intra-picture prediction or using motion estimation / compensation. The block may be provided to an intra-prediction processor 210 or a motion estimation / compensation processor 220. If the block is processed by intra-prediction, the intra-prediction processor 210 may perform a process to output a prediction value. If the block is processed by motion estimation / compensation, the motion estimation and compensation processor 220 may perform a process (including using geometric partitioning) to output a prediction value.
[0041] Still refer to Figure 2, a residual may be formed by subtracting the predicted value from the input video. The residual may be received by a transform / quantization processor 225, which may determine whether the prediction error (e.g., the size of the residual) is considered to be a "high" error or a "low" error (e.g., by comparing the size of the residual or an error measure to a threshold). Based on the determination, the transform / quantization processor 225 may select a transform type (including B-DCT and SA-DCT). In some embodiments, the transform / quantization processor 225 selects a B-DCT transform type when the residual is considered to have a high error, and selects an SA-DCT transform type when the residual is considered to have a low error. Based on the selected transform type, the transform / quantization processor 225 may perform a transform process (e.g., SA-DCT or B-DCT) to generate quantizable coefficients. The quantized coefficients and any associated signaling information (including the selected transform type and / or the number of coefficients used) may be provided to the entropy coding processor 245 for entropy coding and inclusion in the output bitstream 250. The entropy coding processor 245 may support the coding of signaling information related to the SA-DCT for geometric partitioning with adaptive number of regions. In addition, the quantized coefficients may be provided to the inverse quantization / inverse transform processor 230 which may reproduce the pixels, where the pixels are combined with the prediction values and processed by the loop filter 235, the output of which is stored in the decoded picture buffer 240 for use by the motion estimation / compensation processor 220 which may support geometric partitioning with adaptive number of regions.
[0042] Now refer to Figure 3 , a process flow diagram of an exemplary process 300 for encoding a video using SA-DCT for geometric partitioning with adaptive number of regions, wherein the geometric partitioning can improve the complexity and processing performance of video encoding and decoding. In step 310, an initial block partitioning is performed on the video frame, for example, using a tree-structured macroblock partitioning scheme, including partitioning the picture frame into CTUs and CUs. In 320, blocks are selected for geometric partitioning. The selection includes identifying blocks to be processed in a geometric partitioning mode based on a metric rule. In step 330, the selected block is partitioned into three or more non-rectangular regions based on the geometric partitioning mode.
[0043] In step 340, and still referring to Figure 3, determine the transform (transformation or transformation) type for each geometric partition area. This may include determining whether the prediction error (e.g., the size of the residual) is considered to be a "high" error or a "low" error (e.g., by comparing the size of the residual or the error measure to a threshold). Based on the determination, a transform type is selected, for example, using a quadtree plus binary decision tree process as described below, where the transform types include but are not limited to B-DCT and SA-DCT. In some embodiments, when the residual is considered to have a high error, a B-DCT transform type is selected; and when the residual is considered to have a low error, an SA-DCT transform type is selected. Based on the selected transform type, a transform process (e.g., SA-DCT or B-DCT) is performed to generate quantizable coefficients.
[0044] In step 350, and continuing with reference to Figure 3 , the determined transform type is signaled in the bitstream. The residual of the transform and quantization may be included in the bitstream. In some embodiments, the number of transform coefficients may be signaled in the bitstream.
[0045] Figure 4 4 is a system block diagram showing a non-limiting example of a decoder 400 capable of decoding a bitstream 470 using DCT (including but not limited to SA-DCT and / or B-DCT) for geometric partitioning with adaptive number of regions, wherein the geometric partitioning can improve the complexity and processing performance of video encoding and decoding. The decoder 400 includes an entropy decoder processor 410, an inverse quantization and inverse transform processor 420, a deblocking filter 430, a frame buffer 440, a motion compensation processor 450, and an intra-frame prediction processor 460. In some embodiments, the bitstream 470 includes parameters that signal a geometric partitioning mode and a transform type. In some embodiments, the bitstream 470 includes parameters that signal a number of transform coefficients. The motion compensation processor 450 can reconstruct pixel information using the geometric partitioning described herein.
[0046] In operation, and still refer to Figure 4 , the bitstream 470 may be received by the decoder 400 and input to the entropy decoder processor 410, which entropy decodes the bitstream into quantized coefficients. The quantized coefficients are provided to the inverse quantization and inverse transform processor 420, which may determine the type of coding transformation (e.g., B-DCT or SA-DCT) and perform inverse quantization and inverse transformation according to the determined coding transformation type to generate a residual signal. In some embodiments, the inverse quantization and inverse transform processor 420 may determine the number of transform coefficients and perform inverse transformation according to the determined number of transform coefficients.
[0047] Still refer to Figure 4, the residual signal may be added to the output of the motion compensation processor 450 or the intra prediction processor 460 depending on the processing mode. The output of the motion compensation processor 450 and the intra prediction processor 460 may include a block prediction value based on a previously decoded block. The sum of the prediction value and the residual may be processed by the deblocking filter 430 and stored in the frame buffer 440. For a given block (e.g., CU or PU), when the signal sent by the bitstream 470 indicates that the partitioning mode is a block-level geometric partitioning, the motion compensation processor 450 may construct a prediction value based on the geometric partitioning method described herein.
[0048] Figure 5 5 is a process flow diagram illustrating an exemplary process 500 for decoding a bitstream using SA-DCT for geometric partitioning with region number adaptation, wherein the geometric partitioning can improve the complexity and processing performance of video encoding and decoding. In step 510, a bitstream that may include a current block (e.g., CTU, CU, PU) is received. Receiving may include extracting and / or parsing the current block and associated signaling information from the bitstream. The decoder may extract or determine one or more parameters characterizing the geometric partitioning. For example, these parameters may include indices of the start and end points of a line segment (e.g., P1, P2, P3, P4); extracting or determining may include identifying and retrieving parameters from the bitstream (e.g., parsing the bitstream).
[0049] In step 520, and still referring to Figure 5 , a first region, a second region, and a third region of the current block may be determined according to a geometric partitioning mode. The determination may include determining whether a geometric partitioning mode for the current block is activated (e.g., true). If the geometric partitioning mode is not activated (e.g., false), the decoder may process the current block using an alternative partitioning mode. If the geometric partitioning mode is activated (e.g., true), three or more regions are determined and / or processed.
[0050] In optional step 530, and continuing with reference to Figure 5 , determining a coding transformation type. The coding transformation type may be signaled in the bitstream. For example, the bitstream is parsed to determine a coding transformation type (which may be specified as B-DCT or SA-DCT). The determined coding transformation type may be used to decode the first region, the second region, and / or the third region.
[0051] In 540, and still referring to Figure 5 , decoding the current block. Decoding the current block may include inverse transforming each of the first region, the second region, and / or the third region using the determined transform type. Decoding may include determining associated motion information for each region based on the geometric partitioning mode.
[0052] Although some variations have been described in detail above, other modifications or additions are possible. For example, the geometric partitioning can be signaled in the bitstream based on rate-distortion decisions in the encoder. Encoding can be performed based on conventional predefined partitions (e.g., templates), temporal and spatial predictions of the partitions, and other offset combinations. Each geometric partitioned region can use motion compensated prediction or intra-frame prediction. The boundaries of the prediction regions can be smoothed before adding residuals.
[0053] In some embodiments, a quadtree plus binary decision tree (QTBT) may be implemented. In QTBT, at the coding unit level, the partitioning parameters of the QTBT are dynamically derived to adapt to local features without any transmission overhead. Then, at the coding unit level, the joint classifier decision tree structure can eliminate unnecessary iterations and control the risk of false predictions. In some embodiments, geometric partitioning with adaptive number of regions can be used as an additional partitioning option available at each leaf node of the QTBT.
[0054] In some embodiments, the decoder may include a partition processor that generates a geometric partition for the current block and provides all partition related information to the relevant processes. The partition processor may directly affect motion compensation, as it may be performed segment by segment if the block is geometrically partitioned. In addition, the partition processor may provide shape information to the intra prediction processor and the transform coding processor.
[0055] In some embodiments, additional syntax elements may be signaled at different levels of the bitstream. To activate geometric partitioning with region number adaptation over the entire sequence, an activation flag may be encoded in the Sequence Parameter Set (SPS). In addition, at the Coding Tree Unit (CTU) level, a CTU flag is encoded to indicate whether any Coding Unit (CU) uses geometric partitioning with region number adaptation. A CU flag is encoded to indicate whether the current Coding Unit uses geometric partitioning with region number adaptation. Parameters for line segments on a specified block are encoded. For each region, a flag is encoded that specifies whether the current region is intra-predicted or inter-predicted.
[0056] In some implementations, the size of the minimum region may be specified.
[0057] The subject matter of the invention described herein has many technical advantages. For example, some embodiments of the present subject matter provide block partitioning that can reduce complexity while improving compression efficiency. In some embodiments, the block effect of object boundaries can be eliminated.
[0058] It should be noted that any one or more of the aspects and embodiments described in this specification are easy to implement with digital electronic circuits, integrated circuits, specially designed application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof. It is obvious to a person of ordinary skill in the computer field that it is implemented and / or implemented in one or more machines (e.g., one or more computing devices, one or more server devices, such as document servers) programmed according to the teachings of this specification. These various aspects or features may be implemented in one or more computer programs and / or software executed and / or interpreted on a programmable system, the programmable system including at least one programmable processor, which may be dedicated or general purpose, coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to transmit data and instructions to a storage system, at least one input device, and at least one output device. A skilled programmer can easily prepare corresponding software coding based on the teachings of this disclosure, which is obvious to a person of ordinary skill in the software field. The software and / or software modules used in the above aspects and embodiments may also include corresponding hardware for assisting the machine in executing software and / or software module instructions.
[0059] Such software may be a computer program product using a machine-readable storage medium. A machine-readable storage medium may be any medium capable of storing and / or encoding a sequence of instructions executable by a machine (e.g., a computing device), and causing the machine to perform the methods and / or embodiments described herein. Examples of machine-readable storage media include, but are not limited to, disks, optical disks (e.g., CDs, CD-Rs, DVDs, and DVD-Rs, etc.), magneto-optical disks, read-only storage devices "ROMs", random access storage devices "RAMs", magnetic cards, optical cards, solid-state storage devices, EPROMs, EEPROMs, programmable logic devices (PLDs), and / or any combination thereof. The machine-readable storage medium used herein is intended to include a single medium and a physically separated set of media, such as a combination of an optical disk set or one or more hard disk drives and a computer memory. The machine-readable storage medium used herein does not include transient storage forms in signal transmission.
[0060] Such software may also include information (e.g., data) carried as a data signal on a data carrier such as a carrier wave. For example, machine executable information may be included as a data signal carried on a data carrier, wherein the signal encodes a sequence of instructions or a portion thereof for execution by a machine (e.g., a computing device), as well as any related information (e.g., data structures and data) that enables the machine to perform any of the methods and / or embodiments described herein.
[0061] Examples of computing devices include, but are not limited to, electronic book reading devices, computer workstations, terminal computers, server computers, handheld devices (e.g., tablet computers and smart phones, etc.), network devices, network routers, network switches, bridges, any machine capable of executing a sequence of instructions for instructing the machine to take an action, and any combination thereof. In one example, the computing device may include and / or be included in a public information kiosk.
[0062] Figure 6 A schematic diagram of one embodiment of a computing device in the example form of a computer system 600 is shown, wherein a set of instructions for causing a control system to perform any one or more of the aspects and / or methods of the present disclosure may be executed. It is also contemplated that a plurality of computing devices may be utilized to execute a set of specially configured instructions to cause one or more of the devices to perform any one or more of the aspects and / or methods of the present disclosure. The computer system 600 includes a processor 604 and a memory 608, which communicate with each other and with other components via a bus 612. The bus 612 may include any of a variety of bus structures, including but not limited to a memory bus, a storage controller, a peripheral bus, a local bus, and any combination thereof using any of a variety of bus architectures.
[0063] The memory 608 may include various components (e.g., machine-readable media), including, but not limited to, random access memory components, read-only components, and any combination thereof. In one example, a basic input / output system 616 (BIOS) may be stored in the memory 608, which includes basic routines that help transfer information between components within the computer system 600 (e.g., during startup). The memory 608 may also include instructions (e.g., software) 620 (e.g., stored on one or more machine-readable media) that implement any one or more of the aspects and / or methods of the present disclosure. In another example, the memory 608 may further include any number of program modules, including, but not limited to, an operating system, one or more application programs, other program modules, program data, and any combination thereof.
[0064] The computer system 600 may also include a storage device 624. Examples of storage devices (e.g., storage device 624) include, but are not limited to, a hard disk drive, a combination of a disk drive and an optical medium, a solid-state storage device, and any combination thereof. The storage device 624 may be connected to the bus 612 via a corresponding interface (not shown). Exemplary interfaces include, but are not limited to, SCSI, Advanced Technology Configuration (ATA), Serial ATA, Universal Serial Bus (USB), IEEE 1394 interface (Firewire), and any combination thereof. In one example, the storage device 624 (or one or more of its components) may be removably connected to the computer system 600, for example, via an external port connector (not shown). Specifically, the storage device 624 and the associated machine-readable medium 628 may provide non-volatile and / or volatile storage for machine-readable instructions, data structures, program modules, and / or other data for the computer system 600. In one example, the software 620 may reside in whole or in part in the machine-readable medium 628. In another example, the software 620 may reside in whole or in part in the processor 604.
[0065] The computer system 600 may also include an input device 632. In one example, a user of the computer system 600 may enter commands and / or other information into the computer system 600 via the input device 632. Examples of the input device 632 include, but are not limited to, an alphanumeric input device (e.g., a keyboard), a pointing device, a joystick, a game controller, an audio input device (e.g., a microphone and voice response system, etc.), a cursor control device (e.g., a mouse), a touch pad, an optical scanner, a video capture device (e.g., a still camera and a video camera), a touch screen, and any combination thereof. The input device 632 may be connected to the bus 612 via any of a variety of interfaces (not shown); the interfaces include, but are not limited to, a serial interface, a parallel interface, a game port, a USB interface, a FireWire interface, a direct interface to the bus 612, and any combination thereof. The input device 632 may include a touch screen interface, which may be part of the display 636 or separate from the display 636, as will be discussed further below. The input device 632 may be used as a user selection device to select one or more graphical representations in a graphical interface as described above.
[0066] The user may also input instructions and / or other information into the computer system 600 via a storage device 624 (e.g., a removable disk drive and a flash drive, etc.) and / or a network interface device 640. A network interface device (e.g., network interface device 640) may be used to connect the computer system 600 to one or more of a plurality of networks (e.g., network 644), and to one or more remote devices 648 connected thereto. Examples of network interface devices include, but are not limited to, network interface cards (e.g., mobile network interface cards, LAN cards), modems, and any combination thereof. Examples of networks include, but are not limited to, wide area networks (e.g., the Internet and enterprise networks), local area networks (e.g., networks associated with offices, buildings, campuses, or other relatively small geographic spaces), telephone networks, data networks associated with telephone / voice providers (e.g., data and / or voice networks of mobile communication providers), direct connections between two computing devices, and any combination thereof. The network may employ wired and / or wireless communication modes, such as network 644. In general, any network topology may be used. Information (e.g., data and software 620, etc.) may be transmitted to and / or from the computer system 600 via the network interface device 640.
[0067] The computer system 600 may further include a video display adapter 652 for transmitting displayable images to a display device, such as a display device 636. Examples of display devices include, but are not limited to, liquid crystal displays (LCDs), cathode ray tubes (CRTs), plasma displays, light emitting diode (LED) displays, and any combination thereof. The display adapter 652 and the display device 636 may be used in conjunction with the processor 604 to provide a graphical representation of aspects of the present invention. In addition to the display device, the computer system 600 may include one or more other peripheral output devices, including, but not limited to, audio speakers, printers, and any combination thereof. The peripheral output device may be connected to the bus 612 via a peripheral interface 656. Examples of peripheral interfaces include, but are not limited to, serial ports, USB interfaces, FireWire interfaces, parallel interfaces, and any combination thereof.
[0068] The illustrative embodiments of the present invention have been described in detail above. Various modifications and additions may be made to the present invention without departing from the spirit and scope of the present invention. The features of each of the above-mentioned multiple embodiments may be combined with the features of other described embodiments as appropriate, so as to provide a variety of feature combinations in related new embodiments. In addition, although a plurality of separate embodiments have been described above, the description in the present invention is merely an illustration of the application of the principles of the present invention. In addition, although the specific methods of the present invention are shown and / or described as being performed in a specific order, the order is highly variable within the ordinary art to implement the embodiments of the present disclosure. Therefore, this specification is intended to be illustrative only and is not intended to limit the scope of the present invention.
[0069] In the above description and claims, phrases such as "at least one" or "one or more" may appear, followed by an associated list of elements or features. The term "and / or" may also appear in a list containing two or more elements or features. Unless otherwise implied or explicitly stated to contradict the phrase used in the context, the phrase is intended to mean any element or feature listed alone, or any listed element or feature combined with other listed elements or features. For example, the phrases "at least one of A and B;", "one or more of A and B;" and "A and / or B" are intended to mean "single A, single B, or A and B", respectively. A similar interpretation also applies to lists containing three or more items. For example, the phrases "at least one of A, B and C", "one or more of A, B and C" and "A, B and / or C" are intended to mean "single A, single B, single C, A and B, A and C, B and C, or A, B and C", respectively. In addition, the term "based on" used in the above and claims is intended to mean "based at least in part on", thereby also allowing the inclusion of unlisted features or elements.
[0070] The subject matter described in the present invention can be implemented in systems, devices, methods and / or articles according to the desired configuration. The embodiments described in the foregoing description do not represent all embodiments consistent with the subject matter described in the present invention. On the contrary, they are only some examples consistent with aspects related to the subject matter. Although some changes have been described in detail above, other modifications or additions are also possible. In addition to the changes described above, other features and / or changes can also be provided in particular. For example, the embodiments described above are intended to provide multiple combinations and sub-combinations of disclosed features and / or combinations and / or combinations and sub-combinations of several other features disclosed above. In addition, the logical flow shown in the drawings and / or described in the present invention does not necessarily need to be in the specific order shown or in a sequential order to achieve the desired results. Other embodiments may also be within the scope of the appended claims.
Claims
1. A decoder, the decoder comprising circuitry configured to: receive a bitstream, the bitstream including encoded pictures and signaling information, the encoded pictures including encoded tree units, the signaling information including a sequence parameter set, the sequence parameter set containing a flag indicating activation of geometric partitioning for the bitstream, a first index for enabling the decoder to determine a first endpoint on the boundary of the encoded tree unit for a first straight-line partition in the encoded tree unit, a second index for enabling the decoder to determine a second endpoint of the first straight-line partition on the boundary of the encoded tree unit, and a third index for enabling the decoder to determine a first endpoint of a second straight-line partition in the encoded tree unit, wherein the first endpoint of the second straight-line partition in the encoded tree unit is located on the boundary of the encoded tree unit, and a fourth index for enabling the decoder to determine a second endpoint of the second straight-line partition, the second endpoint of the second straight-line partition being located on the first straight-line partition; and decode the encoded tree unit using the first index, the second index, the third index, and the fourth index to reconstruct the encoded tree unit, the reconstructed encoded tree unit being divided by the first straight-line partition and the second straight-line partition into three non-rectangular regions.
2. The decoder according to claim 1, further comprising: an entropy decoder processor configured to receive the bitstream and decode the bitstream into quantized coefficients; an inverse quantization and inverse transform processor configured to process the quantized coefficients, including performing an inverse discrete cosine transform according to a determined coding transform type; a deblocking filter; a frame buffer; and an intra prediction processor.
3. A method, comprising: receiving, by a decoder, a bitstream, the bitstream including a current encoded picture and signaling information, the current encoded picture including encoded tree units, the signaling information including a sequence parameter set, the sequence parameter set containing a flag indicating activation of geometric partitioning for the bitstream, a first index for enabling the decoder to determine a first endpoint on the boundary of the encoded tree unit for a first straight-line partition in the encoded tree unit, a second index for enabling the decoder to determine a second endpoint of the first straight-line partition on the boundary of the encoded tree unit, and a third index for enabling the decoder to determine a first endpoint of a second straight-line partition in the encoded tree unit, wherein the first endpoint of the second straight-line partition in the encoded tree unit is located on the boundary of the encoded tree unit, and a fourth index for enabling the decoder to determine a second endpoint of the second straight-line partition, the second endpoint of the second straight-line partition being located on the first straight-line partition; and decoding a current encoded tree unit using the first index, the second index, the third index, and the fourth index to reconstruct the encoded tree unit, the reconstructed encoded tree unit being divided by the first straight-line partition and the second straight-line partition into three non-rectangular regions.
Citation Information
Patent Citations
Inter-layer prediction through texture segmentation for video coding
US20130287109A1