High-Speed DST-7

The method enhances the efficiency of DST Type 7 implementation in video codecs by generating specific tuples of transform core elements and performing transformations on blocks, thereby reducing computational resources and improving performance.

JP7698102B2Active Publication Date: 2025-06-24TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024071040
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-12-12
Filing Date
2024-04-25
Publication Date
2025-06-24
Estimated Expiration
2039-04-11

AI Technical Summary

Technical Problem

The implementation of Discrete Sine Transform Type VII (DST Type 7) is inefficient compared to Discrete Cosine Transform Type 2 (DCT-2), limiting its application in practical video codec implementations.

Method used

A method for decoding a video sequence using a DST Type VII transform core, which involves generating tuples of transform core elements, creating an n-point DST-VII transform core, and performing transformations on blocks using this core, thereby reducing the number of arithmetic operations required.

Benefits of technology

This approach enables significant reduction in computational resources and improves efficiency by requiring fewer multiplication operations while producing substantially the same results as matrix multiplication-based implementations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698102000002
    Figure 0007698102000002
  • Figure 0007698102000003
    Figure 0007698102000003
  • Figure 0007698102000004
    Figure 0007698102000004
Patent Text Reader

Abstract

To provide a fast-speed DST-7.SOLUTION: A method and apparatus for decoding a video sequence using a discrete sine transform (DST) type-VII transform core includes generating a set of tuples of transform core elements associated with an n-point DST-VII transform core. A first sum of a first subset of transform core elements of a first tuple is equal to a second sum of a second subset of remaining transform core elements of the first tuple. The n-point DST-VII transform core is generated based on generating the set of tuples of transform core elements. A transform on a block is performed using the n-point DST-VII transform core.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - Reference to Related Applications] This application claims priority under 35 U.S.C. § 119 to U.S. Patent Application No. 62 / 668,065, filed on May 7, 2018, with the United States Patent and Trademark Office, and incorporates by reference the entire disclosure thereof.

[0002] [Technical Field] The present disclosure relates to next - generation video coding technologies beyond HEVC (High Efficiency Video Coding), such as VVC (Versatile Video Coding) for example. More specifically, the present disclosure is directed to a method for accelerating the implementation of Discrete Sine Transform Type VII (DST Type 7).

Background Art

[0003] ITU - T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the H.265 / HEVC (High Efficiency Video Coding) standard (version 1) in 2013, and provided updated versions in 2014 (version 2), 2015 (version 3), and 2016 (version 4). ITU is studying the potential need for standardizing future video coding technologies with compression capabilities significantly exceeding the HEVC standard (including its extensions).

[0004] In October 2017, ITU issued a Joint Call for Proposals on Video Compression with Capability beyond HEVC (CfP). By February 15, 2018, a total of 22 CfP responses for standard dynamic range (SDR), 12 CfP responses for high dynamic range (HDR), and 12 CfP responses for 360 - video categories were respectively submitted.

[0005] In April 2018, all received CfP responses were evaluated at the 122nd MPEG / 10th JVET (Joint Video Exploration Team - Joint Video Expert Team) meeting. Through careful evaluation, JVET officially started the standardization of next-generation video coding beyond HEVC, namely, the so-called VVC (Versatile Video Coding). The current version is VTM (VVC Test Model), namely VTM 1.

[0006] Compared with DCT-2, for which acceleration methods have been widely studied, the implementation of DST-7 is still very inefficient compared to DCT-2. For example, VTM 1 involves matrix multiplication.

[0007] In JVET-J0066, a method for approximating different types of DCT and DST in JEM7 by applying an adjustment stage to the transforms in the DCT-2 family including DCT-2, DCT-3, DST-2, and DST-3 has been proposed. The adjustment stage represents matrix multiplication using a sparse matrix that requires a relatively small number of arithmetic operations.

[0008] In JVET-J001, a method for implementing n-point DST-7 using a (2n + 1)-point discrete Fourier transform (DFT, Discrete Fourier Transform) has been proposed. SUMMARY OF THE INVENTION

[0009] According to one aspect of the present disclosure, a method for decoding a video sequence using a discrete sine transform (DST) type VII transform core includes generating a set of tuples of transform core elements associated with an n-point DST-VII transform core, wherein a first sum of a first subset of the transform core elements of a first tuple is equal to a second sum of a second subset of the remaining transform core elements of the first tuple; generating an n-point DST-VII transform core based on the generated set of tuples of transform core elements; and performing a transform on a block using the n-point DST-VII transform core.

[0010] According to one aspect of the present disclosure, a device for decoding a video sequence using a discrete sine transform (DST) type VII transform core includes at least one memory configured to store program code, and at least one processor configured to read the program code and operate according to instructions in the program code. The program code includes generation code configured to cause the at least one processor to generate a set of tuples of transform core elements associated with an n-point DST-VII transform core, wherein a first sum of a first subset of the transform core elements of a first tuple is equal to a second sum of a second subset of the remaining transform core elements of the first tuple. The generation code is further configured to cause the at least one processor to generate an n-point DST-VII transform core based on the generated set of tuples of transform core elements. The program code further includes execution code configured to cause the at least one processor to perform a transform on a block using the n-point DST-VII transform core.

[0011] According to one aspect of the present disclosure, in a non-transitory computer-readable medium storing instructions, when the instructions are executed by one or more processors of a device, the one or more processors are caused to perform steps of generating a set of tuples of transform core elements related to an n-point DST-VII transform core, wherein a first sum of a first subset of transform core elements of a first tuple is equal to a second sum of a second subset of the remaining transform core elements of the first tuple; generating an n-point DST-VII transform core based on the set of tuples of transform core elements; and performing a transform on a block using the n-point DST-VII transform core.

Brief Description of the Drawings

[0012] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Modes for Carrying Out the Invention

[0013] [Problems to be Solved] The lack of an efficient high-speed implementation of DST-7 limits the application of DST-7 for practical video codec implementations.

[0014] In different implementation scenarios, matrix multiplication-based implementations are preferred because they involve more regular processing. However, in some cases, a speed-up method that significantly reduces the number of arithmetic operations is preferred. Therefore, it is highly desirable to identify a speed-up method that outputs substantially the same results as compared to matrix multiplication-based implementations, such as the DCT-2 design in HEVC, which supports both matrix multiplication and partial butterfly implementations.

[0015] Existing speed-up methods for DST-7, such as JVET-J0066 and JVET-J0017, cannot support all desirable features of the transform design in video codecs, including 16-bit intermediate operations and integer operations, and / or cannot provide the same results between the implementation of the speed-up method and the matrix multiplication-based implementation.

[0016] [Detailed Description] The present disclosure enables substantially similar results as compared to matrix multiplication-based implementations based on the utilization of individual features / patterns in the transform basis of DST-7. Thus, some implementations herein save computational resources of the encoder and / or decoder and improve efficiency.

[0017] As follows, the 16-point DST-7 integer transform core used in the forward transform can be represented as {a, b, c, d, e, f, g, h, i, j, k, l, m, n, o, p} = {8, 17, 25, 33, 41, 48, 55, 62, 67, 73, 77, 81, 84, 87, 88, 89}.

[0018] As described above, the element values of the DST-7 transform core include the characteristics of a + j = l, b + i = m, c + h = n, d + g = o, and e + f = p.

[0019] According to one embodiment, for a 16-point transform, the input vector is x = {x0, x1, x2,..., x15}, and the output transform coefficient vector is y = {y0, y1, y2,..., y15}.

[0020] Based on the above relationships between the transformation core elements, instead of implementing a·x0 + j·x9 + l·x11 which requires three multiplication operations, one embodiment implements a·(x0 + x11) + j·(x9 + x11) which requires two multiplication operations.

[0021] Thus, instead of performing per-vector operations of y0 = a·x0 + b·x1 + c·x2 + d·x3 + e·x4 + f·x5 + g·x6 + h·x7 + i·x8 + j·x9 + k·x10 + l·x11 + m·x12 + n·x13 + o·x14 + p·x15 which requires 16 multiplication operations to calculate y0, one embodiment performs operations of y0 = a·(x0 + x11) + b·(x1 + x12) + c·(x2 + x13) + d·(x3 + x14) + e·(x4 + x15) + f·(x5 + x15) + g·(x6 + x14) + h·(x7 + x13) + i·(x8 + x12) + j·(x9 + x11) + k·x10 which requires 11 multiplication operations and derives substantially the same result.

[0022] Furthermore, when calculating y2, y3, y6, y8, y9, y11, y12, y14, y15, similar implementations can be realized, and the intermediate results of (x0 + x11), (x1 + x12), (x2 + x13), (x3 + x14), (x4 + x15), (x5 + x15), (x6 + x14), (x7 + x13), (x8 + x12), (x9 + x11) and k·10 can be reused respectively.

[0023] According to one embodiment, each third transformation basis vector starting from the second basis vector includes several replication patterns. For example, the second basis vector can be represented as follows.

[0024]

Number

[0025] Furthermore, the calculations when calculating y1, y4, y7, y10, and y13 can be performed in the same way, and the intermediate results of (x0 + x9 - x11), (x1 + x8 - x12), (x2 + x7 - x13), (x3 + x6 - x14), and (x4 + x5 - x15) can be reused.

[0026] Regarding the inverse transformation, the transformation core matrix is the transpose of the transformation core matrix used for the forward transformation, and the above two features are also applicable to the inverse transformation. Furthermore, it should be noted that the tenth basis vector is {k, 0, -k, k, 0, -k, k, 0, -k, k, 0, -k, k, 0, -k, k}, which contains only a single unique absolute value (i.e., k). Thus, instead of using the per-vector multiplication of y10 = k·x0 - k·x2 + k·x3 - k·x5 + k·x6 - k·x8 + k·x9 - k·x11 + k·x12 - k·x14 + k·x15 which requires 11 multiplication operations to calculate y10, one embodiment may perform the operation of y10 = k·(x0 - x2 + x3 - x5 + x6 - x8 + x9 - x11 + x12 - x14 + x15) which requires a single multiplication operation while deriving substantially the same result.

[0027] For the 64-point forward and reverse DST-7, the above two features are also applicable, and thus the same acceleration method as above is also applicable.

[0028] For the 32-point forward and reverse DST-7, the second feature is applicable, that is, there are replicated or inverted segments in the part of the basis vectors. Therefore, the same acceleration method as described above is also applicable based on the second feature. Also, the first feature can be utilized in the 32-point forward and reverse DST-7 transform cores with different formulations as described below.

[0029] The elements of the 32-point transform core include 32 different numbers of {a, b, c, d, e, f, g, h, i, j, k, l, m, n, o, p, q, r, s, t, u, v, w, x, y, z, A, B, C, D, E, F} (without considering sign changes).

[0030] An example of the fixed-point assignment of the elements is {a, b, c, d, e, f, g, h, i, j, k, l, m, n, o, p, q, r, s, t, u, v, w, x, y, z, A, B, C, D, E, F} = {4, 9, 13, 17, 21, 26, 30, 34, 38, 42, 46, 49, 53, 56, 60, 63, 66, 69, 71, 74, 76, 78, 81, 82, 84, 85, 87, 88, 89, 89, 90, 90}.

[0031] It should be noted that the element values of the 32-point floating-point DST-7 transform core have the characteristics of #0: a + l + A = n + y; #1: b + k + B = o + x; #2: c + j + C = p + w; #3: d + i + D = q + v; #4: e + h + E = r + u; and #5: f + g + F = s + t.

[0032] The elements included in each of the above six equations form a quintuple. For example, {a, l, A, n, y} is quintuple #0, and {b, k, B, o, x} is another quintuple #1.

[0033] According to one embodiment, the input vector for the 32-point transform is x = {x0, x1, x2,..., x31}, and the output transform coefficient vector is y = {y0, y1, y2,..., y31}.

[0034] Based on the above relationship between the transform core elements, instead of implementing a·x0 + l·x11 + n·x13 + y·x24 + A·x26 which requires five multiplication operations, one embodiment performs the operation of a·x0 + l·x11 + n·x13 + y·x24 + (n + y - a - l)·x26 while deriving substantially the same result. Also, this can be implemented as a·(x0 - x26) + l·(x11 - x26) + n·(x13 + x26) + y·(x24 + x26) which requires four multiplication operations.

[0035] Similarly, the above reduced-form multiplication operations are also applicable to the other five quintuples. For example, the intermediate results for the quintuple #0, (x0 - x26), (x11 - x26), (x13 + x26), (x24 + x26) are pre-calculated and can be reused to calculate each of the transform coefficients.

[0036] However, it should be noted that the assignment of integer values to the elements {a, b, c,..., F} may not exactly follow the above formula due to rounding errors. For example, in the case of {a, b, c, d, e, f, g, h, i, j, k, l, m, n, o, p, q, r, s, t, u, v, w, x, y, z, A, B, C, D, E, F} = {4, 9, 13, 17, 21, 26, 30, 34, 38, 42, 46, 49, 53, 56, 60, 63, 66, 69, 71, 74, 76, 78, 81, 82, 84, 85, 87, 88, 89, 89, 90, 90} assigned by scaling 64·√32 and rounding to the nearest integer.

[0037] For quintuple #1, b + k + B = 9 + 46 + 88 = 143, but o + x = 60 + 82 = 142.

[0038] Thus, to achieve substantially the same results between matrix multiplication and the acceleration method, one embodiment adjusts each of the five-term elements to execute substantially the same as the equations defined herein. For example, the five-term #1 is adjusted such that {b, k, o, x, B} = {9, 46, 60, 82, 87}. Alternatively, the five-term #1 is adjusted such that {b, k, o, x, B} = {9, 46, 60, 83, 88}. Alternatively, the five-term #1 is adjusted such that {b, k, o, x, B} = {9, 46, 61, 82, 88}. Alternatively, the five-term #1 is adjusted such that {b, k, o, x, B} = {9, 45, 60, 82, 88}. Alternatively, the five-term #1 is adjusted such that {b, k, o, x, B} = {8, 46, 60, 82, 88}.

[0039] Compared to a matrix multiplication-based implementation that requires 256 multiplication operations and 256 addition / subtraction operations, or the JVET-J0066 implementation that requires 152 multiplication operations and 170 addition / subtraction operations, some implementations herein enable 126 multiplication operations and 170 addition / subtraction operations while providing substantially the same results. Thus, some implementations herein enable improved efficiency and conserve the computational resources of the encoder and / or decoder.

[0040] FIG. 1 is a flowchart of an exemplary process 100 for decoding a video sequence using a discrete sine transform (DST) type VII transform core. In some implementations, one or more of the processing blocks in FIG. 1 may be executed by a decoder. In some implementations, one or more of the processing blocks in FIG. 1 may be executed by another device or group of devices separate from the decoder, such as an encoder, or by another device or group of devices that includes the decoder.

[0041] As shown in FIG. 1, process 100 may include generating a set of tuples of transform core elements associated with an n-point DST-VII transform core, and a first sum of a first subset of transform core elements of a first tuple is equal to a second sum of a second subset of the remaining transform core elements of the first tuple (block 110).

[0042] As further shown in FIG. 1, process 100 may include generating an n-point DST-VII transform core based on generating a set of tuples of transform core elements (block 120).

[0043] As further shown in FIG. 1, process 100 may include performing a transform on a block using the n-point DST-VII transform core.

[0044] According to one embodiment, the distinct absolute element values present in one transform core can be divided into a plurality of tuples along with some predetermined constants, and for each tuple, the sum of some of the absolute element values is the same as the sum of the remaining absolute element values within the same tuple.

[0045] For example, in an embodiment, the tuple is ternary including three elements, and the sum of the absolute values of two elements is the same as the absolute value of the remaining one element. Alternatively, the tuple is quaternary including four elements, and the sum of the absolute values of two elements is the same as the sum of the absolute values of the remaining two elements. Alternatively, the tuple is quaternary including four elements, and the sum of the absolute values of three elements is the same as the sum of the absolute value of the remaining one element. Alternatively, the tuple is quinary including five elements, and the sum of the absolute values of three elements is the same as the sum of the absolute values of the remaining two elements.

[0046] According to one embodiment, in addition to the existing distinct absolute element values present in one transform core, predetermined constants that are powers of two (e.g., 1, 2, 4, 8, 16, 32, 64, etc.) can also be considered as elements within the tuple.

[0047] According to one embodiment, the distinct absolute element values present in the 16-point and 64-point DST-7 transform cores are partitioned into a plurality of ternaries each containing three elements. Further, also, or alternatively, the distinct absolute element values present in the 32-point DST-7 transform core are partitioned into a plurality of quinary each containing five elements.

[0048] According to one embodiment, for the integer transform core, the integer elements of the transform core may be further adjusted to exactly satisfy the above characteristics, that is, the sum of the absolute values of some of the elements within one tuple is the same as the sum of the absolute values of the remaining elements within the same tuple while maintaining good orthogonality of the transform core. For example, for the 32-point integer DST-7 core, the second tuple {b, k, o, x, B} is adjusted to {9, 46, 60, 82, 87}. Or alternatively, for the 32-point integer DST-7 core, the second tuple {b, k, o, x, B} is adjusted to {9, 46, 60, 83, 88}. Or alternatively, for the 32-point integer DST-7 core, the second tuple {b, k, o, x, B} is adjusted to {9, 46, 61, 82, 88}. Or alternatively, for the 32-point integer DST-7 core, the second tuple {b, k, o, x, B} is adjusted to {9, 45, 60, 82, 88}. Or alternatively, for the 32-point integer DST-7 core, the second tuple {b, k, o, x, B} is adjusted to {8, 46, 60, 82, 88}. Or alternatively, for the 32-point integer DST-7 core, the sixth tuple {f, g, s, t, B} is adjusted to {26, 30, 71, 74, 89}. Or alternatively, for the 32-point integer DST-7 core, the sixth tuple {b, k, o, x, B} is adjusted to {26, 30, 71, 75, 90}. Or alternatively, for the 32-point integer DST-7 core, the sixth tuple {b, k, o, x, B} is adjusted to {26, 30, 72, 74, 90}. Or alternatively, for the 32-point integer DST-7 core, the sixth tuple {b, k, o, x, B} is adjusted to {26, 29, 71, 74, 90}. Or alternatively, for the 32-point integer DST-7 core, the sixth tuple {b, k, o, x, B} is adjusted to {25, 30, 71, 74, 90}.

[0049] According to one embodiment, the elements of the conversion core may be further adjusted only by +1 or -1 with respect to the element values derived by scaling with a predetermined constant and rounding to the nearest integer.

[0050] According to one embodiment, the elements of the conversion core may be further adjusted only by +1, -1, +2, and -2 with respect to the element values derived by scaling with a predetermined constant and rounding to the nearest integer.

[0051] According to one embodiment, the orthogonality of the adjusted conversion core is T measured by the sum of the absolute values of A·A

[0052] Figure 1 shows exemplary blocks of process 100. However, in some implementations, process 100 may include more blocks, fewer blocks, different blocks, or differently arranged blocks than those shown in Figure 1. Additionally, or alternatively, two or more of the blocks of process 100 may be executed in parallel.

[0053] Figure 2 shows a simplified block diagram of a communication system (200) according to an embodiment of the present disclosure. The communication system (200) may include at least two terminals (210 - 220) interconnected via a network (250). For one-way data transmission, the first terminal (210) may encode local-position video data for transmission to another terminal (220) via the network (250). The second terminal (220) may receive encoded video data of another terminal from the network (250), decode the encoded data, and display the restored video data. One-way data transmission may be common in media-providing applications and the like.

[0054] FIG. 2 shows a pair (230, 240) of second terminals provided to support two-way transmission of encoded video that may occur during a video conference, for example. For two-way transmission of data, each terminal (230, 240) may encode video data captured at a local location for transmission to other terminals via a network (250). Also, each terminal (230, 240) may receive encoded video data transmitted by other terminals, may decode the encoded data, and may display the restored video data on a local display device.

[0055] In FIG. 2, the terminals (210-240) may be shown as servers, personal computers, and smartphones, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure have applications in laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (250) represents any number of networks that transmit encoded video data between terminals (210-240), including, for example, wired and / or wireless communication networks. The communication network (250) may exchange data in a circuit-switched channel and / or a packet-switched channel. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this description, the architecture and topology of the network (250) may not be important to the operation of the present disclosure, unless otherwise described below.

[0056] FIG. 3 shows the arrangement of a video encoder and a video decoder in a streaming environment as an example of an application of the subject matter of the disclosure. The subject matter of the disclosure is similarly applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storage of compressed video on digital media (including CDs, DVDs, memory sticks, etc.).

[0057] The streaming system may include a capture subsystem (313), which may include, for example, a video source (301) (e.g., a digital camera) that generates an uncompressed video sample stream (302). This sample stream (302) is shown in thick lines to emphasize the high data volume when compared to the encoded video bitstream and may be processed by an encoder (303) coupled to the camera (301). The encoder (303) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video bitstream (304) is shown in thin lines to emphasize the low data volume when compared to the sample stream and may be stored in a streaming server (305) for future use. One or more streaming clients (306, 308) may access the streaming server (305) to obtain copies (307, 309) of the encoded video bitstream (304). The client (306) may include a video decoder (310) that decodes an input copy of the encoded video bitstream (307) and generates an output video sample stream (311) that can be rendered on a display (312) or other rendering device (not shown). In some streaming systems, the video bitstreams (304, 307, 309) may be encoded according to a particular video encoding / compression standard. Examples of these standards include ITU-T Recommendation H.265. A video encoding standard under development is informally known as VVC (Versatile Video Coding). The disclosed subject matter may be used in the context of VVC.

[0058] FIG. 4 may be a functional block diagram of a video decoder (310) according to an embodiment of the present invention.

[0059] The receiver (410) may receive one or more encoded video sequences to be decoded by the decoder (310), and in the same or other embodiments, may receive one encoded video sequence at a time, and the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequence may be received from a channel (412), and the channel (412) may be a hardware / software link to a storage device storing the encoded video data. The receiver (410) may receive the encoded video data together with other data (e.g., encoded audio data and / or auxiliary data streams), and these data may be transferred to their respective usage entities (not shown). The receiver (410) may separate the encoded video sequence from other data. To prevent network jitter, a buffer memory (415) may be coupled between the receiver (410) and the entropy decoder / parser (420) (hereinafter referred to as "parser"). If the receiver (410) is receiving data from a storage / transfer device with sufficient bandwidth and controllability, or from an isochronous network, the buffer (415) may not be necessary or may be made small. For use in a best-effort packet network such as the Internet, a buffer (415) may be required, may be relatively large, and advantageously may be of an adaptive size.

[0060] The video decoder (310) may include a parser (420) for reconstructing symbols (421) from an entropy-coded video sequence. The categories of these symbols include information used to manage the operation of the decoder (310) and potential information for controlling a rendering device such as a display (312). The rendering device, as shown in FIG. 4, is not an essential part of the decoder but may be coupled to the decoder. The control information for the rendering device may be in the form of supplementary enhancement information (SEI message) or a video usability information (VUI) parameter set fragment (not shown). The parser (420) may parse / entropy-decode the received encoded video sequence. The encoding of the encoded video sequence may follow a video encoding technique or standard and may follow principles well-known to those skilled in the art, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (420) may extract a set of subgroup parameters for at least one of the subgroups of pixels within the video decoder from the encoded video sequence based on at least one parameter corresponding to the group. The subgroups may include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The entropy decoder / parser may also extract information such as transform coefficients, quantizer parameter (QP) values, motion vectors, etc. from the encoded video sequence.

[0061] The parser (420) may perform an entropy decoding / analysis operation on the video sequence received from the buffer (415) to generate symbols (421). The parser (420) may receive the encoded data and selectively decode specific symbols (421). Further, the parser (420) may determine whether a specific symbol (421) should be provided to the motion compensation prediction unit (453), the scaler / inverse transform unit (451), the intra prediction unit (452), or the loop filter (456).

[0062] The reconstruction of the symbol (421) may involve multiple different units depending on the type of the encoded video picture or a portion thereof (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors. How each unit is involved may be controlled by subgroup control information parsed from the encoded video sequence by the parser (420). Such a flow of subgroup control information between the parser (420) and the following multiple units is not shown for clarity.

[0063] In addition to the above functional blocks, the decoder (310) may conceptually be subdivided into multiple functional units as described below. In a practical implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for the purpose of explaining the subject matter of the disclosure, it is appropriate to conceptually subdivide into the following functional units.

[0064] The first unit is the scaler / inverse transform unit (451). The scaler / inverse transform unit (451) receives the quantized transform coefficients as symbols (621) from the parser (420) along with control information (including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc.). The scaler / inverse transform unit (451) may output a block including sample values that can be input to the aggregator (455).

[0065] In some cases, the output samples of the scaler / inverse transform (451) may be related to intra-coded blocks, i.e., blocks that do not use prediction information from previously reconstructed pictures but may use prediction information from the previously reconstructed part of the current picture. Such prediction information may be provided by the intra-picture prediction unit (452). In some cases, the intra-picture prediction unit (452) uses the surrounding already reconstructed information taken from the current picture (partially reconstructed picture) (456) to generate blocks of the same size and shape as the block being reconstructed. In some cases, the aggregator (455) adds, for each sample, the prediction information generated by the intra-prediction unit (452) to the output sample information provided by the scaler / inverse transform unit (451).

[0066] In other cases, the output samples of the scaler / inverse transform unit (451) may be related to inter-coded and potentially motion-compensated blocks. In such cases, the motion-compensation prediction unit (453) may access the reference picture memory (457) to retrieve the samples used for prediction. According to the symbol (421) related to the block, after motion-compensating the retrieved samples, these samples may be added by the aggregator (455) to the output of the scaler / inverse transform unit (451) (in this case, called the residual samples or residual signal) to generate the output sample information. The addresses in the reference picture memory (457) from which the motion-compensation prediction unit retrieves the prediction samples, which are available to the motion-compensation prediction unit, may be controlled by the motion vector, for example, in the form of a symbol (421) that can have X, Y, and reference picture components. Also, motion compensation may include interpolation of the sample values retrieved from the reference picture memory when an exact motion vector of sub-samples is used, a motion vector prediction mechanism, etc.

[0067] The output samples of the aggregator (455) may undergo various loop filtering techniques within the loop filter unit (456). The video compression technique may include in-loop filter techniques, and the in-loop filter techniques are controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream), and are made available to the loop filter unit (456) as symbols (421) from the parser (420), but respond to the meta information obtained during the decoding of the (in decoding order) previous part of the encoded picture or encoded video sequence, and may also respond to the previously reconstructed and loop filtered sample values.

[0068] The output of the loop filter unit (456) may be a sample stream, and the sample stream is output to the rendering device (312) and may also be stored in the reference picture memory (457) for use in future inter-picture prediction.

[0069] When a particular encoded picture is fully reconstructed, it may be used as a reference picture for future prediction. When the encoded picture is fully reconstructed and the encoded picture is identified as a reference picture (e.g., by the parser (420)), the current reference picture (656) may become part of the reference picture buffer (457), and a new current picture memory may be reallocated before starting the reconstruction of subsequent encoded pictures.

[0070] The video decoder (310) may perform a decoding operation according to a predetermined video compression technique that may be documented in a standard such as ITU-T Rec. H.265. The encoded video sequence conforms to the syntax of the video compression technique or standard, in the sense that it is specified in the video compression technique document or standard, particularly in the document of its profile. Also, for compliance, it is necessary that the complexity of the encoded video sequence is within the range defined by the level of the video compression technique or standard. In some cases, the level restricts the maximum picture size, the maximum frame rate, the maximum reconstruction sample rate (e.g., measured in megasamples per second), the maximum reference picture size, etc. In some cases, the restrictions set by the level may be further restricted through the Hypothetical Reference Decoder (HRD) specification and the metadata about the HRD buffer management transmitted in the encoded video sequence.

[0071] In one embodiment, the receiver (410) may receive additional (redundant) data together with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (310) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, a temporal, spatial, or signal noise ratio (SNR) enhancement layer, a redundant slice, a redundant picture, a forward error correction code, etc.

[0072] FIG. 5 may be a functional block diagram of a video encoder (303) according to an embodiment of the present invention.

[0073] The video encoder (303) may receive video samples from a video source (301) (which is not part of the encoder), and the video source (301) may capture a video image to be encoded by the encoder (303).

[0074] The video source (301) may provide a source video sequence to be encoded by the encoder (303) in the form of a digital video sample stream, and the digital video sample stream may be of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, etc.), any suitable color space (e.g., BT.601 Y CrCB, RGB, etc.) and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media providing system, the video source (301) may be a storage device storing pre-prepared video. In a video conferencing system, the video source (303) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that convey motion when viewed in sequence. Each picture itself may be composed of a spatial array of pixels, and each pixel may include one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0075] According to one embodiment, the encoder (303) may encode and compress pictures of the source video sequence into an encoded video sequence (543) in real time or under any other time constraints required by the application. Achieving an appropriate encoding speed is one function of the controller (550). The controller controls other functional units and is functionally coupled to other functional units as will be described below. The coupling is not shown for clarity. Parameters set by the controller may include rate control related parameters (such as picture skip, quantization, lambda value of rate distortion optimization techniques, etc.), picture size, layout of group of pictures (GOP), maximum motion vector search range, etc. Those skilled in the art can easily recognize other functions of the controller (550) related to the video encoder (303) optimized for a specific system design.

[0076] Some video encoders operate in what those skilled in the art recognize as an "encoding loop." As a very simplified explanation, the encoding loop may include an encoder (530), herein referred to as the "source coder" (which is responsible for generating symbols based on, for example, input pictures and reference pictures to be encoded), and a (local) decoder (533) embedded in the encoder (303). The decoder (533) reconstructs the symbols to generate sample data in the same way as a (remote) decoder does (such that any compression between the symbols and the encoded video bitstream is reversible in the video compression techniques considered in the subject matter of the disclosure). This reconstructed sample stream is input into the reference picture memory (534). Since the decoding of the symbol stream results in an exact bit-by-bit result independent of the location of the decoder (local or remote), the content of the reference picture buffer is also bit-by-bit exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the same sample values as reference picture samples as what the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchronization (including drift that may occur as a result of, for example, being unable to maintain synchronization due to channel errors) is well known to those skilled in the art.

[0077] The operation of the "local" decoder (533) may be the same as that of the "remote" decoder (310), which has already been described in detail above in relation to FIG. 4. However, referring briefly to FIG. 5, since the symbols are available and the encoding / decoding of the symbols into the encoded video sequence by the entropy encoder (545) and the parser (420) can be reversible, the entropy decoding part of the decoder (310) including the channel (412), the receiver (410), the buffer (415), and the parser (420) may not be fully implemented in the local decoder (533).

[0078] One consideration that can be made at this point is that any decoder technology other than the parsing / entropy decoding that exists within the decoder must necessarily exist in a substantially identical functional form within the corresponding encoder. Since the description of the encoder technology is the reverse of the decoder technology described comprehensively, it can be omitted. More detailed explanations are necessary only in certain areas and are provided below.

[0079] As part of its operation, the source coder (530) may perform motion-compensated predictive coding, which predictively encodes an input p-frame by referring to one or more previously encoded frames from a video sequence designated as a "reference frame". In this way, the coding engine (532) encodes the difference between a pixel block of the input frame and a pixel block of a reference frame that can be selected as a predictive reference for the input frame.

[0080] The local video decoder (533) may decode the encoded video data of a frame that can be designated as a reference frame based on the symbols generated by the source coder (530). The operation of the coding engine (532) may advantageously be an irreversible process. If the encoded video data can be decoded by a video decoder (not shown in FIG. 5), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (533) may replicate the decoding process that can be performed by the video decoder for the reference frame and store the reconstructed reference frame in the reference picture cache (534). In this way, the encoder (303) may locally store a copy of the reconstructed reference picture having common content as the reconstructed reference picture obtained by the remote video decoder (without transmission errors).

[0081] Predictor (535) may perform prediction search for the encoding engine (532). That is, for a new frame to be encoded, predictor (535) may search the reference picture memory (534) for sample data (as candidate reference pixel blocks) or specific metadata (reference picture motion vectors, block shapes, etc.). These may function as appropriate prediction references for the new picture. Predictor (535) may operate on a sample block-by-pixel block basis to detect appropriate prediction references. In some cases, the input picture determined by the search result obtained by predictor (535) may have prediction references drawn from a plurality of reference pictures stored in the reference picture memory (534).

[0082] Controller (550) may manage the encoding operation of video coder (530), including, for example, setting parameters and subgroup parameters used for encoding video data.

[0083] The outputs of all the above functional units may undergo entropy encoding in entropy coder (545). The entropy coder converts symbols generated by various functional units into an encoded video sequence by reversibly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.

[0084] Transmitter (540) may buffer the encoded video sequence generated by entropy coder (545) and prepare it for transmission via communication channel (560), which may be a hardware / software link to a storage device storing the encoded video data. Transmitter (540) may merge the encoded video data from video coder (303) with other data to be transmitted (e.g., encoded audio data and / or an auxiliary data stream (not shown)).

[0085] The controller (550) may manage the operation of the encoder (303). During encoding, the controller (550) may assign a specific encoded picture type to each encoded picture. The encoded picture type may affect the encoding technique applicable to each picture. For example, a picture may often be assigned as one of the following frame types.

[0086] An intra picture (I picture) may be encoded and decoded without using other pictures in the sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh pictures. Those skilled in the art will recognize these variations of I pictures and their respective uses and characteristics.

[0087] A predicted picture (P picture) may be encoded and decoded using intra prediction or inter prediction using at most one motion vector and a reference index to predict the sample values of each block.

[0088] A bi - directionally predicted picture (B picture) may be encoded and decoded using intra prediction or inter prediction using at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures may use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0089] Generally, a source picture may be spatially subdivided into a plurality of sample blocks (e.g., blocks of samples of 4×4, 8×8, 4×8, or 16×16 respectively) and encoded block by block. The blocks may be encoded predictively with reference to other (already encoded) blocks as determined by an encoding assignment applied to each picture of the block. For example, blocks of an I picture may be encoded non-predictively, or may be encoded predictively with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be encoded predictively via spatial prediction or temporal prediction with reference to a previously encoded reference picture. Blocks of a B picture may be encoded predictively via spatial prediction or temporal prediction with reference to one or two previously encoded reference pictures.

[0090] The video coder (303) may perform an encoding operation according to a predetermined video encoding technique or standard such as ITU-T Rec. H.265. In that operation, the video encoder (303) may perform various compression operations including a predictive encoding operation that utilizes temporal and spatial redundancies in the input video sequence. Thus, the encoded video data may conform to a syntax specified by the video encoding technique or standard being used.

[0091] In one embodiment, the transmitter (540) may transmit additional data along with the encoded video. The video coder (530) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, supplementary enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, etc.

[0092] Furthermore, the proposed method may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium to execute one or more of the proposed methods.

[0093] The above technology may be implemented as computer software using computer-readable instructions and may be physically stored on one or more computer-readable media. For example, FIG. 6 shows a computer system 1200 suitable for implementing a particular embodiment of the disclosed subject matter.

[0094] The computer software may be encoded using any suitable machine code or computer language, and the machine code or computer language may be subject to assembly, compilation, linking, or similar mechanisms to generate code containing instructions, and the instructions may be executed directly by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or through an interpreter, microcode execution, etc.

[0095] The instructions may be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0096] The components shown in FIG. 6 for the computer system 1200 are exemplary in nature and are not intended to suggest any limitation as to the scope or functionality of the computer software implementing the embodiments of the present disclosure. Also, the configuration of the components should not be construed as having any dependency or requirement regarding any one or combination of the components shown in the exemplary embodiment of the computer system 1200.

[0097] The computer system 1200 may include a specific human interface input device. Such a human interface input device may respond to input by one or more human users, for example, through tactile input (keystrokes, swipes, movements of a data glove, etc.), audio input (voice, clapping, etc.), visual input (gestures, etc.), and olfactory input (not shown). Also, the human interface device may be used to capture specific media that is not necessarily directly related to conscious human input, such as audio (conversation, music, ambient sound, etc.), images (scanned images, photographic images obtained from a still camera, etc.), and video (2D video, 3D video including stereoscopic pictures, etc.).

[0098] The input human interface device may include one or more of a keyboard 601, a mouse 602, a trackpad 603, a touch screen 610, a data glove 1204, a joystick 605, a microphone 606, a scanner 607, and a camera 608.

[0099] In addition, computer system 1200 may include a specific human interface output device. Such a human interface output device may stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by touch screen 610, data glove 1204, or joystick 605, provided that there may be a tactile feedback device that does not function as an input device), audio output devices (such as speaker 609, headphones (not shown), etc.), visual output devices (each of which may or may not have a touch screen input function, each of which may or may not have a tactile feedback function, and some of which may output three-dimensional or higher-dimensional output through means such as two-dimensional visual output or stereoscopic output, including cathode ray tube (CRT) screens, liquid crystal display (LCD) screens, plasma screens, organic light emitting diode (OLED) screens of screen 610, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0100] In addition, computer system 1200 may include a human-accessible storage device such as an optical medium including a CD / DVD ROM / RW 620 having a CD / DVD or similar medium 621, and related media, a thumb drive 622, a removable hard drive or solid state drive 623, legacy magnetic media such as tapes and floppy disks (not shown), devices based on special ROM / ASIC / PLD such as security dongles (not shown), etc.

[0101] Also, those skilled in the art should understand that the term "computer-readable medium" used in connection with the subject matter disclosed herein does not include a transmission medium, a carrier wave, or other non-temporary signals.

[0102] In addition, the computer system 1200 may include an interface to one or more communication networks. The network may be, for example, wireless, wired, or optical. The network may be local, wide area, metropolitan, vehicle and industrial, real-time, delay-tolerant, etc. Examples of networks include Ethernet, wireless LAN, cellular networks (including GSM (global systems for mobile communications), 3rd generation (3G), 4th generation (4G), 5th generation (5G), Long Term Evolution (LTE), etc.), TV wired or wireless wide area digital networks (including cable TV, satellite TV, and terrestrial broadcast TV), vehicle and industrial (including CANBus), etc. A specific network generally requires an external network interface adapter (e.g., a universal serial bus (USB) port of the computer system 1200, etc.) attached to a specific general-purpose data port or peripheral bus 649, and other network interface adapters are generally integrated into the core of the computer system 1200 by being attached to the system bus described below (e.g., an Ethernet interface to a PC computer system or a cellular network to a smartphone computer system). Using any of these networks, the computer system 1200 can communicate with other entities. Such communication may be only one-way reception (e.g., broadcast TV), only one-way transmission (e.g., CANbus to a specific CANbus device), or two-way to other computer systems using, for example, local or wide area digital networks. Specific protocols and protocol stacks may be used in each of the networks and network interfaces as described above.

[0103] The above human interface device, human-accessible storage device, and network interface may be attached to the core 640 of the computer system 1200.

[0104] The core 640 may include one or more central processing units (CPUs) 641, a graphics processing unit (GPU) 642, a special programmable processing unit in the form of a field programmable gate array (FPGA) 643, a hardware accelerator 644 for specific tasks, etc. These devices may be connected through a system bus 1248 together with a read-only memory (ROM) 645, a random access memory (RAM) 646, and an internal mass storage device (an internal hard drive inaccessible to the user, a solid state drive (SSD), etc.) 647. In some computer systems, the system bus 1248 may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the system bus 1248 of the core, or may be attached through a peripheral bus 649. The architecture of the peripheral bus includes a peripheral component interconnect (PCI), USB, etc.

[0105] The CPU 641, GPU 642, FPGA 643, and accelerator 644 may execute specific instructions, and the specific instructions may constitute the above computer code by combination. The computer code may be stored in the ROM 645 or the RAM 646. Also, temporary data may be stored in the RAM 646, while persistent data may be stored in the internal mass storage device 647, for example. By using a cache memory that may be closely related to one or more CPUs 641, GPUs 642, mass storage devices 647, ROM 645, RAM 646, etc., fast storage and retrieval to any of the memory devices may be enabled.

[0106] The computer-readable medium may have computer code for executing operations implemented in various computers. The medium and the computer code may be those specially designed and constructed for the purposes of the present disclosure, or may be those well-known and available to those skilled in the art of computer software.

[0107] By way of example and not limitation, an architecture 1200, specifically, a computer system having a core 640, can provide functions as a result of software embodied on one or more tangible computer-readable media being executed by a processor (including a CPU, GPU, FPGA, accelerator, etc.). Such a computer-readable media may be a media associated with a mass storage device accessible by a user as described above, as well as a specific storage device of the core 640 of a non-transitory nature such as a mass storage device 647 inside the core or a ROM 645. The software implementing various embodiments of the present disclosure may be stored in such a device and executed by the core 640. The computer-readable media may include one or more memory devices or chips according to specific needs. The software may cause the core 640, specifically, a processor (including a CPU, GPU, FPGA, etc.) therein, to define a data structure stored in the RAM 646 and modify such a data structure according to the processing defined by the software, and may cause the core 640 to execute a specific processing or a specific part of a specific processing described herein. Further, alternatively, or in addition, the computer system may provide functions as a result of logic wired in a circuit (e.g., an accelerator 644) or logic embodied in another way, and the circuit may operate instead of or together with software to execute a specific processing or a specific part of a specific processing described herein. References to software include logic and, if necessary, vice versa. References to a computer-readable media may, if necessary, include a circuit (such as an integrated circuit (IC), etc.) for storing software for execution, a circuit for embodying logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.

[0108] Although this disclosure describes some exemplary embodiments, there are changes, substitutions, and various equivalent alternatives that fall within the scope of this disclosure. Accordingly, it will be recognized by those skilled in the art that, although not explicitly illustrated or described herein, many systems and methods can be devised that embody the principles of this disclosure and are thus within the spirit and scope of this disclosure.

Claims

1. 1. A method for transmitting a video sequence encoded by an encoder, comprising: obtaining the video sequence encoded using a Discrete Sine Transform (DST) Type VII transform core; transmitting the video sequence; Including, a set of tuples of transform core elements associated with the n-point DST-VII transform core is generated such that a sum of a subset of the transform core elements of the first tuple is equal to the remaining transform core elements of the first tuple; a transform is performed on the video sequence using the set of tuples of the generated transform core elements associated with the n-point DST-VII transform core; the set of tuples are represented as {a,j,l}, {b,i,m}, {c,h,n}, {d,g,o} and {e,f,p}, the n-point DST-VII transform core is a 16-point DST-VII transform core that includes the transform core elements and element k included in the set of tuples, and a transform vector of the 16-point DST-VII transform core is represented as {a,b,c,d,e,f,g,h,i,j,k,l,m,n,o,p}.

2. 2. The method of claim 1, wherein a+j=l, b+i=m, c+h=n, d+g=o and e+f=p.

3. 1. A device for transmitting an encoded video sequence, comprising: at least one memory configured to store program code; at least one processor configured to read the program code and to operate according to the instructions in the program code; 3. A device comprising: a processor for executing a program code for executing a program for at least one processor;

4. A computer program which, when executed by one or more processors of a device, causes said one or more processors to carry out the method of claim 1 or 2.

5. A video encoding method in an encoder, comprising: generating a bitstream; storing the generated bitstream; Including, the bitstream is encoded using a Discrete Sine Transform (DST) Type VII transform core; a set of tuples of transform core elements associated with an n-point DST-VII transform core such that a sum of a subset of the transform core elements of a first tuple is equal to the remaining transform core elements of said first tuple; The set of tuples is represented as {a,j,l}, {b,i,m}, {c,h,n}, {d,g,o} and {e,f,p}, the n-point DST-VII transform core is a 16-point DST-VII transform core that includes the transform core elements included in the set of tuples and element k, and transform vectors of the 16-point DST-VII transform core are represented as {a,b,c,d,e,f,g,h,i,j,k,l,m,n,o,p}.

Citation Information

Patent Citations

  • Enhanced multiple transforms for prediction residual

    WO2016123091A1