Method and apparatus for signaling decoded data using high-level syntax elements
By using explicit or inferred syntax elements to indicate decoded data in the video bitstream, the problem of excessive bit requirements for high-level syntax elements is solved, encoding and decoding efficiency is improved, and the video encoding and decoding process is optimized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-17
- Publication Date
- 2026-04-07
AI Technical Summary
Existing video encoding and decoding schemes require a large number of bits to activate signaling encoding and decoding tools at high compression efficiency, which leads to a decrease in the encoding and decoding efficiency of video bitstreams.
By using syntax elements in the video bitstream that indicate whether the decoded data is explicitly encoded or inferred from previous data, the bit requirements for high-level syntax elements are reduced, enabling more efficient encoding and decoding.
It improves the encoding and decoding efficiency of video bitstreams, reduces the number of bit repetitive signaling bits, and optimizes the video encoding and decoding process.
Smart Images

Figure CN121814960A_ABST
Abstract
Description
Technical Field
[0001] This embodiment generally involves using advanced syntax elements to signal decoded data. Background Technology
[0002] This section aims to introduce the reader to various aspects of the art that may relate to at least one aspect of the embodiments described below and / or claimed herein. It is believed that this discussion will help provide the reader with background information to facilitate a better understanding of the aspects of at least one embodiment. Therefore, these statements should be understood to be read in this context.
[0003] To achieve high compression efficiency, video encoding and decoding schemes typically employ prediction and transform to fully utilize the spatial and temporal redundancy in the video content. Generally, intra-frame or inter-frame prediction is used to leverage intra-frame or inter-frame correlations. The differences between the original and predicted images of the video (often represented as prediction error or prediction residuals) are then transformed, quantized, and entropy encoded / decoded. To reconstruct the images, the compressed data is decoded through the inverse process corresponding to prediction, transform, quantization, and entropy encoding / decoding. Summary of the Invention
[0004] This section provides a simplified summary of at least one of the embodiments in order to provide a basic understanding of some aspects of this disclosure. This summary is not a broad overview of the embodiments. It is not intended to identify key or essential elements of the embodiments. The following summary presents only some aspects of at least one of the embodiments in a simplified form as a prelude to a more detailed description provided elsewhere in the document.
[0005] According to at least one general aspect of this embodiment, a method for signaling decoded data in a video bitstream is provided, wherein the method includes the step of using syntax elements indicating whether the decoded data is explicitly encoded in the video bitstream or inferred from previous data in the video bitstream.
[0006] According to another general aspect of at least one of the present embodiments, an apparatus is provided for signaling decoded data in a video bitstream, wherein the apparatus includes components for using syntax elements indicating whether the decoded data is explicitly encoded in the video bitstream or inferred from previous data in the video bitstream.
[0007] In one embodiment, the set of decoded data is split into subsets of decoded data, and steps or components for using syntax elements to indicate whether the decoded data is explicitly encoded in the video bitstream or inferred from previous data in the video bitstream.
[0008] In one embodiment, the decoded data is a constraint flag that controls the activation of the encoding / decoding tool, and wherein the syntax elements indicate whether the constraint flag is explicitly encoded in the video bitstream or inferred from previous data in the video bitstream.
[0009] In one embodiment, the decoded data is the decoded data of the header of a set of picture portions, and wherein the syntax elements indicate whether the decoded data of the header of the set of picture portions is explicitly encoded in the video bitstream or inferred from the decoded data of another header of a set of picture portions in the video bitstream.
[0010] In one embodiment, the decoded data of the header of a set of picture portions is a list of reference pictures and the preceding data is a GOP structure and multiple reference frames.
[0011] In one embodiment, the decoded data is stripe parameters and the preceding data is the GOP structure and the order of the decoded images.
[0012] In one embodiment, the decoded data is the decoded data of a chip group and the previous data is the decoded data of another chip group.
[0013] In one embodiment, the decoded data is a GOP structure and the preceding data is a collection of some first decoded images in that GOP and predefined GOP structures.
[0014] According to another general aspect of at least one embodiment, a method and apparatus for encoding or decoding video are provided, including signaling decoded data in a video bitstream according to one of the methods described above.
[0015] According to other general aspects of at least one embodiment, a non-transitory computer-readable storage medium, a computer program product, and a bit stream are provided.
[0016] The specific properties of at least one of the embodiments, as well as other objects, advantages, features, and uses of said at least one of the embodiments, will become apparent from the following description of examples taken in conjunction with the accompanying drawings. Attached Figure Description
[0017] The accompanying drawings illustrate several examples of embodiments. The drawings show: Figure 1 A simplified block diagram of an exemplary encoder according to an embodiment is illustrated; Figure 2 A simplified block diagram of an exemplary decoder according to at least one embodiment is illustrated; Figure 3 The illustration shows an example of an HLS element signaling constraint flag according to the prior art; Figure 4 The illustration shows examples of strips, strips, and bricks according to the prior art; Figure 5 The illustration shows an example of a syntax element for slices according to the prior art; Figure 6-7 An example of a piece-set according to the prior art is illustrated; Figure 7a Examples of sheets, strips, and bricks according to at least one embodiment are illustrated; Figure 8 The illustration shows a flowchart of a method 800 for signaling a set of decoded data according to at least one embodiment; Figure 9 The illustration shows an embodiment according to at least one of the embodiments. Figure 8 Or a flowchart of a variant of method 9; Figure 10 The illustration shows an example of splitting a set into subsets according to at least one embodiment; Figure 11 The illustration shows an example of two hierarchical levels of splitting a set of 16 decoded data. Figure 12 An example of an HLS element signaling constraint flag according to at least one embodiment is illustrated; Figure 13 An example of an HLS element according to at least one embodiment is illustrated; Figure 14-15 The illustration shows an example of a strip segment according to the prior art; Figure 16 The illustration shows a flowchart 1600 of an embodiment of method 800 or 900 when the set of decoded data includes decoded data of the header of a set of image portions; Figure 17 An example of temporal organization of syntax elements associated with a GOP according to at least one embodiment is illustrated; Figure 18 An example of the syntax of stripes according to at least one embodiment is illustrated; Figure 19 A flowchart 1900 illustrates an embodiment of method 800 or 900 when the set of decoded data includes a structure of image groups; Figure 20 The diagram shows Figure 19 Variations of the method; and Figure 21 A block diagram of a computing environment in which aspects of this disclosure can be implemented and executed is illustrated. Detailed Implementation
[0018] This detailed description illustrates the principles of this embodiment. It will therefore be appreciated that those skilled in the art will be able to design various arrangements, although not explicitly described or shown herein, to implement the principles of this embodiment and are included within its scope.
[0019] All examples and conditional language described herein are intended for educational purposes to help readers understand the principles of this embodiment and the concepts contributed by the inventors to the field, and should be interpreted as not being limited to these specific examples and conditions.
[0020] Furthermore, all statements herein recounting the principles, aspects, and embodiments of this disclosure and their specific examples are intended to cover their structural and functional equivalents. Moreover, such equivalents are intended to include both currently known equivalents and those developed in the future, i.e., any element developed that performs the same function, regardless of its structure.
[0021] Therefore, for example, those skilled in the art will recognize that the block diagrams presented herein represent conceptual diagrams of illustrative circuitry illustrating the principles of implementing these embodiments. Similarly, it will be recognized that any flowchart, flow diagram, state transition diagram, pseudocode, etc., represents various processes that can be substantially represented in a computer-readable medium and thus executed by a computer or processor, whether or not such a computer or processor is explicitly shown.
[0022] The present embodiments will be described more fully below with reference to the accompanying drawings, which illustrate examples of the embodiments. However, the embodiments may be implemented in many alternative forms and should not be construed as limited to the examples set forth herein. Thus, it should be understood that the embodiments are not intended to be limited to the specific forms disclosed. Rather, the present embodiments are intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of this application.
[0023] When the accompanying drawings are presented as flowcharts, it should be understood that they also provide block diagrams of the corresponding devices. Similarly, when the accompanying drawings are presented as block diagrams, it should be understood that they also provide flowcharts of the corresponding methods / processes.
[0024] The functionality of the various elements shown in the figure can be provided using dedicated hardware and hardware capable of executing software in association with appropriate software. When provided by a processor, the functionality can be provided by a single dedicated processor, a single shared processor, or multiple separate processors, some of which may be shared. Moreover, the explicit use of the terms "processor" or "controller" should not be construed as specifically referring to hardware capable of executing software, and may implicitly include, but is not limited to, digital signal processor (DSP) hardware, read-only memory (ROM) for storing software, random access memory (RAM), and non-volatile storage devices.
[0025] Other conventional and / or custom hardware may also be included. Similarly, any switches shown in the figure are merely conceptual. Their functionality can be performed through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, with the specific technology to be chosen by the implementer, as understood more specifically from the context.
[0026] Similar or identical elements in the accompanying drawings are referenced using the same reference numerals.
[0027] Some diagrams represent syntax tables widely used in the specifications of video compression standards to define the structure of bitstreams conforming to said video compression standard. In those syntax tables, the term "..." indicates an unchanged portion of the syntax relative to a well-known definition given in the specification of the video compression standard and removed from the diagram for ease of reading. Bold entries in the syntax table indicate that the value used for that entry was obtained by parsing the bitstream. The right column of the syntax table indicates the number of bits used to encode the data for the syntax elements. For example, u(4) indicates that 4 bits are used to encode the data, u(8) indicates 8 bits, and ae(v) indicates a syntax element for context-adaptive arithmetic entropy encoding / decoding.
[0028] In its claims, any element expressed as a component performing a specified function is intended to cover any manner in which that function is performed, including, for example, a) a combination of circuit elements performing that function or b) any form of software, thus including firmware, microcode, etc., combined with appropriate circuitry for performing that software to perform the function. This embodiment as defined by such claims rests on the fact that the functions provided by the various described components are combined and assembled in the manner claimed in the claims. Therefore, any component that can provide those functions is considered equivalent to those shown herein.
[0029] It should be understood that the accompanying drawings and descriptions have been simplified to illustrate elements relevant to a clear understanding of this embodiment, while many other elements found in typical encoding and / or decoding devices have been omitted for clarity.
[0030] It will be understood that while the terms "first" and "second" may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. Various methods have been described above, and each method includes one or more steps or actions for implementing the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of a particular set of steps and / or actions may be modified or combined.
[0031] In the following sections, the terms "reconstructed" and "decoded" are used interchangeably. "Reconstructed" is typically, but not necessarily, used at the encoder end, while "decoded" is used at the decoder end. Furthermore, the terms "encode / decode" and "encode" are used interchangeably. Also, the terms "image," "picture," and "frame" are used interchangeably. Additionally, the terms "encode / decode," "source encode / decode," and "compression" are used interchangeably.
[0032] It should be understood that the terms "one embodiment" or "embodiment" or "one implementation" or "implementation" in this disclosure, as well as other variations thereof, refer to specific features, structures, characteristics, etc., described in connection with the embodiment, which are included in at least one embodiment of this disclosure. Therefore, the appearance of the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in one implementation" and any other variations in different places throughout the specification does not necessarily refer to the same embodiment.
[0033] Furthermore, this embodiment or its claims may refer to "determining" various types of information. Determining or deriving information may include, for example, one or more of the following: estimated information, calculated information, predicted information, or information retrieved from memory.
[0034] Furthermore, this application or its claims may refer to "providing" various types of information. Providing information may include one or more of, for example, output information, stored information, transmitted information, sent information, displayed information, shown information, or moved information.
[0035] Furthermore, this application or its claims or claims can refer to "accessing" various types of information. Accessing information can include, for example, receiving information, retrieving information (e.g., from memory), storing information, processing information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information, or one or more of these.
[0036] Additionally, this application or its claims may refer to "receiving" various types of information. Like "access," receiving is intended to be a broad term. Receiving information may include, for example, accessing information or retrieving information (e.g., from memory) or one or more of these. Furthermore, "receiving" is generally referred to in one or more ways during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0037] Additionally, this application or its claims may refer to "inferring" various data from prior data of a video bitstream. The term "prior data" should be understood as data that has been constructed, reconstructed, and parsed prior to the current data to be inferred. Note that data can be inferred from previously inferred data (recursive inference). A special case of inference is default inference, which uses default data as the inferred data; that is, the "prior" data is the default data. Inferring and deriving from prior data of a video bitstream may include one or more of the following, for example, copying or combining prior data to obtain the inferred data, or accessing information from the prior data that provides instructions for deriving the components used to derive the inferred data.
[0038] Additionally, this application or its claims may refer to "explicit encoding / decoding of decoded data in a video bitstream" and "explicit decoding of decoded data from a video bitstream." Explicit encoding / decoding (decoding) of decoded data refers to adding (obtaining) a dedicated syntax element in (from) the video bitstream. Explicitly encoded data may refer to inputting / inputting at least one binary bit (bin) / information representing this data into an entropy encoding engine (e.g., CABAC), while inferred data refers to not inputting any binary bit / information representing this data into the entropy encoding / decoding engine. Similarly, explicitly decoded data refers to outputting at least one binary bit / information representing this data from an entropy decoding engine (i.e., CABAC), while inferred data refers to not outputting any binary bit / information representing this data into the entropy decoding engine. Note that the information indicating whether data is explicitly encoded / decoded or inferred can be another syntax element that can also be explicitly encoded / decoded or inferred.
[0039] For example, the decoded data can be a stripe parameter or an index of a table of predefined GOP structures. Explicitly encoding a stripe parameter then involves adding syntax elements to the stripe parameter or the index of a table of predefined GOP structures in the video bitstream each time this decoded data is used during the decoding process.
[0040] It should be recognized that the various features shown and described are interchangeable. Unless otherwise stated, a feature shown in one embodiment may be incorporated into another embodiment. Furthermore, features described in the various embodiments may be combined or separated unless otherwise indicated as inseparable or non-combinable.
[0041] As previously stated, the functionality of the various elements shown in the figure can be provided using dedicated hardware and hardware capable of executing software in association with appropriate software. Furthermore, when provided by a processor, the functionality can be provided by a single dedicated processor, a single shared processor, or multiple separate processors, some of which may be shared.
[0042] It should also be understood that, because some of the system components and methods depicted in the accompanying drawings are preferably implemented in software, the actual connections between system components or process functional blocks may vary depending on how the processes of this disclosure are programmed. In view of the teachings herein, those skilled in the art will be able to conceive of these and similar embodiments or configurations of this disclosure.
[0043] While illustrative embodiments have been described herein with reference to the accompanying drawings, it should be understood that this disclosure is not limited to those precise embodiments, and various changes and modifications can be made therein by those skilled in the art without departing from the scope of this disclosure. Furthermore, various embodiments can be combined without departing from the scope of this disclosure. All such changes and modifications are intended to be included within the scope of this disclosure as set forth in the appended claims.
[0044] It should be recognized that the use of any of the following “ / ”, “and / or”, and “at least one of”, for example, in the cases of “A / B”, “A and / or B”, and “at least one of A and B”, is intended to cover selecting only the first listed option (A), or only the second listed option (B), or selecting both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, such wording is intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or selecting all three options (A, B, and C). As will be readily apparent to those skilled in the art and related fields, this can be extended to many of the listed items.
[0045] As will be apparent to those skilled in the art, implementations can generate various signals formatted to carry information, such as information that can be stored or transmitted. This information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described embodiments. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of a spectrum) or baseband signals. Formatting may include, for example, encoding the data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal can be transmitted via a variety of different wired or wireless links. The signal may be stored on a processor-readable medium.
[0046] It should be understood that an image (also referred to as a picture or frame) can be an array of luminance samples in monochrome format, or an array of luminance samples and two corresponding arrays of chrominance samples in 4:2:0, 4:2:2 or 4:4:4 color format, or three arrays of three color components (e.g., RGB).
[0047] In video compression standards, images are divided into blocks that may have different sizes and / or different shapes.
[0048] It should be understood that blocks are two-dimensional arrays or matrices. The horizontal or x-direction (or axis) represents the width, and the vertical or y-direction (or axis) represents the height. Indexing starts from 0. The x-direction represents columns, and the y-direction represents rows. The maximum x-index is width – 1. The maximum y-index is height – 1.
[0049] coding Figure 1 A simplified block diagram of an exemplary encoder 100 according to at least one embodiment is illustrated.
[0050] The encoder 100 can be included in the transmitter or headend of a communication system.
[0051] To encode a video sequence using one or more images, the images are segmented into blocks that may be of different sizes and / or different shapes (module 110). For example, in HEVC ( "ITU-T H.265 TELECOMMUNICATION STANDARDIZATION SSECTOR OF ITU (10 / 2014), SERIES H: AUDIOVISUAL AND MULTIMEDIA SYSTEMS, Infrastructure of audiovisual services–Coding of moving Video, High Effective Video coding, Recommendation ITU-T H.265 (ITU-T H.265) ITU Telecommunications Standardization Sector (10 / 2014), H Series: Audiovisual and Multimedia Systems, Audiovisual Service Infrastructure – Motion Video Coding, High-Efficiency Video Coding, ITU-T H.265 Recommendation), images can be segmented into square CTUs (Codec Tree Units) with configurable sizes. Consecutive sets of CTUs can be grouped into stripes. A CTU is the root of a quadtree, which is divided into blocks represented by codec units (CUs).
[0052] In the exemplary encoder 100, the image is encoded by a block-based encoding module as described below.
[0053] Each block is encoded using either intra-frame prediction mode or inter-frame prediction mode.
[0054] When encoding blocks in intra-prediction mode (module 160), encoder 100 performs intra-prediction (also known as spatial prediction) based on at least one sample of a block in the same image (or based on a predefined value of the first block of the image or strip). As an example, a predicted block is obtained by performing intra-prediction on the block from reconstructed neighboring samples.
[0055] When encoding blocks in inter-frame prediction mode, encoder 100 performs inter-frame prediction (also known as temporal prediction) based on at least one reference block of at least one reference picture or strip (stored in a reference picture buffer).
[0056] Inter-frame predictive coding and decoding are performed by performing motion estimation (module 175) and motion compensation (module 170) on reference blocks stored in reference image buffer 180.
[0057] In inter-frame prediction (also known as one-way prediction) mode, the predicted blocks can generally (but not necessarily) be based on an earlier reference image.
[0058] In dual-frame prediction (also known as dual-prediction) mode, the predicted blocks are typically (but not necessarily) based on earlier and later images.
[0059] Encoder 100 (module 105) determines (whether to use intra-prediction mode or inter-prediction mode) to encode the block and indicates the intra- / inter-prediction decision via prediction mode syntax elements.
[0060] The prediction residual block is calculated by subtracting the prediction block (also known as the predictor) from the block (module 120).
[0061] The predicted residual block is transformed (module 125) and quantized (module 130). Transformation module 125 can transform the predicted residual block from the pixel (spatial) domain to the transform (frequency) domain. This transformation can be, for example, a cosine transform, a sine transform, a wavelet transform, etc. Quantization can be performed according to, for example, a rate distortion criterion (module 130).
[0062] The quantized transform coefficients, along with motion vectors and other syntax elements, are entropy encoded (module 145) to output a bitstream. Entropy encoding can be, for example, context-adaptive binary arithmetic codec (CABAC), context-adaptive variable-length codec (CAVLC), Huffman, arithmetic, exponential Golomb, etc.
[0063] The encoder can also skip the transform and apply quantization directly to the untransformed prediction residual block. Alternatively, the encoder can bypass both the transform and quantization, meaning the prediction residual block is directly encoded and decoded without applying either the transform or quantization process.
[0064] In direct PCM encoding and decoding, no prediction is applied and block samples are directly encoded and decoded into the bitstream.
[0065] Encoder 100 includes a decoding loop and thus decodes the encoded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (also known as inverse quantization) (module 140) and inverse transformed (module 150) to decode the prediction residual block. The block is then reconstructed by combining (module 155) the decoded prediction residual block and the prediction block. One or more in-loop filters (165) can be applied to the reconstructed image, for example, performing deblocking / sample adaptive offset (SAO) filtering to reduce encoding / decoding artifacts. The filtered image is stored in a reference image buffer 180.
[0066] The encoder 100 module can be implemented in software and executed by a processor, or it can be implemented using circuit components well known to those skilled in the art of compression. In particular, the video encoder 100 can be implemented as an integrated circuit (IC).
[0067] decoding Figure 2 A simplified block diagram of an exemplary decoder 200 according to at least one embodiment is shown.
[0068] The decoder 200 can be included in the receiver of the communication system.
[0069] Decoder 200 generally performs and is Figure 1 The encoder 100 described in the document performs the encoding traversal inversely to the decoding traversal, but not all operations in the decoder are the inverse of the encoding process (e.g., intra-frame and inter-frame prediction).
[0070] Specifically, the input to the decoder 200 includes a video bitstream that can be generated by the encoder 100.
[0071] First, entropy decoding (module 230) is performed on the video bitstream to obtain, for example, transform coefficients, motion vectors (MV), image segmentation information, possible prediction mode flags, syntax elements, and other decoded data.
[0072] For example, in HEVC, the image segmentation information indicates the size of the CTU and how the CTU is split into CUs. Therefore, the decoder can divide the image into (235) CTUs based on the image segmentation information, and divide each CTU into CUs.
[0073] The transform coefficients are dequantized (module 240) and inverse transformed (module 250) to decode the predictive residual block. The decoded predictive residual block is then combined with the predictive block (also known as the predictor) (module 255) to obtain the decoded / reconstructed block.
[0074] Depending on the prediction mode flag, prediction blocks (module 205) can be obtained from intra-frame prediction (module 260) or motion-compensated prediction (i.e., inter-frame prediction) (module 270). In-loop filters (module 265) can be applied to the reconstructed image. In-loop filters may include deblocking filters and / or SAO filters. The filtered image is stored in a reference image buffer 280.
[0075] The decoder 200 module can be implemented in software and executed by a processor, or it can be implemented using circuit components well known to those skilled in the art of compression. In particular, the decoder 200 can be implemented as an integrated circuit (IC), either alone or in combination with the encoder 100 as a codec.
[0076] This embodiment aims to provide a mechanism to improve the High-Level Syntax (HLS) for block-based video coding by using fewer bits than the current syntax used in conventional block-based video compression standards.
[0077] HLS is a signaling mechanism that enables bitstream interoperability points defined by parties other than MPEG or ITU, such as DVB, ATSC, and 3GPP. For example, in the VVC standard draft (B. Bross, J. Chen, S. Liu, “Versatile Video Coding (Draft 4)”, JVET document JVET-N1001, 14th meeting: Geneva, March 19-27, 2019), HLS constraint flags allow control over the activation of codec tools in the VVC decoder. HLS constraint flags are grouped into the syntax element “general_constraint_info()” included in the SPS (Sequence Parameter Set) and / or profile and level, VPS (Video Parameter Set) and DPS (Decoding Parameter Set) bitstream sections, such as... Figure 3 As described in [the document]. Constraint flags indicate properties that cannot be violated throughout the video bitstream. When constraint flags are set in an SPS (or similar) syntax structure, the decoder can safely assume that the tool will not be used in the video bitstream.
[0078] The current syntax for encoding / decoding constraint flags uses a large number of bits that could jeopardize the efficiency of video bitstream encoding / decoding. In fact, this information can be repeated at several locations in the video bitstream, as it is typically present in the SPS sent at each Random Access Point (RAP), and can be repeated at other locations, such as also in the Decoder Parameter Set (DPS).
[0079] Another example of HLS is the High-Level Syntax, which groups non-codec image portions such as sequences, pictures, strip or slice headers, Supplemental Enhancement Information (SEI), and Video Usability Information (VUI), excluding codec macroblocks, CTUs, or CUs. In the following text, we will refer to this as "HLS Segmentation," which allows information associated with subsets of macroblocks, CTUs, or CUs to be grouped into pictures, strips, slices, strip headers, or brick headers.
[0080] In the AVC and HEVC video codec standards, advanced segmentation is performed using stripes (continuous CTUs in raster scanning) or slices (uniform or non-uniform slice grids). In the VVC draft standard, advanced segmentation is performed using slices and possibly bricks that vertically divide the slices into rectangular subregions. Figure 4 In the case of uneven tile spacing, the tiles are distributed according to an uneven grid (top right). In this case, in HLS, the values of "tile-column-width" and "tile-column-height" can be signaled. Figure 5 ).
[0081] Several slices can be grouped into a "slice-group". To define a slice-group, the slices are implicitly marked with their raster scan order, such as... Figure 6 As depicted, the top-left tile index (top_left_tile_idx[i]) and bottom-right tile index (bottom_right_tile_idx[i]) are signaled in each tile-group "i" HLS. For example, 8 tiles ( Figure 4 The top right corner of the image has been grouped into 4 "piece-groups" ( Figure 6 For example, the first piece-group 0 consists of pieces 0, 1, and 2, where top_left_tile_idx[0] = 0, bottom_right_tile_idx[0] = 2, and bottom_right_tile_idx_delta[0] = 2 ( bottom_right_ tile_idx_delta[i]=bottom_right_tile_idx[i]- top_left_tile_idx[i] () Figure 7 ).
[0082] In the variant, a slice can be divided into one or more bricks. Each brick consists of multiple CTU rows within the slice. A strip can contain one or more slices or multiple bricks. Two strip modes are supported: raster scan strip mode and rectangular strip mode. In raster scan strip mode, the strip can contain a sequence of slices in the raster scan of the image. In rectangular strip mode, the strip can contain multiple bricks of the image, which together form a rectangular area of the image. The bricks within the rectangular strip can be arranged according to the raster scan order of the strip's bricks. Bricks can be implicitly marked according to the raster scan order, with the brick index in the upper left corner (…). top_left_brick_idx[i] ) and the bottom right brick index ( bottom_right_brick_ idx[i] This can be signaled in the HLS used for each stripe "i". For example, the first stripe 0 ( Figure 7a The left side of the middle) contains 2 pieces and 3 bricks. It consists of "brick_idx" 0, 1 and 2, top_left_brick_idx[0]=0, bottom_right_brick_idx[0]=2 and bottom_right_brick_idx_delta[0]=2 ( bottom_right_ brick_idx_delta[i]=bottom_right_brick_idx[i]- top_left_brick_idx_brick ). second Band 1 ( Figure 7a The top right corner contains 2 bricks. It consists of "brick_idx" 3 and 4, top_left_brick_idx[0]=3, bottom_right_brick_idx[0]=4 and bottom_right_brick_idx_delta[0]=1. The third article With 2 ( Figure 7a The middle right (in the middle) contains 3 bricks. It consists of “brick_idx” 5, 6 and 7, top_left_brick_idx[0]=5, bottom_right_brick_idx[0]=7 and bottom_right_brick_idx_delta[0]=2. The fourth strip 3 ( Figure 7a The bottom right corner contains 3 bricks. It consists of "brick_idx" 8, 9, and 10, top_left_brick_idx[0]=8, bottom_right_brick_idx[0]=10, and bottom_right_brick_idx_delta[0]=2. Note that the names of these pieces, bricks, and groups of pieces can be changed, but the basic principles remain the same.
[0083] In a video bitstream, the strip header, slice header, etc., are repeated at least once for each encoded and decoded image, even though they share the same values in the temporal domain for most of the time domain.
[0084] This embodiment signals the decoded data in the video bitstream by using a syntax element that indicates whether the decoded data is encoded or decoded in the video bitstream or inferred from previous data in the video bitstream.
[0085] Figure 8 A flowchart illustrating a method 800 for signaling a set of decoded data according to at least one embodiment is shown.
[0086] In step 810, the method can check whether the decoded data of the set is explicitly encoded in the video bitstream or inferred from previous data PDs in the video bitstream. A syntax element F can be added to the video bitstream. If the decoded data of the set is inferred from previous data PDs (step 830), then the syntax element F indicates that the decoded data of the set is inferred from previous data PDs. If the decoded data of the set is explicitly encoded in the video bitstream, then (step 820) the syntax element F indicates that the decoded data of the set is explicitly encoded in the video bitstream.
[0087] In step 840, the method can access syntax element F from the video bitstream.
[0088] In step 850, the method may check whether the syntax element F indicates that the decoded data of the set of decoded data is inferred from previous data PD of the video bitstream. In that case (step 870), the decoded data of the set is inferred from the previous data PD. If the syntax element F indicates that the decoded data of the set is explicitly encoded and decoded in the video bitstream, then (step 860), the decoded data of the set is explicitly decoded from the video bitstream.
[0089] Figure 9 The diagram illustrates when a syntax element indicates that the decoded data for that set is explicitly encoded or decoded in the video bitstream. Figure 8 A flowchart of a variant of the method.
[0090] In step 910, as Figure 10 As shown, the set of decoded data is split into subsets, which are further divided into 4 subsets i.
[0091] Step 910 is followed by method 800 ( Figure 8 The iterations of steps 810-830 are performed. In each iteration, the current subset i of the decoded data is considered as input to step 810 (step 920). During this process, the method can check whether the decoded data of the current subset i is explicitly encoded in the video bitstream or inferred from previous data PDi in the video bitstream. A syntax element Fi can be added to the video bitstream. If the decoded data of the current subset i is inferred from previous data PDi (step 830), then the syntax element Fi indicates that the decoded data of the current subset i is inferred from previous data PDi. If the decoded data of the current subset i is explicitly encoded in the video bitstream, then (step 820) the syntax element Fi indicates that the decoded data of the subset i is explicitly encoded in the video bitstream.
[0092] When considering all subsets i, method 800 ( Figure 8Steps 840-870 are iterated. In each iteration, the current subset i of the decoded data is considered as input to step 840 (step 930), during which the method can access the syntax element Fi of the current subset i from the video bitstream.
[0093] In step 850, the method checks whether the syntax element Fi indicates that the decoded data of subset i is inferred from previous data PDi of the video bitstream. In that case (step 870), the decoded data of subset i is inferred from previous data PDi of the video bitstream. If the syntax element Fi indicates that the decoded data of subset i is explicitly encoded in the video bitstream, then (step 860) the decoded data of subset i is explicitly decoded from the video bitstream.
[0094] Index i (Fi, FDi) indicates the syntax element and previous data dedicated to the decoded data of subset i. A certain Fi and / or FDi can be the same.
[0095] Figure 10 The illustration shows the use of method 800 when the set of decoded data is split into subsets.
[0096] In a variant, method 800 can be used when the set of decoded data is recursively split to create multiple hierarchical levels as follows: When the decoded data of a subset at hierarchical level L, as indicated by the syntax element, is explicitly encoded and decoded in the video bitstream, the subset at hierarchical level L is split into subsets at hierarchical level (L+1) (step 830). The recursive splitting stops when conditions such as, for example, the maximum number of hierarchical levels or the minimum cardinality of the subsets are met.
[0097] Once all subsets of the decoded data for hierarchical level L have been considered, method 900 is iteratively applied to a subset of the new hierarchical level (L+1) (if it exists).
[0098] Figure 11 The illustration shows an example of two hierarchical levels of splitting a set of 16 decoded data.
[0099] At the first hierarchical level 0, the set is first split into two subsets (i=0 and i=1) of 8 decoded data (810). At step 820, a syntax element is added to the bitstream to indicate that the subset of decoded data (i=0) is split into, for example, two sub-subsets (i=00 and i=01). Another syntax element is also added to the video bitstream to indicate that the decoded data of the second subset (i=1) is inferred from previous data in the video bitstream (step 830) and the second subset (i=1) is not split.
[0100] In the second layer level 1, another syntax element is added to the video bitstream to indicate that the decoded data of the subset (i=00) is inferred from previous data in the video bitstream (step 830) and the subset (i=00) is not split. Another syntax element is also added to the video bitstream to indicate that the decoded data of the subset (i=01) is inferred from previous data in the video bitstream (step 830) and the subset (i=01) is not split.
[0101] In one embodiment, the set of decoded data may include constraint flags that control the activation of the encoding / decoding tools. The syntax element F (or Fi) can then indicate whether the constraint flags of the set (or a subset of the set) are explicitly encoded / decoded in the video bitstream or inferred from previous data PD of the video bitstream.
[0102] Figure 12 An example of the HLS element "general_constraint_info" according to at least one embodiment is illustrated.
[0103] The syntax element "general_constraint_info" includes constraint flags grouped together under the syntax element "use_default_constraint_flag" (syntax element F or Fi). This syntax element "use_default_constraint_flag" can be false when these constraint flags are explicitly encoded or decoded in the video bitstream (step 820) or explicitly decoded from the video bitstream (step 860), or true when all these constraint flags are inferred from previous data PD of the video bitstream (steps 830, 870).
[0104] Note that when constraint flags are inferred, other constraint flags can be explicitly encoded / decoded (not shown here).
[0105] In the variant, the previous data PD is a single binary value (false or true) and all inferred constraint flag values are equal to this single binary value FD.
[0106] In the variant, the preceding data PD depends on the video bitstream's decoding profile, level, and / or at least one decoding parameter. For example, for a "high profile," all inferred constraint flag values are equal to false (all tools may be enabled), while for a "low profile," a predetermined subset of constraint flag values are inferred as true (tools are disabled) and other constraint flag values are inferred as false (tools may be enabled).
[0107] In the variant, the state of syntax element F (or Fi) determines whether the decoding parameters exist in the video bitstream.
[0108] For example, Figure 13 The set of syntax elements "profile_tier_level()" includes the constraint flag "general_constraint_info()" under the syntax element "use_default_constraint-flag" (syntax element F or Fi). Figure 13 In the context of the code, if the syntax element "use_default_constraint_flag" is true, then the set of syntax elements "general_constraint_info()" is explicitly encoded into the video bitstream. If "use_default_constraint_flag" is false, then this set of syntax elements does not exist and the syntax elements are inferred.
[0109] exist Figure 12 In the variant shown, the syntax element "use_default_constraint_flag" is encoded in "general_constraint_info()" instead of "profile_tier_level()". If the syntax element "use_default_constraint_flag" is false, then a set of syntax elements exists in the video bitstream (explicitly encoded) indicating whether some tools may be enabled or disabled in the bitstream; otherwise, this set of syntax elements does not exist and the syntax elements are inferred.
[0110] In block-based video codec specifications, the video bitstream is divided into encoded and decoded video sequences (CVS), which contain one or more access units (AUs) in decoding order. An AU contains exactly one encoded and decoded picture.
[0111] For example, in HEVC, an AU is segmented into stripes. Some decoded data in the headers (or slice headers, or slice group headers) of different stripes is the same for several stripes. For that reason, in HEVC, the first strip in the image must be independent, while subsequent stripes can be subordinate, such as... Figure 14 As depicted in [the image]. The subordinate strip header shares common decoding data with the first independent strip in the same image ([the image is described in]). Figure 15 ).
[0112] Figure 16 The illustration shows a flowchart 1600 of an embodiment of method 800 or 900 when the set of decoded data includes headers of a set of image portions (such as, for example, strip headers and / or slice headers and / or slice-group headers of an image). In HEVC, strips have headers and may contain slices, but slices do not have headers. In other codecs, strips are replaced by slices, and slices have headers.
[0113] In HEVC, examples of such decoding can include at least one of the following decoded data (strip header): - Post filter flags, such as (slice_sao_luma_flag, slice_sao_chroma_flag, slice_alf_luma_flag, slice_alf_chroma_flag).
[0114] - See the list of reference images. - weighted_pred / bipred_flag: Enable or disable WP (weighted prediction) in the stripe. - Divide into strips, slices or slice-groups, or bricks (e.g., num_bricks_in_slice_minus1). - Slice_type: I (intra-frame only), P (one-way), or B (two-way prediction). - pic_output_flag: Whether the current image will be output. - colour_plane_flag: Whether the color channels are encoded and decoded separately. - POC: Image Sequence Counting - Long-term reference information: image index and POC - slice_temporal_mvp_enabled_flag: Whether the temporal motion vector prediction tool is enabled. - long-term-refinfo (POC) - slice_qp_delta: Starts with a QP (quantization parameter) that differs from the init_QP specified in PPS. - Deblocking filter parameters (beta, tc).
[0115] In other codecs, this data can be included in the slice header. In the following text, we consider segmenting an image into multiple regions (or image parts), each region consisting of several blocks (e.g., CTUs or macroblocks). Depending on the codec's segmentation topology, the region can be, for example, a strip, a slice, a slice-group, or a brick. Some of these can be associated with a header containing information related to the blocks that make up the region.
[0116] In step 1610 ( Figure 8In the embodiment of step 810, the method can check whether the decoded data of the header of a group of picture portions (or regions) is explicitly encoded and decoded in the video bitstream or inferred from previous data IPH (“inter_picture-header”) in the video bitstream. A syntax element F (or Fi) denoted as “infer_from_inter_picture_header_flag” can be added to the video bitstream. If the decoded data of the header of the picture portion group is inferred from previous data IPH, then in step 1630 (… Figure 8 In the embodiment of step 830), the syntax element F (or Fi) indicates that the header of the picture portion group is inferred from the decoded data (previous data IPH) of another header of a group of picture portions. If the decoded data of the header of the picture portion group is explicitly encoded and decoded in the video bitstream, then in step 1520 ( Figure 8 In the embodiment of step 820, the syntax element F (or Fi) can indicate that the decoded data of the header of the picture portion group is explicitly encoded and decoded in the video bitstream.
[0117] Previous data IPH grouped decoded data, for example, multiple strip headers and / or multiple common slice headers and / or multiple slice-group headers of the same image.
[0118] It is inferred that the common decoding data defined in the header of a typical AU, CTU / CU, chip, or chip group reduces the bandwidth used for video transmission.
[0119] For example, in each AU, and possibly in each stripe within that AU, the syntax element "infer_from_inter_picture_header_flag" can indicate whether the decoded data is explicitly defined in the AU (in each stripe within that AU), such as in the stripe header, slice, or slice-group header, or whether it is inferred from previous data IPH and thus can be shared by several picture segments that may be several pictures.
[0120] In step 1640 ( Figure 8 In the embodiment of step 840, the method can access syntax element F (or Fi) from the video bitstream.
[0121] In step 1650 ( Figure 8 In step 850 of the embodiment, the method may check whether the syntax element F (or Fi) indicates that the decoded data of the header of a set of picture portions is inferred from previous data IPH in the video bitstream. Then, in step 1670 ( Figure 8In the embodiment of step 870, the decoded data of the header of the picture portion group is inferred from the previous data IPH. If the syntax element F (or Fi) indicates that the decoded data of the header of the picture portion group is encoded and decoded in the video bitstream, then in step 1660 ( Figure 8 In the embodiment of step 860, the decoding data of the header of the image portion group is explicitly decoded from the video bitstream.
[0122] Advantageously, the preceding data IPH is located in the bitstream associated with the GOP (Group of Pictures) after the RAP syntax element and before the AU syntax element relative to the RAP. The decoded data for the AU syntax element can then be inferred from the preceding data IPH, such as... Figure 17 As shown in the image.
[0123] In one embodiment of method 1600, a list of reference pictures (decoded data of the header of a set of picture portions) can be inferred from the GOP structure and multiple reference frames (previous data IPH).
[0124] In one embodiment of method 1600, stripe parameters (e.g., stripe type) can be inferred from the GOP structure and the order of the decoded images (previous data IPH).
[0125] In one embodiment, the decoded data of the slice group controlling the encoding / decoding of the slice group can be inferred from the decoded data of another slice group (previous data IPH) in the video bitstream.
[0126] Inferring fragment information from previous data in the video bitstream reduces the bandwidth used to transmit images.
[0127] In the variant, a slice group comprises multiple slices of a picture or strip (decoded data denoted as "num-tiles"). Syntax elements (F, Fi, ...) can then indicate whether the multiple slices i are encoded or decoded in the video bitstream or inferred from previous data IPH representing multiple slices per column (denoted as "num_tile_columns_minus1") and multiple slices per row (denoted as "num_tile_rows_minus1").
[0128] For example, "num-tiles" can be inferred from "num_tile_columns_minus1" and "num_tile_rows_minus1" as follows: num_tiles=(num_tile_columns_minus1 + 1) × (num_tile_rows_minus1 +1).
[0129] In an embodiment, the tile group also includes the "bottom-right-tile-idx" value or "bottom-right-tile-idx-delta" value of the last tile group of the image.
[0130] The final slice-group (decoded data) can be inferred from the number of slices (previous data in the video bitstream). bottom- right-tile-idx value or bottom-right-tile-idx-delta The values (decoded data) are as follows: bottom_right_tile_idx[num_tile_groups_in_pic_minus1]=num_tile– 1 bottom_right_tile_idx_delta[num_tile_groups_in_pic_minus1]=num_tile- 1 - top_left_tile_idx[num_tile_groups_in_pic_minus1].
[0131] like Figure 18 As shown, one of the advantages of this embodiment is that it stores the bits used for encoding and decoding "bottom_right_tile_idx_delta[num_tile_groups_in_pic_minus1]" in the syntax element "pic_parameter_set_rbsp".
[0132] Similarly, the number of bricks in an image (NumBricksInPic) can be inferred from previous IPH data. In the case of rectangular stripe mode, the stripes can also include the "bottom-right-brick-idx" value or the "bottom-right-brick-idx-delta" value of the stripes in the image.
[0133] The last strip (decoded data) bottom-right-brick-idx value or bottom-right-brick- idx-delta The value (decoded data) can be inferred from the number of bricks (previous data from the video bitstream) as follows: bottom_right_brick_idx[num_slices_in_pic_minus1]=NumBricksInPic–1 bottom_right_brick_idx_delta[num_slices_in_pic_minus1]=NumBricksInPic- 1 - top_left_brick_idx[num_slices_in_pic_minus1] One advantage of this embodiment is that it stores the bits used for encoding and decoding "bottom_right_brick_idx_delta[num_slices_in_pic_minus1]" in the syntax element "pic_parameter_set_rbsp".
[0134] Figure 19 The flowchart 1900 illustrates an embodiment of method 800 or 900 when the set of decoded data includes a structure of picture groups (GOPs).
[0135] In step 1910 ( Figure 8 In the embodiment of step 810, the method can check whether the GOP structure is explicitly encoded in the video bitstream or inferred from previous data PD in the video bitstream. A syntax element F (or Fi) denoted as "GOP_structure_indicator" can be added to the video bitstream. If the GOP structure is inferred from previous data PD, then in step 1930 (… Figure 8 In step 830 (an embodiment), the syntax element F (or Fi) indicates that the GOP structure is inferred from the previous data PD. If the GOP structure is explicitly encoded and decoded in the video bitstream, then in step 1920 ( Figure 8 In the embodiment of step 820, the syntax element F (or Fi) can indicate that the GOP structure is explicitly encoded and decoded in the video bitstream.
[0136] In the variant, when the GOP structure is inferred from previous data PD in the video bitstream, the syntax element F is not added to the video bitstream.
[0137] It is inferred that the GOP structure reduces the bandwidth used for video transmission.
[0138] In step 1940 ( Figure 8 In the embodiment of step 830, the method can access syntax element F (or Fi) from the video bitstream.
[0139] In a variant, the method checks whether the syntax element F (or Fi) exists in the video bitstream.
[0140] In step 1950 ( Figure 8In step 850 of the embodiment, the method may check whether the syntax element F (or Fi) indicates that the GOP structure is inferred from previous data PD of the video bitstream, or, depending on the variant, whether the syntax element F (or Fi) is not present in the video bitstream. Then, in step 1970 ( Figure 8 In step 870 (an embodiment), the GOP structure is inferred from the previous data PD. If syntax element F (or Fi) indicates that the GOP structure is explicitly encoded in the video bitstream, or depending on the variant, syntax element F (or Fi) is present in the video bitstream, then in step 1960 ( Figure 8 In step 860 of the embodiment, the GOP structure is explicitly decoded from the video bitstream.
[0141] In one embodiment, the syntax element F (or Fi) can signal the index of the predefined GOP structure of the set of predefined GOP structures.
[0142] Examples of predefined GOP structures can be random access (RA: layered encoding / decoding), low-latency B (LB), and low-latency P (LP), such as... Figure 20 As shown in the image.
[0143] exist Figure 20 In the variant of method 1900 shown, additional syntax elements such as GOP length can be used to precisely signal which of the predefined GOP structures.
[0144] In one embodiment, the prior data PD includes a GOP and some first-decoded images from the predefined set of GOP structures.
[0145] For example, based on Figure 20 The predefined GOP structure is related to the predefined GOP structure with index "1" if the POC of the first decoded image is 0-16-8-4.
[0146] Figure 21 The diagram illustrates a block diagram of an example system that implements various aspects and embodiments.
[0147] System 2100 can be implemented as a device comprising the various components described below and configured to perform one or more of the aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 2100 can be implemented individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 2100 are distributed across multiple ICs and / or discrete components. In various embodiments, system 2100 is communicatively coupled to one or more other similar systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 2100 is configured to implement one or more of the aspects described in this document.
[0148] System 2100 includes at least one processor 2110 configured to execute instructions loaded therein for implementing various aspects, such as those described in this document. Processor 2110 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 2100 includes at least one memory 2120 (e.g., a volatile memory device and / or a non-volatile memory device). System 2100 includes a storage device 2140, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 2140 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.
[0149] System 2100 includes an encoder / decoder module 2130 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 2130 may include its own processor and memory. The encoder / decoder module 2130 represents one or more modules that can be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both encoding and decoding modules. Furthermore, the encoder / decoder module 2130 may be implemented as a separate element of system 2100, or may be incorporated into processor 2110 as a combination of hardware and software known to those skilled in the art. Program code that can be loaded onto processor 2110 or encoder / decoder 2130 to perform the various aspects described in this document may be stored in storage device 2140 and subsequently loaded onto memory 2120 for execution by processor 2110. According to various embodiments, during the execution of the processes described in this document, one or more of the processor 2110, memory 2120, storage device 2140, and encoder / decoder module 2130 may store one or more of a variety of items. Such stored items may include, but are not limited to, input video, decoded video or a portion of decoded video, bitstream, matrix, variable, and intermediate or final results of equations, formulas, operations, and operational logic.
[0150] In some embodiments, the memory within the processor 2110 and / or encoder / decoder module 2130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, external memory (e.g., the processing device may be either the processor 2110 or the encoder / decoder module 2130) is used for one or more of these functions. External memory may be memory 2120 and / or storage device 2140, such as volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory (such as RAM) is used as working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG stands for Moving Picture Experts Group; MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC stands for High Efficiency Video Codec, also known as H.265 and MPEG-H Part 2), or VVC (Universal Video Codec, a new standard developed by the Joint Video Experts Group JVET).
[0151] As shown in box 2230, input to the components of system 2100 can be provided through various input devices. Such input devices include, but are not limited to, (i) an RF section that receives radio frequency (RF) signals transmitted over the air, for example by a broadcaster, (ii) a component (COMP) input terminal (or a collection of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Figure 21 Other examples not shown include composite video.
[0152] In various embodiments, the input device of block 2230 has associated corresponding input processing elements, as known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) further limiting the frequency band to a narrower band to select, for example, a signal band that may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section in various embodiments includes one or more elements performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include a tuner performing various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to a desired frequency band. Various embodiments rearrange the order of the aforementioned (and other) components, remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as, for example, inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.
[0153] Furthermore, the USB and / or HDMI terminals may include corresponding interface processors for connecting system 2100 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented as needed, for example, within a separate input processing IC or within processor 2110. Similarly, various aspects of USB or HDMI interface processing may be implemented as needed within a separate interface IC or within processor 2110. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 2110, and an encoder / decoder 2130 operating in conjunction with memory and storage elements, to process the data stream as needed for presentation on the output device.
[0154] Various components of system 2100 can be provided within an integrated housing. Within the integrated housing, various components can be interconnected and data can be transferred between them using suitable connection arrangements (e.g., internal buses known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards).
[0155] System 2100 includes a communication interface 2150 that enables communication with other devices via a communication channel 2160. The communication interface 2150 may include, but is not limited to, a transceiver configured to transmit and receive data on the communication channel 2160. The communication interface 2150 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 2160 may be implemented, for example, within a wired and / or wireless medium.
[0156] In various embodiments, a wireless network such as Wi-Fi (e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)) is used to stream or otherwise provide data to system 2100. In these embodiments, Wi-Fi signals are received via a communication channel 2160 and a communication interface 2150 suitable for Wi-Fi communication. The communication channel 2160 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box that delivers data via an HDMI connection of input block 2130 to provide streaming data to system 2100. Still other embodiments use an RF connection of input block 2130 to provide streaming data to system 2100. As indicated above, various embodiments provide data in a non-streaming manner. Furthermore, various embodiments use wireless networks other than Wi-Fi (e.g., cellular networks or Bluetooth networks).
[0157] System 2100 can provide output signals to various output devices, including display 2200, speaker 2210, and other peripheral devices 2220. Display 2200 in various embodiments includes one or more of, for example, touchscreen displays, organic light-emitting diode (OLED) displays, flexible displays, and / or foldable displays. Display 2200 can be used in televisions, tablets, laptops, cellular phones (mobile phones), or other devices. Display 2200 can also be integrated with other components (e.g., as in a smartphone) or separate (e.g., as in an external monitor for a laptop computer). In various examples of embodiments, other peripheral devices 2220 include stand-alone digital video discs (or digital multifunction discs) (DVR, for both terms), disc players, stereo systems, and / or lighting systems. Various embodiments use one or more peripheral devices 2220 that provide functionality based on the output of system 2100. For example, a disc player performs the function of playing the output of system 2100.
[0158] In various embodiments, control signals communicate between system 2100 and display 2200, speaker 2210, or other peripheral devices 2220 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols enabling device-to-device control, with or without user intervention. Output devices may be communicatively coupled to system 2100 via dedicated connections through corresponding interfaces 2170, 2180, and 2190. Alternatively, output devices may be connected to system 2100 via communication interface 2150 using communication channel 2160. Display 2200 and speaker 2210 may be integrated with other components of system 2100 into a single unit in an electronic device, such as a television set. In various embodiments, display interface 2170 includes a display driver, such as, for example, a timing controller (TCon) chip.
[0159] Display 2200 and speaker 2210 may alternatively be separate from one or more other components, for example, if the RF section of input 2230 is part of a separate set-top box. In various embodiments where display 2200 and speaker 2210 are external components, output signals may be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
[0160] The embodiments described herein can be implemented, for example, as methods or processes, apparatus, software programs, data streams, or signals. Even if discussed only in a single form of implementation (e.g., discussed only as a method), the embodiments of the discussed features can also be implemented in other forms (e.g., apparatus or program). Apparatus can be implemented, for example, in suitable hardware, software, and firmware. Methods can be implemented in, for example, apparatus (such as, for example, a processor), which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, computers, mobile phones, portable / personal digital assistants (PDAs), and other devices that facilitate information communication between end users.
[0161] According to one aspect of this embodiment, an apparatus 2100 for video encoding and / or decoding is provided, the apparatus including a processor 2110 and at least one memory 2120, 2140 coupled to the processor, the processor 2110 being configured to perform any of the embodiments of methods 800, 900, 1600 and / or 1700 described above.
[0162] According to one aspect of this disclosure, an apparatus for video encoding and / or decoding is provided, the apparatus including a component for using a syntax element indicating whether decoded data is explicitly encoded or decoded in a video bitstream or inferred from previous data in the video bitstream. Figure 1 The video encoder may include the structure or components of the device. The device for video encoding may perform any of the embodiments of any of methods 800, 900, 1600, and 1700.
[0163] As will be apparent to those skilled in the art, implementations can generate various signals formatted to carry, for example, information that can be stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described embodiment. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding the data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal can be transmitted via a variety of different wired or wireless links. The signal may be stored on a processor-readable medium.
[0164] Furthermore, any of methods 800, 900, 1600, and / or 1700 can be implemented as a computer program product (independently or jointly) including computer-executable instructions that can be executed by a processor. The computer program product having computer-executable instructions can be stored in the corresponding temporary or non-temporary computer-readable storage medium of system 2100, encoder 100, and / or decoder 200.
[0165] It is important to note that one or more elements in processes 800, 900, 1600, and / or 1700 may be combined, executed in a different order, or excluded in some embodiments, while still achieving aspects of this disclosure. Other steps may be executed in parallel, wherein the processor does not wait for one step to be fully completed before starting another step.
[0166] Furthermore, aspects of this embodiment can take the form of a computer-readable storage medium. Any combination of one or more computer-readable storage media can be utilized. A computer-readable storage medium can take the form of a computer-readable program product implemented in one or more computer-readable media and having computer-executable computer-readable program code implemented thereon. Given its inherent ability to store information therein and to provide information retrieval therefrom, a computer-readable storage medium as used herein is considered a non-transitory storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof.
[0167] It should be recognized that while the following list provides more specific examples of computer-readable storage media to which the present disclosure may be applied, it is merely illustrative and not exhaustive, as will be readily apparent to those skilled in the art. The list of examples includes portable computer floppy disks, hard disks, ROMs, EPROMs, flash memory, portable optical disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0168] According to one aspect of this embodiment, a computer-readable storage medium carrying a software program is provided, including program code instructions for performing any embodiment of any method of this embodiment (including method 800, 900, 1600 and / or 1700).
Claims
1. A method for decoding video data, comprising: Decode the syntax element, which indicates whether the data in the stripe header is explicitly encoded in the stripe header or inferred from the data in another header shared by multiple stripes; Obtain the data of the stripe header based on the syntax elements; as well as The block is decoded based on the data in the strip header.
2. The method according to claim 1, wherein, In the bitstream, the other header precedes the strip header.
3. The method according to claim 1, wherein, The other head is the image header.
4. The method according to claim 1, wherein, The data of the strip head is inferred from the data of the other head by copying the data of another head shared by multiple stripes into the data of the strip head.
5. An apparatus comprising one or more processors and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to perform: Decode the syntax element, which indicates whether the data in the stripe header is explicitly encoded in the stripe header or inferred from the data in another header shared by multiple stripes; Obtain the data of the stripe header based on the syntax elements; as well as The block is decoded based on the data in the strip header.
6. The apparatus according to claim 5, wherein, In the bitstream, the other header precedes the strip header.
7. The apparatus according to claim 5, wherein, The other head is the image header.
8. The apparatus according to claim 5, wherein, The data of the strip head is inferred from the data of the other head by copying the data of another head shared by multiple stripes into the data of the strip head.
9. A method for encoding video data, comprising: Obtain the data from the strip header; Encode blocks based on data from the strip header; as well as The syntax element is encoded to indicate whether the data in the strip header is explicitly encoded in the strip header or inferred from the data in another header shared by multiple stripes.
10. The method according to claim 9, wherein, In the bitstream, the other header precedes the strip header.
11. The method according to claim 9, wherein, The other head is the image header.
12. The method according to claim 9, wherein, The data of the strip head is inferred from the data of the other head by copying the data of another head shared by multiple stripes into the data of the strip head.
13. An apparatus comprising one or more processors and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to perform: Obtain the data from the strip header; Encode blocks based on data from the strip header; as well as The syntax element is encoded to indicate whether the data in the strip header is explicitly encoded in the strip header or inferred from the data in another header shared by multiple stripes.
14. The apparatus according to claim 13, wherein, In the bitstream, the other header precedes the strip header.
15. The apparatus according to claim 13, wherein, The other head is the image header.
16. The apparatus according to claim 13, wherein, The data of the strip head is inferred from the data of the other head by copying the data of another head shared by multiple stripes into the data of the strip head.
17. A non-transitory computer-readable storage medium having instructions stored thereon for performing the following operations: Decode the syntax element, which indicates whether the data in the stripe header is explicitly encoded in the stripe header or inferred from the data in another header shared by multiple stripes; Obtain the data of the stripe header based on the syntax elements; as well as The block is decoded based on the data in the strip header.
18. The non-transitory computer-readable storage medium according to claim 17, wherein, In the bitstream, the other header precedes the strip header.
19. The non-transitory computer-readable storage medium according to claim 17, wherein, The other head is the image header.
20. The non-transitory computer-readable storage medium according to claim 17, wherein, The data of the strip head is inferred from the data of the other head by copying the data of another head shared by multiple stripes into the data of the strip head.
21. A non-transitory computer-readable storage medium having instructions stored thereon for performing the following operations: Obtain the data from the strip header; Encoding blocks based on data from the stripe header; and The syntax element is encoded to indicate whether the data in the strip header is explicitly encoded in the strip header or inferred from the data in another header shared by multiple stripes.
22. The non-transitory computer-readable storage medium according to claim 21, wherein, In the bitstream, the other header precedes the strip header.
23. The non-transitory computer-readable storage medium according to claim 21, wherein, The other head is the image header.
24. The non-transitory computer-readable storage medium according to claim 21, wherein, The data of the strip head is inferred from the data of the other head by copying the data of another head shared by multiple stripes into the data of the strip head.