Tile and sub-picture partitioning
By independently decoding sub-images with unique tile groups and filtering within rectangular boundaries, the method addresses decoding challenges in video coding standards, improving efficiency and parallel processing in video streaming applications.
Patent Information
- Application Number
- JP2025067774
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-03-11
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2040-03-11
AI Technical Summary
Existing video coding standards face challenges in efficiently decoding and encoding sub-images and tiles within a video stream, particularly in scenarios requiring independent decoding and parallel processing without inter-dependencies across tile boundaries.
The method involves decoding sub-images independently using sub-image and tile partitioning, with each sub-image having its own tile groups and scan order, and implementing specific filtering and prediction operations only within rectangular tile groups to avoid complexity.
This approach enables efficient decoding and encoding of sub-images and tiles, supporting parallel processing and reducing dependencies across tile boundaries, enhancing performance in applications like immersive media and VR360 content.
Smart Images

Figure 2025105650000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the priority of U.S. Provisional Patent Application No. 62 / 816,846, filed on March 11, 2019, the disclosure of which is hereby incorporated by reference in its entirety into this specification.
[0002] This disclosure relates to a series of advanced video coding techniques. More particularly, it relates to tile and sub - image segmentation designs.
Background Art
[0003] A video coding standard, informally known as Versatile Video Coding (VVC), is under development.
Summary of the Invention
Problems to be Solved by the Invention
[0004] Some embodiments of this disclosure address problems related to VVC and other problems.
Means for Solving the Problems
[0005] In some embodiments, a method is provided. The method may be executed by at least one processor to decode a sub - bitstream of an encoded video stream, where the encoded video stream includes an encoded version of a first sub - image and a second sub - image of an image. The method includes receiving the sub - bitstream and decoding the first sub - image of the image independently of the second sub - image using sub - image and tile partitioning, wherein (i) the first sub - image includes a first rectangular region of the image, the second sub - image includes a second rectangular region of the image, the second rectangular region is different from the first rectangular region, (ii) the first sub - image and the second sub - image each include at least one tile, and (iii) the first sub - image and the second sub - image do not share a common tile.
[0006] In one embodiment, the method further includes the step of decoding a second sub-image of the image using the sub-image and tile partitioning, independent of the first sub-image, wherein at least one tile of the first sub-image is one of a first plurality of tiles, and at least one tile of the second sub-image is one of a second plurality of tiles. In one embodiment, the step of decoding the first sub-image is performed in a tile scan order different from the step of decoding the second sub-image. In one embodiment, the steps of decoding the first sub-image and decoding the second sub-image are performed using the sub-image and tile partitioning, the first plurality of tiles of the first sub-image are classified into at least one first tile group, the second plurality of tiles of the second sub-image are classified into at least one second tile group, and no tile of at least one first tile group is placed in at least one second tile group. In one embodiment, one of the at least one first tile groups is a non-rectangular tile group. In one embodiment, the first sub-image is decoded according to a decoding technique, and tile group level loop filtering control at the boundary between two tile groups of at least one first tile group is possible only if each of the two tile groups is rectangular.
[0007] In one embodiment, the step of receiving the sub-bitstream includes receiving an encoded video stream, the encoded video stream including a sequence parameter set (SPS) having information about a method of partitioning sub-images of an image, the image including the first sub-image and the second sub-image. In one embodiment, the received encoded video stream includes a picture parameter set (PPS) having information about a method of partitioning tiles of an image, the image including at least one tile of the first sub-image and at least one tile of the second sub-image. In one embodiment, the received encoded video stream includes an active parameter set (APS) signaling the adaptive loop filter (ALF) coefficients of the first sub-image.
[0008] In some embodiments, a decoder is provided. The decoder is for decoding a sub-bitstream of an encoded video stream, where the encoded video stream includes an encoded version of a first sub-image and a second sub-image of an image. The decoder may include a memory configured to store computer program code, and at least one processor configured to receive the sub-bitstream, access the computer program code, and operate as instructed by the computer program code. The computer program code may include decoding code configured to cause the at least one processor to decode the first sub-image of the image independently of the second sub-image using the sub-image and tile partitioning, where (i) the first sub-image includes a first rectangular region of the image, the second sub-image includes a second rectangular region of the image, the second rectangular region is different from the first rectangular region, (ii) the first sub-image and the second sub-image each include at least one tile, and (iii) the first sub-image and the second sub-image do not share a common tile.
[0009] In one embodiment, the decoding code is further configured to cause the at least one processor to decode the second sub-image of the image independently of the first sub-image using the sub-image and tile partitioning and At least one tile of the first sub-image is a first plurality of tiles, and at least one tile of the second sub-image is a second plurality of tiles. In one embodiment, the decoding code is configured to cause at least one processor to decode the first sub-image in a tile scan order different from the tile scan order used for decoding the second sub-image. In one embodiment, the decoding code is configured to cause at least one processor to decode the first sub-image and the second sub-image using the sub-image and tile division, the first plurality of tiles of the first sub-image are classified into at least one first tile group, the second plurality of tiles of the second sub-image are classified into at least one second tile group, and no tile of at least one first tile group is placed in at least one second tile group. In one embodiment, one of the at least one first tile groups is a non-rectangular tile group. In one embodiment, the decoding code is configured to cause at least one processor to decode according to a decoding technique, and tile group level loop filtering control at the boundary between two tile groups of at least one first tile group is possible only when each of the two tile groups is rectangular.
[0010] In one embodiment, a decoder is configured to receive an encoded video stream including a sub-bitstream, the encoded video stream including a sequence parameter set (SPS) having first information about a method of dividing sub-images of an image including a first sub-image and a second sub-image, and a decoding code is configured to cause at least one processor to divide sub-images of the image according to the first information. In one embodiment, the encoded video stream includes a picture parameter set (PPS) having second information about a method of dividing tiles of an image including at least one tile of the first sub-image and at least one tile of the second sub-image, and the decoding code is configured to cause at least one processor to divide tiles of the image according to the second information. In one embodiment, the encoded video stream includes an active parameter set (APS) signaling adaptation loop filter (ALF) coefficients of the first sub-image, and the decoding code is configured to cause at least one processor to use the APS for decoding the first sub-image of the image.
[0011] In some embodiments, a non-transitory computer-readable medium storing computer instructions is provided. The computer instructions, when executed by at least one processor, cause the at least one processor to decode a first sub-image of an image of an encoded video stream independently of a second sub-image of the image of the encoded video stream using sub-image and tile partitioning, where (i) the first sub-image includes a first rectangular region of the image and the second sub-image includes a second rectangular region of the image, the second rectangular region being different from the first rectangular region, (ii) the first sub-image and the second sub-image each include at least one tile, and (iii) the first sub-image and the second sub-image do not share a common tile.
[0012] In one embodiment, when the computer instructions are executed by at least one processor, the at least one processor is further caused to decode a second sub-image of the image independently of the first sub-image using the sub-image and tile division, at least one tile of the first sub-image being a first plurality of tiles and at least one tile of the second sub-image being a second plurality of tiles.
[0013] Further features, properties, and various advantages of the subject matter of this disclosure will become more apparent from the following detailed description and the accompanying drawings.
Brief Description of the Drawings
[0014]
Fig. 1
Fig. 2
Fig. 3
Fig. 4
Fig. 5A
Fig. 5B
Fig. 5C
Fig. 6
Fig. 7
Fig. 8
Fig. 9
Fig. 10
Best Mode for Carrying Out the Invention
[0015] FIG. 1 shows a simplified block diagram of a communication system 100 according to an embodiment of the present disclosure. The system 100 may include at least two terminals 110 and 120 interconnected via a network 150. In the case of unidirectional data transmission, the first terminal 110 may encode video data to be transmitted to the other terminal 120 via the network 150 at a local location. The second terminal 120 may receive the encoded video data of the other terminal from the network 150, decode the encoded data, and display the recovered video data. Unidirectional data transmission may be common in media supply applications and the like.
[0016] FIG. 1 shows a second pair of terminals 130 and 140 provided to support bidirectional transmission of encoded video, which may occur, for example, during a video conference. In the case of bidirectional data transmission, each terminal 130, 140 may encode video data captured at a local location to be transmitted to the other terminal via the network 150. Each terminal 130, 140 may receive the encoded video data transmitted by the other terminal, may decode the encoded data, and may display the recovered video data on a local display device.
[0017] In FIG. 1, terminals 110 to 140 may be, for example, a server, a personal computer, and a smartphone, and / or any other type of terminal. For example, the terminals (110 to 140) may be a notebook computer, a tablet computer, a media player, and / or dedicated video conferencing equipment. Network 150 represents any number of networks including a wired and / or wireless communication network that transmits the encoded video data between terminals 110 to 140. The communication network 150 may exchange data through a circuit-switched channel and / or a packet-switched channel. Representative networks include a telecommunications network, a local area network, a wide area network, and / or the Internet. For the purposes of this discussion, the architecture and topology of network 150 may be irrelevant to the operation of the present disclosure unless otherwise described below.
[0018] FIG. 2 shows the arrangement of a video encoder and decoder in a streaming environment as an application example of the disclosed subject matter. The disclosed subject matter can be used for other video usage applications including, for example, video conferencing, digital television, and storage of compressed video on digital media such as CDs, DVDs, and memory sticks.
[0019] As shown in FIG. 2, the streaming system 200 may include a capture subsystem (213) that includes a video source 201 and an encoder 203. The streaming system 200 may further include at least one streaming server 205 and / or at least one streaming client 206.
[0020] The video source 201 can generate, for example, an uncompressed video sample stream 202. The video source 201 may be, for example, a digital camera. The sample stream 202, which is shown in bold to emphasize that it has a large data volume compared to the encoded video bitstream, can be processed by an encoder 203 coupled to the camera 201. As will be described in more detail below, the encoder 203 can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter. The encoder 203 may further generate an encoded video bitstream 204. The encoded video bitstream 204, which is shown in thin lines to emphasize that it has a small data volume compared to the uncompressed video sample stream 202, can be stored in a streaming server 205 for later use. One or more streaming clients 206 can access the streaming server 205 to retrieve a video bitstream 209, which may be a copy of the encoded video bitstream 204.
[0021] The streaming client 206 can include a video decoder 210 and a display device 212. The video decoder 210 decodes, for example, a received copy of the encoded video bitstream 204, which is the video bitstream 209, to generate a transmitted video sample stream 211, and the video sample stream 211 can be displayed on the display device 212 or another display device (not shown). In some streaming systems, the video bitstreams 204, 209 can be encoded according to some video encoding / compression standards. Examples of such standards include, but are not limited to, ITU-T Recommendation H.265. A video encoding standard, informally known as Versatile Video Coding (VVC), is under development. Embodiments of the present disclosure may be used in the context of VVC.
[0022] FIG. 3 is an exemplary functional block diagram of a video decoder 210 attached to a display device 212 according to an embodiment of the present disclosure.
[0023] The video decoder 210 may include a channel 312, a receiver 310, a buffer memory 315, an entropy decoder / syntax analyzer 320, a scaler / inverse transform unit 351, an intra prediction unit 352, a motion compensation prediction unit 353, an aggregation device 355, a loop filter unit 356, a reference image memory 357, and a current image memory 358. In at least one embodiment, the video decoder 210 may include an integrated circuit, a series of integrated circuits, and / or other electronic circuits. The video decoder 210 may further be embodied as software executed by one or more CPUs with associated memory, partially or wholly.
[0024] In this embodiment, and in other embodiments, receiver 310 may receive one or more encoded video sequences decoded by decoder 210, may receive one encoded video sequence at a time, and the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequence may be received from channel 312, which may be hardware / software that couples to a storage device storing the encoded video data. Receiver 310 may receive the encoded video data along with other data such as encoded audio data and / or auxiliary data streams, which may each be transferred to an entity (not shown) that uses them. Receiver 310 may separate the encoded video sequence from other data. To counter network jitter, buffer memory 315 may be coupled between receiver 310 and entropy decoder / syntax analyzer 320 (hereinafter referred to as "syntax analyzer"). Buffer 315 may not be used or may be made smaller when receiver 310 is receiving data from a storage / transfer device with sufficient bandwidth and controllability, or from an isochronous network. Buffer 315 may be required for use in a best-effort packet network such as the Internet, may be relatively large, and may be made adaptable in size.
[0025] The video decoder 210 may include a syntax analyzer 320 to reconstruct symbol 321 from the entropy - encoded video sequence. Such classification of symbols includes, for example, information used to manage the operation of the decoder 210 and potential information for controlling display devices such as display device 212 shown in FIG. 2 that may be coupled to the decoder. The control information for the display device may be in the form of, for example, Supplementary Enhancement Information (SEI message) or a Video Usability Information (VUI) parameter set fragment (not shown). The syntax analyzer 320 may perform syntax analysis / entropy decoding on the received encoded video sequence. The code of the encoded video sequence may follow video encoding techniques or standards and may follow principles well - known to those skilled in the art, including variable - length coding, Huffman coding, context - dependent or context - independent arithmetic coding, etc. The syntax analyzer 320 may extract a group of sub - group parameters for at least one of the sub - groups of pixels within the video decoder based on at least one parameter corresponding to the group from the encoded video sequence. The sub - groups can include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. Also, the syntax analyzer 320 may extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.
[0026] The syntax analyzer 320 may perform an entropy decoding / syntax analysis operation on the video sequence received from buffer 315 to generate symbol 321.
[0027] The reconstruction of symbol 321 can include multiple different units depending on the type of the encoded video or a portion thereof (e.g., inter-picture and intra-picture, inter-block and intra-block), as well as other elements. Which units are included and how they are included can be controlled by subgroup control information parsed from the video sequence encoded by syntax analyzer 320. Such a flow of subgroup control information between syntax analyzer 320 and the multiple units described below is not shown for clarity.
[0028] In addition to the function blocks already described, decoder 210 can be conceptually subdivided into several functional units as described below. In actual implementations operating under commercial constraints, many of such units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, it is appropriate to conceptually subdivide into the following functional units.
[0029] One unit may be a scaler / inverse transform unit 351. The scaler / inverse transform unit 351 may receive quantization transform coefficients, as well as control information, which includes, as (a plurality of) symbols 321 from syntax analyzer 320, the transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit 351 can output a block containing sample values, which can be input to aggregator 355.
[0030] In some cases, the output samples of the scaler / inverse transform 351 can be related to intra-coded blocks, i.e., blocks that do not use prediction information from the previously reconstructed image can use prediction information from the previously reconstructed part of the current image. Such prediction information can be provided by the intra-image prediction unit 352. In some cases, the intra-image prediction unit 352 uses the surrounding already reconstructed information taken from the current (partially reconstructed) image in the current image memory 358 to generate a block of the same size and shape as the block being reconstructed. The aggregator 355 optionally adds the prediction information generated by the intra-prediction unit 352 to the output sample information provided by the scaler / inverse transform unit 351 on a sample-by-sample basis.
[0031] In other cases, the output samples of the scaler / inverse transform unit 351 can be related to inter-coded and potentially motion-compensated blocks. In such cases, the motion-compensation prediction unit 353 can access the reference image memory 357 to extract the samples to be used for prediction. After motion-compensating the extracted samples according to the symbols 321 related to the block, these samples can be added by the aggregator 355 to the output of the scaler / inverse transform unit 351 (in this case called the residual samples or residual signal) to generate the output sample information. The address in the reference image memory 357 from which the motion-compensation prediction unit 353 extracts the prediction samples can be controlled by the motion vector. The motion vector can be in the form of the symbols 321 and can be used by the motion-compensation prediction unit 353 and can have, for example, X, Y, and reference image components. Also, motion compensation can include interpolation of the sample values taken from the reference image memory 357 when an exact motion vector of sub-samples is used, a motion vector prediction mechanism, etc.
[0032] The output samples of the aggregation device 355 can be subjected to various loop filtering techniques of the loop filter unit 356. The video compression technology can include in-loop filter technology that is controlled by parameters included in the encoded video bitstream and can be made available to the loop filter unit 356 as symbol 321 from the syntax analyzer 320. Further, it can also respond to meta information obtained during the decoding of a previous (in decoding order) portion of the encoded image or encoded video sequence, and can similarly respond to previously reconstructed and loop-filtered sample values.
[0033] The output of the loop filter unit 356 may be a sample stream that can be output to a display device such as the display device 212 and can be stored in the reference image memory 357 for use in subsequent inter-picture prediction.
[0034] Some encoded images can be used as reference images for subsequent prediction once they are fully reconstructed. When an encoded image is fully reconstructed and the encoded image is specified as a reference image (e.g., by the syntax analyzer 320), the current reference image stored in the current image memory 358 can become part of the reference image memory 357, and a new current image memory can be reallocated before starting the reconstruction of subsequent encoded images.
[0035] The video decoder 210 may perform a decoding operation according to a predetermined video compression technique that can be described in standards such as ITU-T Rec.H.265. The encoded video sequence is used according to the syntax specified by the video compression technique or standard in the sense of complying with the syntax of the video compression technique or standard, and specifically in the profile described therein. Also, in order to comply with some video compression techniques or standards, the complexity of the encoded video sequence may be within the range defined by the level of the video compression technique or standard. In some cases, the maximum image size, maximum frame rate, maximum reconstructed sample rate (e.g., measured in megasamples / second), maximum reference image size, etc. are limited by the level. The limitations set by the level may, in some cases, be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.
[0036] In one embodiment, the receiver 310 may receive additional (redundant) data together with the encoded video. The additional data may be included as part of the (plural) encoded video sequence. The additional data may be used by the video decoder 210 to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant images, forward error correction codes, etc.
[0037] FIG. 4 is an exemplary functional block diagram of a video encoder 203 associated with a video source 201 according to an embodiment of the present disclosure.
[0038] The video encoder 203 may include, for example, an encoder that is a source encoder 430, an encoding engine 432, a (local) decoder 433, a reference image memory 434, a predictor 435, a transmitter 440, an entropy encoder 445, a controller 450, and a channel 460.
[0039] The symbolizer 203 may receive video samples from a video source 201 (not part of the symbolizer) that can capture the video to be symbolized by the symbolizer 203.
[0040] The video source 201 may provide the source video sequence to be symbolized by the symbolizer 203 in the form of a digital video sample stream that can be of any suitable bit depth (x-bit, 10-bit, 12-bit, etc.), any color space (BT.601 Y CrCb, RGB, etc.), and any suitable sampling structure (Y CrCb 4:2:0, Y CrCb 4:4:4, etc.). In a media supply system, the video source 201 may be a storage device that stores previously prepared video. In a video conferencing system, the video source 203 may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual images that convey motion when viewed in sequence. The image itself may be organized as a spatial array of pixels, and each pixel may contain one or more samples depending on the sampling structure, color space, etc. at the time of use. A person skilled in the art will be able to easily understand the relationship between pixels and samples. Hereinafter, the explanation will be centered around samples.
[0041] According to one embodiment, the encoder 203 can encode and compress the images of the source video sequence in real time or under other time constraints required by the application to obtain an encoded video sequence 443. Achieving an appropriate encoding speed may be a function of the controller 450. The controller 450 may further control other functional units as described later and may be functionally coupled to these units. For clarity, the couplings are not shown. The parameters set by the controller 450 can include rate control related parameters (such as picture skip, quantization, lambda value of rate-distortion optimization techniques), picture size, layout of groups of pictures (GOP), maximum motion vector search range, etc. Those skilled in the art can easily identify other functions of the controller 450 because they may be related to the video encoder 203 optimized for some system designs.
[0042] Some video encoders operate in what is readily recognized by those skilled in the art as an "encoding loop". Briefly, in some video compression techniques, when the compression between the symbols and the encoded video bitstream is reversible, the encoding loop may consist of an encoding section of the source encoder 430 (involved in generating symbols based on the input image to be encoded and reference images), and a (local) decoder 433 incorporated into the encoder 203. The decoder 433 reconstructs the symbols to generate sample data that the (remote) decoder also generates. The reconstructed sample stream may be input into the reference image memory 434. When the decoding of the symbol stream is bit-exact regardless of the position of the decoder (local or remote), the content of the reference image memory is also bit-exact between the local encoder and the remote encoder. In other words, the prediction section of the encoder "assumes" the same sample values for the reference image samples as those "assumed" by the decoder when using prediction during decoding. This basic principle of reference image synchronization (and the resulting drift if synchronization cannot be maintained, for example due to channel errors) is known to those skilled in the art.
[0043] The operation of the "local" decoder 433 may be substantially the same as that of the "remote" decoder 210, which has already been described in detail above in connection with FIG. 3. However, since the symbol is available and can be reversibly encoded / decoded into the symbol-encoded video sequence by the entropy coder 445 and the syntax analyzer 320, the entropy decoding section of the decoder 210, including the channel 312, the receiver 310, the buffer 315, and the syntax analyzer 320, need not be fully implemented in the local decoder 433.
[0044] As can be considered at present, all the decoder technologies existing in the decoder, except for syntax analysis / entropy decoding, will need to exist in substantially the same functional form in the corresponding encoder. For this reason, the disclosed subject matter focuses on the operation of the decoder. Since the description of the encoder technology may be the reverse of the decoder technology described comprehensively, it can be omitted. More detailed descriptions are required and will be described below only in some areas.
[0045] The source encoder 430 may perform motion-compensated predictive encoding as part of its operation, and predictively encode the input frame with respect to one or more previously encoded frames from the video sequence designated as the "reference frame". In this method, the encoding engine 432 encodes the difference between the pixel blocks of the input frame and the pixel blocks of the (multiple) reference frames that can be selected as the (multiple) prediction references for the input frame.
[0046] The local video decoder 433 may decode the encoded video data of a frame that can be specified as a reference frame based on the symbols generated by the source encoder 430. The operation of the encoding engine 432 may preferably be an irreversible process. When the encoded video data is decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence may be a copy of the source video sequence, usually with some errors. The local video decoder 433 may perform a decoding process that may be executed on the reference frame by the video decoder and may store the reconstructed reference frame in the reference image memory 434. In this way, the encoder 203 may locally store a copy of the reconstructed reference frame having the same content as the reconstructed reference frame obtained by a remote video decoder (without transmission errors).
[0047] The predictor 435 may perform a prediction search for the encoding engine 432. That is, the predictor 435 may search the reference image memory 434 for sample data (as candidate reference pixel blocks), or some metadata such as reference image motion vectors and block shapes, for the new frame to be encoded, which functions as an appropriate prediction reference for the new image. The predictor 435 may operate on samples based on block × pixel blocks to find an appropriate prediction reference. In some cases, the input image may have a prediction reference drawn from a plurality of reference images stored in the reference image memory 434 as determined by the search result obtained by the predictor 435.
[0048] The controller 450 may manage the encoding operation of the video encoder 430, including setting parameters and subgroup parameters used for encoding video data, for example.
[0049] The output of all the above-described functional units may be entropy-encoded by an entropy encoder 445. When generated by various functional units, the entropy encoder converts symbols into an encoded video sequence by reversibly compressing the symbols using techniques known to those skilled in the art, such as Huffman coding, variable-length coding, arithmetic coding, etc.
[0050] When generated by the entropy encoder 445, the transmitter 440 may buffer the encoded video sequence in preparation for transmission via the communication channel 460, which may be a hardware / software cooperation with a storage device for storing the encoded video data. The transmitter 440 may merge the encoded video data of the video encoder 430 with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).
[0051] The controller 450 may manage the operation of the encoder 203. During encoding, the controller 450 may assign several encoded image types to each of the encoded images, which may affect the encoding technique applicable to each image. For example, an image may often be assigned as an intra picture (I picture), a predicted picture (P picture), or a bi-directionally predicted picture (B picture).
[0052] An intra picture (I picture) can be encoded and decoded without using other frames in the sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, an Independent Decoder Refresh (IDR) picture. Those skilled in the art are aware of such variations of I pictures, as well as their respective uses and characteristics.
[0053] A predicted picture (P picture) can be encoded and decoded using intra prediction or inter prediction using at most one motion vector and a reference index to predict the sample values of each block.
[0054] The bidirectional predicted picture (B picture) can be encoded and decoded using intra prediction or inter prediction, using up to two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures can use more than two reference pictures and related metadata to reconstruct one block.
[0055] The source image is usually spatially subdivided into a plurality of sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be encoded block by block. The blocks may be encoded predictively with reference to other (already encoded) blocks when determined by the coding assignment applied to each image of the block. For example, blocks of an I picture may be encoded non-predictively, or predictively with reference to already encoded blocks of the same image (spatial prediction or intra prediction). Pixel blocks of a P picture may be encoded non-predictively by spatial prediction or temporal prediction with reference to one previously encoded reference image. Blocks of a B picture may be encoded non-predictively by spatial prediction or temporal prediction with reference to one or two previously encoded reference images.
[0056] The video coder 203 may perform an encoding operation according to a predetermined video encoding technique or standard such as ITU-T Rec.H.265. In that operation, the video coder 203 may perform various compression operations, including predictive encoding operations that utilize temporal and spatial redundancies in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.
[0057] In one embodiment, transmitter 440 may transmit additional data along with the encoded video. Video encoder 430 may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, and the like.
[0058] The encoders and decoders of the present disclosure may encode and decode video streams according to tile and sub-picture segmentation designs. Embodiments including methods using tile and sub-picture segmentation designs may be used individually or in combination in any order. Also, the methods, encoders, and decoders of the embodiments may each be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In an embodiment, one or more processors execute a program stored in a non-transitory computer-readable medium to perform encoding and decoding of a video stream according to the tile and sub-picture segmentation designs of one or more embodiments. Some aspects of the tile and sub-picture segmentation designs in the present disclosure are described below.
[0059] It is useful to include a motion constraint tile set (MCTS)-like function with the following characteristics in Versatile Video Coding (VVC) and some other video compression standards. (1) A sub-bitstream includes a video coding layer (VCL), a network abstraction layer (NAL), and non-VCL NAL units necessary for decoding a subset of tiles, and the sub-bitstream can be extracted from one VVC bitstream that covers all the tiles constituting the entire image. (2) The extracted sub-bitstream can be independently decoded by a VVC decoder without referring to NAL units.
[0060] Such a segmentation mechanism with the above-described features may be referred to as MCTS (Motion Constrained Tile Set) or a sub-image. To avoid confusion between the terms "tile group" and "tile set", in the present disclosure, the term "sub-image" is used to describe the above-described segmentation mechanism.
[0061] A non-limiting, exemplary structure of the tile and sub-image segmentation design in the embodiments will be described below with reference to FIGS. 5A-5C. In one embodiment, a video stream includes a plurality of images 500. Each image 500 may include one or more sub-images 510, and each sub-image 510 may include one or more tiles 530, as shown, for example, in FIGS. 5A-5B. The size and shape of the sub-images 510 and tiles 530 are not limited by FIGS. 5A-5B and can be of any shape or size. The tiles 510 of each sub-image 510 may be divided into one or more tile groups 520, as shown, for example, in FIGS. 5B-5C. The size and shape of the tile groups 520 are not limited by FIGS. 5B-5C and can be of any shape or size. In one embodiment, one or more tiles 530 may not be provided in any of the tile groups 520. In one embodiment, the tile groups 520 within the same sub-image 510 may share one or more tiles 530. Embodiments of the present disclosure may decode and encode a video stream according to the tile and sub-image segmentation design, and the sub-images 510, tile groups 520, and tiles 530 are defined and used.
[0062] Embodiments of the sub-image design of the present disclosure may include aspects of the sub-image design according to JVET-M0261.
[0063] In addition, embodiments of the sub-image design of the present disclosure enhance the usefulness of the sub-image design of the present disclosure when used in immersive media by the following features.
[0064] (1) Each sub-image 510 may have different random access periods and different inter-prediction structures. Such features can be used to give unequal random access periods to viewport-dependent 360-degree streaming. In viewport-dependent 360-degree streaming, only the concentrated areas containing dynamic content can be noticed by the viewer, while the other background areas may change slowly. By having different GOP structures and different random access periods, it is possible to assist in quickly accessing specific areas while improving the overall visual quality.
[0065] (2) The sub-images 510 may have different resampling rates from each other. Due to such features, the quality of the background areas (for example, the upper and lower parts in 360 degrees, non-dynamic objects in PCC) can be efficiently sacrificed for the overall bit efficiency.
[0066] (3) The plurality of sub-images 510 constituting the image 500 may or may not be encoded by a single encoder and decoded by a single decoder. For example, while some sub-images 510 can be independently encoded by an encoder, other sub-images 510 can be encoded by another encoder. Next, the sub-bitstreams corresponding to these sub-images 510 can be combined into one bitstream, which can be decoded by a decoder. Such features may be used, for example, in e-sports content. For example, in one embodiment, there are several players (Player 1 to Player N) participating in a game (such as a video game), and the game views and camera views of each of these players may be captured and transmitted. Depending on the player selected by each viewer, the media processor can group the relevant game views (for example, of the next player) and convert the group into a video.
[0067] In an embodiment, there may be a camera that captures an image of each player playing a game. For example, as shown in FIG. 6, the embodiment may include a camera 610 that captures a video image 611 of player 1, a camera 610 that captures a video image 612 of player 2, and the like. Also, one or more processors having a memory may capture a video image including an observer view 620 and a player view 630. The observer views 620 are, respectively, views such as a viewer view of the game, and the players of the game may not be seen while actively playing the game. For example, the viewer view may be an overall image of the video game world that is different from the perspective seen by the player, and / or may include information that assists the viewer watching the game, and the information will not be such that the player actively playing can be seen. The observer view 620 may include one or more observer views including a first observer view 621. The player views 630 may be, respectively, views seen by each player playing the game. For example, the first player view 631 may be an image of the video game seen by player 1 during the play of the game, and the second player view 632 may be an image of the video game seen by player 2 during the play of the game.
[0068] The video image of the camera 610, the observer view 620, and the competitor view 630 may be received by the compositor 640. In an embodiment, any number of such video images may be individually encoded by each encoder, and / or one or more video images may be encoded together by a single encoder. The compositor 640 may receive the video image after the video image has been encoded by one or more encoders, before the video image is encoded, or after the video image has been decoded. The compositor 640 may also function as an encoder and / or a decoder. Based on an input such as the layout & stream selection 660, the compositor 640 may provide a specific composite of two or more video images from the camera 610, the observer view 620, and the competitor view 630 as the composite video 690 shown in FIG. 7 to the transcoder 650. The transcoder 650 transcodes the composite video 690 and outputs the composite video 690 to the media sink 680 that may include one or more display devices for displaying the composite video 690. The transcoder 650 may be monitored by the network monitor 670.
[0069] As shown in FIG. 7, the transcoded composite video 690 may include, for example, a composite of the video image 611 of competitor 1, the video image 612 of competitor 2, the first competitor view 631, the second competitor view 632, and the first observer view 621. However, any combination of any number of any video images provided to the compositor 640 may be composited together as the composite video 690.
[0070] In one embodiment, the sub-image design includes the following features and technical details: (1) The sub-image 510 is a rectangular area. (2) The sub-image 510 may or may not be divided into a plurality of tiles 530. (3) When divided into a plurality of tiles 530, each sub-image 510 has its own tile scan order. (4) The tiles 530 within the sub-image 510 can be combined into rectangular or non-rectangular tile groups 520, but tiles 530 belonging to another sub-image 510 cannot be grouped together.
[0071] At least the following aspects (1)-(7) of the sub-image design in the embodiments of the present disclosure are distinguished from JVET-M0261.
[0072] (1) In one embodiment, each sub-image 510 may refer to its own SPS, PPS, and APS, provided that each SPS may include all sub-image division information.
[0073] (a) In one embodiment, the sub-image division and layout (notification) information may be signaled in the SPS. For example, the decoded sub-image and the output size of the sub-image 510 may be signaled. For example, the reference picture list (RPL) information of the sub-image 510 may be signaled. In one embodiment, each sub-image 510 may or may not have the same RPL information within the same picture. (b) In one embodiment, the tile division information of the sub-image 510 may be signaled in the PPS. (c) In one embodiment, the ALF coefficients of the sub-image 510 may be signaled in the APS. (d) In one embodiment, a plurality of sub-images 510 can refer to any parameter set or SEI message.
[0074] (2) In one embodiment, the sub-image ID may be signaled in the NAL unit header.
[0075] (3) In one embodiment, all decoding processes (e.g., in-loop filtering, motion compensation) across sub-image boundaries can be rejected.
[0076] (4) In one embodiment, the boundary of the sub-image 510 may be extended and padded for motion compensation. In one embodiment, a flag indicating whether the boundary is extended may be signaled in the SPS.
[0077] (5) In one embodiment, the decoded sub-image 510 may or may not be resampled for output. In one embodiment, the spatial ratio between the decoded sub-image size and the output sub-image size, signaled in the SPS, may be used to calculate the resampling rate.
[0078] (6) In one embodiment, for sub-bitstream extraction, the VCL NAL unit corresponding to the sub-image ID is extracted and the rest are removed. The parameter set referred to by the VCL NAL unit having the sub-image ID is extracted and the rest are removed.
[0079] (7) In one embodiment, for sub-bitstream assembly, all VCL NAL units having the same POC value may be interleaved with the same access unit (AU). The sub-image 510 splitting information in the SPS is rewritten as necessary. The sub-image ID and any parameter set ID are rewritten as necessary.
[0080] The sequence parameter set RBSP syntax of the embodiments of the present disclosure is shown in Table 1 below.
[0081]
Table 1
[0082] "num_sub_pictures_in_pic" specifies the number of sub-images 510 in each image 500 that refers to the SPS.
[0083] If "signalled_sub_pic_id_length_minus1" is 1, it specifies the number of bits used to represent the syntax element "sub_pic_id[i]" that exists, and the syntax element "tile_group_sub_pic_id[i]" within the tile group header. The value of "signalled_sub_pic_id_length_minus1" may be in the range from 0 to 15.
[0084] "dec_sub_pic_width_in_luma_samples[i]" specifies the width of the i-th decoded sub-picture 510 within the unit of luma samples in the encoded video sequence. "dec_sub_pic_width_in_luma_samples[i]" may not be 0 and may be an integer multiple of "MinCbSizeY".
[0085] "dec_sub_pic_height_in_luma_samples[i]" specifies the height of the i-th decoded sub-picture 510 within the unit of luma samples in the encoded video sequence. "dec_sub_pic_height_in_luma_samples[i]" may not be 0 and may be an integer multiple of "MinCbSizeY".
[0086] "output_sub_pic_width_in_luma_samples[i]" specifies the width of the i-th output sub-picture 510 within the unit of luma samples. "output_sub_pic_width_in_luma_samples" may not be 0.
[0087] "output_sub_pic_height_in_luma_samples[i]" specifies the height of the i-th output sub-picture 510 within the unit of luma samples. "output_sub_pic_height_in_luma_samples" may not be 0.
[0088] 「sub_pic_id[i]」specifies the sub-image identifier of the i-th sub-image 510. The length of the 「sub_pic_id[i]」 syntax element is 「sub_pic_id_length_minus1」 + 1 bit. When it does not exist, the value of 「sub_pic_id[i]」 is set to 0.
[0089] 「left_top_pos_x_in_luma_samples[i]」 specifies the column position of the first pixel of the i-th sub-image 510.
[0090] 「left_top_pos_y_in_luma_samples[i]」 specifies the row position of the first pixel of the i-th sub-image 510.
[0091] In one embodiment, each sub-image 510 is resampled to its corresponding output sub-image size after decoding. In one embodiment, a sub-image 510 cannot be overlaid with the area of another sub-image 510A. In one embodiment, the width and height of the image size, which are composed of the output sizes and positions of all sub-images, may be equal to 「pic_width_in_luma_samples」 and 「pic_height_in_luma_samples」, but a partial image area composed of a subset of sub-images can be decoded.
[0092] The syntax of the tile group header of the embodiments of the present disclosure is shown in Table 2 below.
[0093]
Table 2
[0094] 「tile_group_sub_pic_id」 specifies the sub-image identifier of the sub-image to which the current tile group belongs.
[0095] Tiles in HEVC are designed to support two main use cases: (1) parallel decoding processes and (2) partial transmission and partial decoding. The first use case is basically realized by using the original tiles in HEVC, but there are still some dependencies on inter-prediction operations. The second use case is realized by using an additional SEI message called motion constraint tile set (MCTS), although in an arbitrary way. In VVC, the same tile design is inherited from HEVC, but a new approach, so-called tile groups, is applied to support multiple use cases.
[0096] This disclosure provides embodiments including sub-image designs that can be used for viewport-based transmission in VR360 and other immersive content that supports 3 / 6 degrees of freedom. Such functionality would be useful for the widespread use of VVC in future immersive content services. Also, it is desirable to have full functionality to enable complete parallel decoding without dependencies across tile boundaries.
[0097] To provide better parallel processing capabilities, two syntax elements ("loop_filter_across_tiles_enabled_flag" and "loop_filter_across_tile_groups_enabled_flag") may be used. These syntax elements indicate that the in-loop filtering operation is not performed across tile boundaries or tile group boundaries respectively. In embodiments of the present disclosure, the above two syntax elements may be included together with two additional syntax elements. The two additional syntax elements may indicate whether the inter-prediction operation is performed or not across tile boundaries or tile group boundaries respectively. In these semantics, the inter-prediction operation includes, for example, temporal motion compensation, reference to the current picture, temporal motion vector prediction, and any parameter prediction operation between pictures. Therefore, in embodiments of the present disclosure, in an encoding / decoding standard such as HEVC, instead of a tile set with restricted motion, tiles with restricted motion may be used. This function can be compatible with sub-images or the MCTS method.
[0098] The above syntax elements are useful for at least the following two use cases. (1) A complete parallel decoding process without dependencies across tile / tile group boundaries, (2) Reconfiguration of the tile group layout without transcoding VCL NAL units.
[0099] Regarding the first use case, even when the image 500 is divided into two tile groups 520 for transmission, when the tile group 520 includes a plurality of tiles 530, a plurality of decoders can decode the tile group 520 in parallel. In this way, the complete parallel processing capability becomes useful.
[0100] The second use case relates to a use case of processing that relies on a VR360 viewport. When the viewport in question transitions across the boundary between two tile groups, in order to display the viewport in question, the decoder may need to receive and decode two tile groups. However, in embodiments of the present disclosure having the described syntax elements, the split information of tile group 520 in the PPS can be updated, for example, by a server or a cloud processor, so that only one tile group 520 includes the entire viewport in question, and one tile group 520 can be transmitted to the decoder. This instantaneous re-splitting becomes possible without changing the VLC level when all tiles are encoded as motion-limited tiles.
[0101] Examples of the syntax elements of embodiments of the present disclosure are shown below.
[0102] The image parameter set RBSP syntax of embodiments of the present disclosure is shown in Table 3 below.
[0103] [Table 3]
[0104] If "full_parallel_decoding_enabled_flag" is 1, it is specified that any processing and prediction across tile boundaries are rejected.
[0105] If "loop_filter_across_tiles_enabled_flag" is 1, it specifies that within the picture referred to by the PPS, the in-loop filtering operation is performed across tile boundaries. If "loop_filter_across_tiles_enabled_flag" is 0, it specifies that within the picture referred to by the PPS, the in-loop filtering operation is not performed across tile boundaries. The in-loop filtering operation includes, for example, non-blocking filter operation, sample adaptive offset filter operation, and adaptive loop filter operation. If it does not exist, the value of loop_filter_across_tiles_enabled_flag can be presumed to be 0.
[0106] If "loop_filter_across_tile_groups_enabled_flag" is 1, it specifies that within the picture referred to by the PPS, the in-loop filtering operation is performed across tile group boundaries. If "loop_filter_across_tile_group_enabled_flag" is 0, it specifies that within the picture referred to by the PPS, the in-loop filtering operation is not performed across tile group boundaries. The in-loop filtering operation includes, for example, non-blocking filter operation, sample adaptive offset filter operation, and adaptive loop filter operation. If it does not exist, the value of "loop_filter_across_tile_groups_enabled_flag" can be presumed to be 0.
[0107] If the "inter_prediction_across_tiles_enabled_flag" is 1, it specifies that in the picture referred to by the PPS, the inter prediction operation is performed across tile boundaries. If the "inter_prediction_across_tiles_enabled_flag" is 0, it specifies that in the picture referred to by the PPS, the inter prediction operation is not performed across tile boundaries. The inter prediction operation includes, for example, temporal motion compensation, current picture reference, temporal motion vector prediction, and any parameter prediction operation between pictures. If it does not exist, the value of the "inter_prediction_across_tiles_enabled_flag" can be presumed to be 0.
[0108] If the "inter_prediction_across_tile_groups_enabled_flag" is 1, it specifies that in the picture referred to by the PPS, the inter prediction operation is performed across tile group boundaries. If the "inter_prediction_across_tile_groups_enabled_flag" is 0, it specifies that in the picture referred to by the PPS, the inter prediction operation is not performed across tile group boundaries. The inter prediction operation includes, for example, temporal motion compensation, current picture reference, temporal motion vector prediction, and any parameter prediction operation between pictures. If it does not exist, the value of the "inter_prediction_across_tile_groups_enabled_flag" can be presumed to be 0.
[0109] In VVC, the tile design is the same as that in HEVC, but in order to support multiple embodiments, a new method, also called "tile group", which includes, but is not limited to, (1) partial transmission and decoding, (2) parallel decoding processes, and (3) matching with the tile granularity of the MTU size, is applied.
[0110] For parallel processing, two syntax elements ("loop_filter_across_tiles_enabled_flag" and "loop_filter_across_tile_groups_enabled_flag") may be used to selectively permit in-loop filtering processes that cross tile group boundaries. However, with such syntax elements, while in-loop filtering across tile group boundaries is disabled, filtering across tile boundaries is enabled, which can complicate the filtering process when the tile groups are non-rectangular. As shown in FIG. 8, an adaptive loop filter (ALF) is processed with a 7x7 rhombic filter 840 across the tile group boundaries of two tile groups 820. If the tile groups 820 are non-rectangular, the filtering process becomes chaotic. A boundary check process is essential for each pixel and each filter coefficient to identify which pixels belong to the current tile group. There is no difference between the case of temporal motion compensation and the case of reference to the current image. If boundary extension is applied at the tile group boundary for motion compensation, the padding process also becomes complicated.
[0111] According to an embodiment of the present disclosure, loop filter control (or any other arbitrary operation) at the tile group level at the boundary of a tile group is permitted only for rectangular tile groups. Therefore, confusion in the filtering process and the padding process can be successfully avoided.
[0112] Furthermore, the flag "loop_filter_across_tile_groups_enabled_flag" may not be signaled when single_tile_per_tile_group_flag is 1. Therefore, an embodiment of the present disclosure may include, for example, the following picture parameter set RBSP syntax as shown in Table 4.
[0113]
Table 4
[0114] If the "loop_filter_across_tile_groups_enabled_flag" is 1, it specifies that the in-loop filtering operation is performed across tile group boundaries within the picture referred to by the PPS. If the "loop_filter_across_tile_group_enabled_flag" is 0, it specifies that the in-loop filtering operation is not performed across tile group boundaries within the picture referred to by the PPS. The in-loop filtering operation includes, for example, non-blocking filter operation, sample adaptive offset filter operation, and adaptive loop filter operation. If it does not exist, the value of the "loop_filter_across_tile_groups_enabled_flag" may be presumed to be equal to the value of the loop_filter_across_tiles_enabled_flag.
[0115] In one embodiment, the method may be performed by at least one processor to decode a sub-bitstream of an encoded video stream. The encoded video stream may include an encoded version of a plurality of sub-pictures 510 for one or more pictures 500.
[0116] As shown in FIG. 9, the method may include at least one processor that receives a sub-bitstream (850). In one embodiment, the sub-bitstream may include information about one or more sub-images 510 of one or more images 500 and may not include information about other sub-images 510 of the one or more images 500. After receiving the sub-bitstream, the at least one processor may decode the sub-images 510 included in the sub-bitstream. For example, assuming that the sub-bitstream includes information about a first sub-image and a second sub-image of an image, the at least one processor may decode the first sub-image of the image independently of the second sub-image (860) according to the decoding technique of the present disclosure that uses tile and sub-image segmentation designs. The step of decoding the first sub-image may include decoding the tiles of the first sub-image. The at least one processor may further decode the second sub-image of the image (870) according to the decoding technique of the present disclosure that uses tile and sub-image segmentation designs. The step of decoding the second sub-image may include decoding the tiles of the second sub-image.
[0117] In an embodiment, the at least one processor may encode the sub-images 510 of the one or more images 500 according to the tile and sub-image segmentation designs of the present disclosure and transmit one or more sub-bitstreams of the video bitstream, including one or more encoded sub-images 510, to one or more decoders for decoding according to the tile and sub-image segmentation designs of the present disclosure.
[0118] In an embodiment, a decoder (e.g., video decoder 210) of the present disclosure may execute the decoding method of the present disclosure by accessing computer program code stored in a memory and operating as instructed by the computer program code. For example, in one embodiment, the computer program code may include decoding code configured to cause the decoder to decode a first sub-image of an image independently of a second sub-image according to a decoding technique using the tile and sub-image segmentation design of the present disclosure, and further cause the decoder to decode a second sub-image of the image independently of the first sub-image according to a decoding technique using the tile and sub-image segmentation design of the present disclosure.
[0119] In an embodiment, an encoder (e.g., encoder 203) of the present disclosure may execute the encoding method of the present disclosure by accessing computer program code stored in a memory and operating as instructed by the computer program code. For example, in one embodiment, the computer program code may include encoding code configured to cause the encoder to encode sub-images 510 of one or more images 500 according to the tile and sub-image segmentation design of the present disclosure. The computer program code may further include transmission code configured to transmit one or more sub-bitstreams of a video bitstream including one or more encoded sub-images 510 to one or more decoders for decoding according to the tile and sub-image segmentation design of the present disclosure.
[0120] The foregoing technology can be implemented as computer software using computer-readable instructions and is physically stored on one or more computer-readable media. For example, FIG. 10 shows a computer system 900 suitable for implementing some embodiments of the present disclosure.
[0121] Computer software can be encoded using any suitable machine code or computer language, which may follow mechanisms such as assembly, compilation, and linking to create code containing instructions that can be executed directly by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or through interpretation, microcode execution, etc.
[0122] Instructions can be executed on various types of computers or their components, such as personal computers, tablet computers, servers, smartphones, game consoles, IoT (Internet of Things) devices, etc.
[0123] The components of the computer system 900 shown in FIG. 10 are exemplary in nature and are not intended to suggest any limitation with respect to the use or functionality of computer software implementing embodiments of the present disclosure. The configuration of the components should not be construed as having dependencies or requirements with respect to any one of the components shown in the non-limiting embodiments of the computer system 900 or combinations of components.
[0124] The computer system 900 may include several human interface input devices. Such human interface input devices can respond to input by one or more users, for example, through tactile input (pressing keys, swiping, moving a data glove, etc.), voice input (voice, clapping hands, etc.), visual input (body gestures, etc.), and olfactory input (not shown). The human interface device can also be further used for capturing several media that are not necessarily directly involved in conscious input by humans, such as voice (speech, music, ambient sound, etc.), images (scanned images, photographic images obtained by a still camera, etc.), videos (two-dimensional videos, three-dimensional videos including stereoscopic videos, etc.).
[0125] The input human interface device may include one or more of a keyboard 901, a mouse 902, a trackpad 903, a touch screen 910, a data glove, a joystick 905, a microphone 906, a scanner 907, and a camera 908 (only one of each is shown).
[0126] The computer system 900 may include several human interface output devices. Such human interface output devices can stimulate the senses of one or more users, such as tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (e.g., tactile feedback by a touch screen 910, a data glove, or a joystick 905, although there may also be a tactile feedback device that does not function as an input device). For example, such devices may be audio output devices (speakers 909, headphones (not shown)), visual output devices (each with or without touch screen input capabilities, each with or without tactile feedback capabilities, some of which may be capable of outputting more than three dimensions by means such as two-dimensional video output or three-dimensional output, including screens 910 such as CRT screens, LCD screens, plasma screens, OLED screens, VR glasses (not shown), holographic displays, smoke generation devices (smoke tank (not shown)), and printers (not shown)).
[0127] The computer system 900 can further include its associated media, such as a humanly accessible storage device, and an optical medium including a CD / DVD ROM / RW 920 containing media 921 such as CD / DVDs, a USB memory (thumb-drive) 922, a removable hard drive or solid state drive 923, legacy magnetic media such as tapes and floppy disks (not shown), and devices based on specialized ROM / ASIC / PLD such as security dongles (not shown).
[0128] It should be further understood by those skilled in the art that the term "computer-readable medium" as used herein with respect to the subject matter disclosed herein does not include a transmission medium, a carrier wave, or other transient signals.
[0129] The computer system 900 can further include an interface to one or more communication networks. The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan, in-vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, mobile communication networks including GSM, 3G, 4G, 5G, LTE, etc., wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial TV, and industrial networks including in-vehicle and CANBus. Some networks generally require an external network interface adapter attached to some general-purpose data ports or peripheral buses (949) (e.g., the USB port of the computer system 900), and other networks are generally integrated into the core of the computer system 900 by attaching to the system bus as described below (e.g., an Ethernet interface to a PC computer system or a mobile communication network interface to a computer system of a smartphone). Using any such network, the computer system 900 can communicate with other entities. Such communication can be one-way communication, reception-only communication (e.g., TV broadcast), transmission-only one-way communication (e.g., a device that transmits from CANbus to CANbus), or two-way communication with other computer systems, for example, using a local or wide area digital network. Such communication can include communication with the cloud computing environment 955. As described above, various protocols and protocol stacks can be used for each such network and each network interface.
[0130] The aforementioned human interface device, human-accessible storage device, and network interface 954 can be attached to the core 940 of the computer system 900.
[0131] The core 940 can include one or more central processing units (CPUs) 941, a graphics processing unit (GPU) 942, specialized programmable processing units 943 in the form of FPGAs (Field Programmable Gate Areas), hardware accelerators 944 for several tasks, and the like. Such devices may be connected via a system bus 948 together with internal mass storage devices such as read-only memory (ROM) 945, random access memory 946, an internal hard drive, SSD, etc. 947 that are not accessible to internal users. In some computer systems, the system bus 948 can be accessed in the form of one or more physical plugs so as to be expandable by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the system bus 948 of the core or via a peripheral bus 949. The architecture of the peripheral bus includes PCI, USB, etc. The core 940 may include a graphics adapter 950.
[0132] The CPU 941, GPU 942, FPGA 943, and accelerator 944 can execute in combination several instructions that can create the aforementioned computer code. The computer code can be stored in the ROM 945 or RAM 946. Transient data can also be stored in the RAM 946, whereas persistent data can be stored, for example, in the internal mass storage device 947. By using cache memory, it becomes possible to quickly store and retrieve data in any memory device, and it can be closely associated with one or more CPUs 941, GPUs 942, mass storage device 947, ROM 945, RAM 946, etc.
[0133] A computer-readable medium can have computer code for performing various computer-implemented operations. The medium and the computer code can be specially designed and constructed for the purposes of this disclosure, or they can be of the kind well known and available to those of ordinary skill in the computer software arts.
[0134] As an example, and not for purposes of limitation, a computer system having an architecture 900, specifically a core 940, can provide functionality as a result of (a plurality of) processors (CPUs, GPUs, FPGAs, accelerators, etc.) executing software embodied on one or more tangible computer-readable media. Such computer-readable media can be media associated with some of the memories of core 940 having a non-transitory nature, such as the user-accessible mass storage devices as introduced above, as well as the core internal mass storage device 947 or ROM 945. The software implementing various embodiments of the present disclosure can be stored on such devices and executed by core 940. The computer-readable media can include one or more memory devices or chips according to specific requirements. The software can cause core 940 and specifically the processors (including CPUs, GPUs, FPGAs, etc.) therein to execute specific processes described herein, or specific parts of specific processes, including the definition of data structures stored in RAM 946 and the modification of such data structures according to the processes defined by the software. In addition to, or instead of, this, the computer system can provide functionality as a result of logic wired or embodied in a circuit (e.g., accelerator 944) and can operate to execute specific processes described herein, or specific parts of specific processes, instead of, or in conjunction with, software. References to software can include logic as necessary, and vice versa. References to computer-readable media can include circuits (such as integrated circuits (ICs)) that store software to be executed, circuits that embody the logic to be executed, or both, as necessary. The present disclosure encompasses any suitable combination of hardware and software.
[0135] Although the present disclosure has been described with respect to several non-limiting embodiments, there are changes, substitutions, and various alternative equivalents, which are included within the scope of the present disclosure. Thus, those skilled in the art will understand that, even if not explicitly illustrated or described herein, the principles of the present disclosure can be embodied, and thus many systems and methods within the principles and scope thereof can be devised.
Explanation of Signs
[0136] 100 Communication system 110 First terminal 120 Second terminal 130 Second pair of terminals 140 Second pair of terminals 150 Communication network 200 Streaming system 201 Video source 202 Uncompressed video sample stream 203 Video coder 204 Video bitstream 205 Streaming server 206 Streaming client 209 Video bitstream 210 Video decoder 211 Video sample stream 212 Display device 213 Capture subsystem 310 Receiver 312 Channel 315 Buffer memory 320 Syntax analyzer 321 Symbol 351 Scaler / inverse transform unit 352 Intra-picture prediction unit 353 Motion compensation prediction unit 355 Aggregation device 356 Loop filter unit 357 Reference picture memory 358 Current picture memory 430 Source coder 432 Symbolization Engine 433 Local Decoder 434 Reference Image Memory 435 Predictor 440 Transmitter 443 Source Video Sequence 445 Entropy Coder 450 Controller 460 Communication Channel 500 Image 510 Sub-Image 520 Tile Group 530 Tile 610 Camera 611 Video Image 612 Video Image 620 Observer View 621 First Observer View 630 Competitor View 631 First Competitor View 632 Second Competitor View 640 Compositor 650 Transcoder 660 Layout & Stream Selection 670 Network Monitor 680 Media Sink 690 Composite Video 820 Tile Group 840 Rhombus Filter 900 Computer System 901 Keyboard 902 Mouse 903 Track Pad 905 Joystick 906 Microphone 907 Scanner 908 Camera 909 Speaker 910 Touch Screen 920 Optical Medium 921 Medium 922 USB Memory 923 Solid State Drive 940 Cores 941 Central Processing Unit (CPU) 942 Graphics Processing Unit (GPU) 943 FPGA 944 Accelerator 945 Read-Only Memory (ROM) 946 Random Access Memory 947 Internal Mass Storage Device 948 System Bus 949 Peripheral Bus 950 Graphics Adapter 954 Network Interface 955 Cloud Computing Environment
Claims
Claim 1 A method executed by at least one processor for decoding a sub-bitstream of an encoded video stream, the encoded video stream including an encoded version of a first sub-image and a second sub-image of an image, the method comprising: receiving the sub-bitstream; decoding the first sub-image of the image independently of the second sub-image using sub-images and tile partitioning, wherein: (i) the first sub-image includes a first rectangular region of the image, the second sub-image includes a second rectangular region of the image, and the second rectangular region is different from the first rectangular region; (ii) the first sub-image and the second sub-image each include at least one tile; (iii) the first sub-image and the second sub-image do not share a common tile. A method comprising the above steps.
Citation Information
Patent Citations
Grouping tiles for video coding
JP2014531178A
Instructions for using wavefront parallelism in video coding.
JP2015507906A
Tile and sub-image division
JP7670885B2