Tile-based CTU parallel processing in video coding
By dividing video images into rectangular slices and raster scan slices, processing video encoding and decoding in parallel, and utilizing WPP technology, the problem of unbalanced resource allocation in existing technologies is solved, the encoding and decoding efficiency and flexibility are improved, and various local motion and texture characteristics can be adapted.
Patent Information
- Application Number
- CN202480013158.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-09-25
- Filing Date
- 2024-02-18
- Publication Date
- 2025-09-26
AI Technical Summary
Existing video codec technologies suffer from inefficiency and unbalanced resource allocation when processing slices in parallel, especially when processing slices of different shapes and arrangements, resulting in resource waste and increased latency during the encoding and decoding process.
A slice-based parallel processing method is adopted to divide video images into rectangular slices and raster scan slices, and the boundaries and parameters of the slices are clarified through syntax elements and syntax tables. The wavefront parallel processing (WPP) technology is used to achieve parallel encoding and decoding of different slices, optimizing resource allocation and processing delay.
It improves the efficiency and resource utilization of the video encoding and decoding process, reduces processing delays, enhances the flexibility and adaptability of the encoding and decoding process, and adapts to various local motion and texture characteristics.
Smart Images

Figure CN120712781A_ABST
Abstract
Description
[0001]
Cross-reference
[0002] This disclosure is a part of a non-provisional application that claims the benefit of priority to U.S. Provisional Patent Applications Nos. 63 / 485,558, 63 / 584,512, and 63 / 584,921, filed February 17, 2023, September 22, 2023, and September 25, 2023, respectively. The contents of the above applications are incorporated herein by reference.
Technical field
[0003] The present disclosure relates to a method for encoding and decoding video pictures by processing slices in parallel. [Background Technology]
[0004] Unless otherwise indicated herein, the approaches described in this section are not prior art to the claims listed below and are not admitted to be prior art by inclusion in this section.
[0005] High-Efficiency Video Coding (HEVC) is an international video coding standard developed by the Joint Collaborative Team on Video Coding (JCT-VC). HEVC is based on a hybrid block-based motion compensated DCT-like transform codec architecture. The basic unit used for compression is called a coding unit (CU), which is a 2Nx2N pixel block. Each CU can be recursively divided into four smaller CUs until a predefined minimum size is reached. Each CU contains one or more prediction units (PUs). The encoded video data is organized into network abstraction layer (NAL) units. Each NAL unit is a packet containing an integer number of bytes for transmission or storage.
[0006] Versatile Video Coding (VVC) is the latest international video codec standard developed by the Joint Video Experts Team (JVET) of ITU-T SG16WP3 and ISO / IEC JTC1 / SC29 / WG11. The input video signal is predicted from a reconstructed signal derived from coded picture regions. The prediction residual signal is processed through a block transform. The transform coefficients are quantized and entropy coded in the bitstream along with other additional information. The reconstructed signal is generated from the prediction signal and the reconstructed residual signal, which is obtained by inverse transforming the dequantized transform coefficients. The reconstructed signal is further processed through loop filtering to remove coding and decoding artifacts. The decoded picture is stored in a frame buffer and used to predict future pictures in the input video signal.
[0007] In VVC, coded pictures are divided into non-overlapping square areas represented by related coding tree units (CTUs). The leaf nodes of the codec tree correspond to codec units (CUs). A coded picture can be composed of multiple slices, each of which includes an integer number of CTUs. The coded video data of a slice can be completely contained in a NAL unit for transmission and decoding. The size of the NAL unit can be defined according to the size of the slice.
[0008] The CTUs in a slice are processed in raster scan order. Bi-predictive (B) slices can be decoded using intra or inter prediction, using up to two motion vectors and reference indices to predict the sample values of each block. Predictive (P) slices are decoded using intra or inter prediction, using up to one motion vector and reference indices to predict the sample values of each block. Intra (I) slices are decoded using only intra prediction.
[0009] A CTU can be divided into one or more non-overlapping codec units (CUs) using a quadtree (QT) and nested multi-type-tree (MTT) structure to accommodate various local motion and texture characteristics. A CU can be further split into smaller CUs using one of five partitioning types: quadtree, vertical binary tree, horizontal binary tree, vertical center-side ternary tree, and horizontal center-side ternary tree.
[0010] Each CU contains one or more prediction units (PUs). A prediction unit, along with the associated CU syntax, serves as the basic unit for signaling predictor information. A specified prediction process is used to predict the values of the associated pixel samples within the PU. Each CU may contain one or more transform units (TUs) to represent prediction residual blocks. A transform unit (TU) consists of a transform block (TB) of luma samples and two corresponding transform blocks of chroma samples, with each TB corresponding to a residual block of samples for a color component. Integer transforms are applied to the transform blocks. The level values of the quantized coefficients are entropy coded in the bitstream along with other additional information. The terms coding tree block (CTB), coding block (CB), prediction block (PB), and transform block (TB) are defined to refer to a two-dimensional array of samples that specifies a single color component associated with a CTU, CU, PU, and TU, respectively. Therefore, a CTU consists of a luma CTB, two chroma CTBs, and associated syntax elements. A similar relationship exists between CUs, PUs, and TUs.
[0011] For each inter-predicted CU, the motion parameters include motion vector, reference picture index and reference picture list usage index, as well as additional information for generating inter-predicted samples. Motion parameters can be signaled explicitly or implicitly. When a CU is coded in skip mode, the CU is associated with one PU and has no significant residual coefficients, no coded motion vector increments or reference picture indices. A merge mode is specified, in which the motion parameters of the current CU are obtained from neighboring CUs, including spatial and temporal candidates, as well as additional schemes introduced in VVC. Merge mode can be applied to any inter-predicted CU. An alternative to merge mode is to explicitly transmit motion parameters, where the motion vector, the corresponding reference picture index and reference picture list usage flag for each reference picture list, and other required information are explicitly signaled in each CU. [Summary of the invention]
[0012] The following summary is provided for illustrative purposes only and is not intended to be limiting in any respect. That is, the following summary is intended to introduce the concepts, highlights, benefits, and advantages of the novel and readily apparent technologies described herein. Selected embodiments are further described in the detailed description that follows. Therefore, the following summary is not intended to identify essential features of the claimed subject matter, nor is it intended to be used to determine the scope of the claimed subject matter.
[0013] Some embodiments of the present disclosure provide a method for encoding and decoding a video picture as a slice. A video compiler receives data to be encoded or decoded as a current picture. The current picture is divided into one or more slices, each of which is divided into coding tree units (CTUs). The video compiler signals or receives parameters of a slice of the current picture. When the slice is a rectangular slice, the parameters of the slice indicate the upper left CTU of the slice, the width of the slice, and the height of the slice. When the slice is a raster scan slice, the parameters of the slice indicate the starting CTU and the ending CTU of the slice.
[0014] In some embodiments, the video compiler may signal as a syntax element in the bitstream or receive a flag to indicate whether the slice is a rectangular slice or a raster scan slice. The video compiler may process CTUs in two or more different slices of the current picture in parallel. The video compiler may perform wavefront parallel processing (WPP) by processing different groups (e.g., rows) of CTUs of a slice in parallel with different delays. In some embodiments, the video compiler may derive the right and left boundaries of the slice based on the slice parameters. In some embodiments, when the slice is a raster scan slice, the right boundary of the slice is derived based on the width of the current picture, particularly when the raster scan slice is rectangular in shape. The current picture may also include adjacent but not aligned rectangular slices.
[0015] The CTUs of the slice are encoded and packaged in a network abstraction layer (NAL) unit for transmission or storage. In some embodiments, the size of the NAL unit is defined based on the size of the slice and does not include other slices. The NAL unit transmits the data of the slice without including other slices, and different slices of the current picture are transmitted by different NAL units.
Brief Description of the Drawings
[0016] The accompanying drawings are included to provide a further understanding of the present disclosure and constitute a part of this disclosure. The drawings illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure. It should be understood that the drawings are not necessarily to scale, as some components may be shown out of proportion to their actual dimensions in order to clearly illustrate the concepts of the present disclosure.
[0017] Figure 1A -B describes raster scan slices and rectangular slices.
[0018] Figure 2 Parallel processing threads are shown operating to encode or decode multiple slices within a video picture.
[0019] Figure 3The application of wavefront parallel processing (WPP) within a slice of a video picture is shown.
[0020] Figure 4 illustrates a video picture divided into rectangular slices.
[0021] Figure 5 illustrates a video picture divided into raster scan slices.
[0022] Figure 6 Describes an example video encoder that can implement raster scan and rectangular slices.
[0023] Figure 7 The portion of a video encoder that implements parallel processing of rectangular and raster scan slices is described.
[0024] Figure 8 Provides a conceptual illustration of the process of encoding video pictures into rectangular or raster scan slices.
[0025] Figure 9 Describes an example video decoder that implements raster scanning and rectangular tiles.
[0026] Figure 10 The portion of a video decoder that implements parallel processing of rectangular and raster scan slices is described.
[0027] Figure 11 Conceptually illustrates the process of decoding video pictures into rectangular or raster scan slices.
[0028] Figure 12 An electronic system that implements certain embodiments of the present disclosure is conceptually illustrated. [Specific implementation method]
[0029] In the following detailed description, numerous specific details are provided by way of example to provide a thorough understanding of the relevant technology. Any variations, derivatives, and / or extensions of the technology described herein are intended to be within the scope of this disclosure. In some cases, well-known methods, procedures, components, and / or circuits related to one or more example implementations disclosed herein may be described at a relatively high level to avoid unnecessarily obscuring aspects of the disclosed technology.
[0030] Slice partitions may have slice boundaries that impose constraints on codec tools during encoding and decoding, such as edge checking and neighbor block availability checking, QP setting, context-adaptive binary arithmetic coding (CABAC) initialization, and loop filtering. These constraints may be used by various applications and / or for flexibility control, including enabling parallel processing structures.
[0031] I. Raster Scan and Rectangular Slices
[0032] For some embodiments, the CTU grid is the basic partitioning of a picture, and the CTUs of a slice are specifically contained in a separate network abstraction layer (NAL) unit. In some embodiments, a slice can be specified as an integer number of CTUs arranged consecutively in a raster scan of the picture. Such a slice is called a raster scan slice. In some embodiments, a slice can be specified as an integer number of consecutive complete CTU rows within a rectangular area of the picture. Such a slice is called a rectangular slice. In some embodiments, a video picture can be partitioned into raster scan slices and / or rectangular slices.
[0033] Figure 1A -B describes raster scan slices and rectangular slices. Figure 1A A picture 100 is illustrated as being divided into two (2) raster slices 111 and 112. The raster slices may or may not be rectangular in shape. Figure 1B The same picture 100 is illustrated as being partitioned into four (4) rectangular slices 121 - 124. During encoding and decoding, the CTUs contained in a slice (rectangular or raster scan) are arranged in raster scan order.
[0034] II. Parallel Processing
[0035] CTU-based block data video encoding and decoding may be constrained and controlled by slice boundaries. Multiple different partitions defined by slice boundaries (i.e., slices) within a picture can be processed by different parallel processing threads. Figure 2 and Figure 3 Conceptual illustration of parallel processing threads (arrow lines) encoding and decoding different parts of a video picture.
[0036] Figure 2 An example is shown where four parallel processing threads (arrowed lines labeled Threads 1-4) are operating to encode or decode four slices 211-214 within a video picture 200. In this example, the four slices 211-214 are rectangular slices.
[0037] Figure 3An example is shown in which wavefront parallel processing (WPP) is applied within a slice 310 of a video picture 300. (Slice 310 may be one of multiple rectangular slices of the video picture 300. Slice 310 may also be the only slice of the video picture 300.) As shown, when parallel processing is performed, slice-specific processing threads 1-6 (arrowed lines) are applied to different CTU rows 321-326 of the slice 310. Processing threads 1-6 are assigned to different CTU rows. Parallel processing is performed in a WPP manner (to accommodate logical dependencies of reconstructed CTUs) such that threads 2-6 of different CTU rows 322-326 start at different CTU latencies relative to processing thread 1 of the first (topmost) CTU row 321.
[0038] In some embodiments, a slice of a video picture may be a rectangular slice or a raster scan slice.In some embodiments, the syntax element pps_rect_slice_flag is used to indicate whether a slice is a rectangular slice or a raster scan slice. Figure 4 A video picture 400 is shown partitioned into rectangular slices 411-415 (labeled as slices 1 to 5). For rectangular slices (pps_rect_slice_flag = 1), the left and right boundaries are derived in CTB units by the top left CTU and the last CTU position of the slice. For example, the figure shows the width and height of rectangular slice 414 and its top left CTU.
[0039] The data for rectangular slices 411-415 are conveyed separately and exclusively by NAL units 421-425. For example, NAL unit 423 conveys only the data for slice 413 and no other slices, NAL unit 424 conveys only the data for slice 414 and no other slices, and so on.
[0040] In some embodiments, adjacent rectangular pieces may not be aligned. Figure 4 In the example shown, slices 411 and 414 are adjacent (vertically) but not aligned (horizontally). Similarly, slices 412 and 414 are adjacent but not aligned. In other words, at least some of the left or right vertical boundaries of rectangular slices 411, 412, or 414 may not span the entire height of the image. (Although not shown, in some embodiments, the top or bottom horizontal boundaries of the rectangular slices may not span the entire width of the image.)
[0041] Figure 5A video picture 500 is shown divided into raster scan slices 511-513 (slices 1 to 3). For raster scan slices (pps_rect_slice_flag = 0), the left and right boundaries are derived in CTB units based on the position of the first CTU of the slice and the picture width. For example, the figure shows the first and last CTUs of raster scan slice 512 and their left and right boundaries. Although not shown in the figure, raster scan slices may have a rectangular shape.
[0042] The data of raster scan slices 511-513 are respectively and exclusively conveyed by NAL units 521-523. For example, NAL unit 521 conveys only the data of slice 511 and no other slices, NAL unit 522 conveys only the data of slice 512 and no other slices, and so on.
[0043] III. Syntactic Elements of a Slice
[0044] In some embodiments, video pictures are not partitioned based on tiles, so the tile partitioning in the picture parameter set (PPS) or in the slice header (SH) is removed. Tile-based processing, such as neighboring block availability derivation process, quantization parameter derivation process, and deblocking filter process, is replaced with slice-based processing. In some embodiments, CTU-based slice partitioning of the picture is specified and signaled, for example, for rectangular slices, the upper left CTU position and width / height are specified in the PPS or SPS in units of CTUs, and / or for raster scan slices, syntax elements are specified in a video codec standard such as VVC. For example, the syntax element sh_num_ctus_in_slice_minus1 replaces sh_num_ctus_in_tiles_minus1 in SH. Syntax elements such as checking for left and right slice boundaries may be modified. The derivation of variables such as NumEntryPoints and NumCtusInCurrSlice may also be modified. Here are some examples of modified syntax elements to replace tiles based on slice partitioning of CTUs:
[0045] sps_entropy_codec_sync_enabled_flag equal to 1 specifies that a specific context variable synchronization process is called before decoding a CTU, which includes the first CTB in a row of CTBs in each slice of the SPS in each picture, and a specific context variable storage process is called after decoding a CTU, which includes the first CTB in a row of CTBs in each slice of the SPS in each picture. sps_entropy_codec_sync_enabled_flag equal to 0 specifies that a specific context variable synchronization process does not need to be called before decoding a CTU, which includes the first CTB in a row of CTBs in each slice of the SPS in each picture, and a specific context variable storage process does not need to be called after decoding a CTU, which includes the first CTB in a row of CTBs in each slice of the SPS in each picture. When sps_entropy_codec_sync_enabled_flag is equal to 1, wavefront parallel processing (WPP) is enabled.
[0046] sps_entry_point_offsets_present_flag equal to 1 specifies that signaling of entry point offsets for slice-specific CTU rows may be present in the slice header of the picture referencing the SPS. sps_entry_point_offsets_present_flag equal to 0 specifies that signaling of entry point offsets for slice-specific CTU rows is not present in the slice header of the picture referencing the SPS.
[0047] sh_num_ctus_in_slice_minus1 plus 1, when present, specifies the number of CTUs in the slice. The value of sh_num_ctus_in_slice_minus1 shall be in the range of 0 to PicSizeInCtbs-1, inclusive. When not present, the value of sh_num_ctus_in_slice_minus1 shall be inferred to be equal to 0.
[0048] The variable NumCtusInCurrSlice specifies the number of CTUs in the current slice. The list CtbAddrInCurrSlice[i], for i from 0 to NumCtusInCurrSlice-1, inclusive, specifies the picture raster scan address of the i-th CTB in the slice. The variable NumCtusInCurrSlice and the list CtbAddrInCurrSlice[i] are derived according to the following syntax (modified from the existing VVC syntax):
[0049]
[0050]
[0051] sh_slice_header_extension_data_byte[i] can have any value. Its presence and value do not affect the decoding process. The variable NumEntryPoints specifies the number of entry points in the current slice and can be derived as follows:
[0052]
[0053] sh_entry_offset_len_minus1 plus 1 specifies the length of the sh_entry_point_offset_minus1[i] syntax element in bits. The value of sh_entry_offset_len_minus1 shall be in the range of 0 to 31, inclusive.
[0054] sh_entry_point_offset_minus1[i] plus 1 specifies the offset (in bytes) of the i-th entry point and is represented by sh_entry_offset_len_minus1 plus 1 bit. The slice data immediately following the slice header consists of NumEntryPoints+1 subsets, with subset index values ranging from 0 to NumEntryPoints (inclusive). The first byte of the slice data is considered to be byte 0. When present, the emulation prevention bytes appearing in the slice data portion of the coded slice NAL unit are considered to be part of the slice data for the purpose of subset identification. Subset 0 consists of bytes 0 to sh_entry_point_offset_minus1[0] (inclusive) of the coded slice data, and subset k, where k is in the range 1 to NumEntryPoints-1 (inclusive), consists of bytes firstByte[k] to lastByte[k] (inclusive) of the coded slice data, with firstByte[k] and lastByte[k] being derived as follows:
[0055]
[0056] lastByte[k]=firstByte[k]+sh_entry_point_offset_minus1[k]
[0057] The last subset (with subset index equal to NumEntryPoints) consists of the remaining bytes of the coded slice data. When sps_entropy_codec_sync_enabled_flag is equal to 0, the value of NumEntryPoints shall be equal to 0. The subset shall consist of all coded bits from all CTUs in the slice. When sps_entropy_codec_sync_enabled_flag is equal to 1, each subset k, where k is in the range of 0 to NumEntryPoints (inclusive), shall consist of all coded bits from all CTUs in a CTU row in the slice, and the number of subsets (i.e., the value of NumEntryPoints+1) shall be equal to the total number of CTU rows in the particular slice.
[0058] For certain embodiments, the following is a syntax table for slice data:
[0059]
[0060]
[0061] The lists CtbToSliceLeftBd[CtbAddrX] and CtbToSliceRightBd[CtbAddrX] specify the conversion from the horizontal CTB address to the left and right slice boundaries in units of CTBs. The index ctbAddrX ranges from 0 to PicWidthInCtbsY (inclusive), where PicWidthInCtbsY is the width of the video picture specified in units of CTBs.
[0062] In some embodiments, for rectangular slices, the left and right borders are derived from the specified upper left CTU position and the width / height of the slice in units of CTUs; for raster scan slices, the left and right borders are derived from the left and right borders of the picture and CtbToSliceLeftBd[0] == (CtbAddrInCurrSlice[0] % PicWidthInCtbsY). For raster scan slices with a rectangular shape, CtbToSliceLeftBd[0] == 0, indicating the left border of the picture. The lists CtbToSliceLeftBd[CtbAddrX, CtbAddrY] and CtbToSliceRightBd[CtbAddrX] can be used for parallel processing, in particular for wavefront parallel processing (WPP) during decoding.
[0063] The lists CtbToSliceLeftBd[ctbAddrX] and CtbToSliceRightBd[ctbAddrX] specify the conversion from the horizontal CTB address of the slice to the left and right slice boundaries in units of CTBs. For some embodiments, the lists CtbToSliceLeftBd[CtbAddrX, CtbAddrY] and CtbToSliceRightBd[CtbAddrX] are derived as follows:
[0064] for(i=0;i <NumCtusInCurrSlice;i++){
[0065] ctbAddrX0=(CtbAddrInCurrSlice[0]%PicWidthInCtbsY)
[0066] ctbAddrX=(CtbAddrInCurrSlice[i]%PicWidthInCtbsY)
[0067] if(pps_rect_slice_flag=1){ / / rectangular slice
[0068] ctbAddrXLz=(CtbAddrInCurrSlice[NumCtusInCurrSlice-1]%PicWidthInCtbsY)
[0069] CtbToSliceLeftBd[ctbAddrX]=ctbAddrX0
[0070] CtbToSliceRightBd[ctbAddrX]=ctbAddrXz
[0071] }else{ / / (pps_rect_slice_flag=0), raster scan slice
[0072] ctbAddrY0=(CtbAddrInCurrSlice[0] / PicWidthInCtbsY)
[0073] ctbAddrY=(CtbAddrInCurrSlice[i] / PicWidthInCtbsY)
[0074] CtbToSliceLeftBd[ctbAddrX]=(ctbAddrY==ctbAddrY0)? ctbAddrX0:0
[0075] CtbToSliceRightBd[ctbAddrX]=PicWidthInCtbsY
[0076] }
[0077] }
[0078] In some embodiments, when deriving the lists CtbToSliceLeftBd[CtbAddrX] and CtbToSliceRightBd[CtbAddrX] (containing raster scan slices with a rectangular shape), the vertical CTB address is not used for the derivation of the left boundary. In some embodiments, for rectangular slices (pps_rect_slice_flag=1), the left and right boundaries are derived in units of CTBs from the top left CTU and the last CTU positions of the slice. For raster scan slices with a rectangular shape (pps_rect_slice_flag=0), the left and right boundaries are derived in units of CTBs from the picture width. In some embodiments, the lists CtbToSliceLeftBd[ctbAddrX] and CtbToSliceRightBd[ctbAddrX] specify the conversion from the horizontal CTB address of the slice to the left and right boundaries in units of CTBs. Such CtbToSliceLeftBd and CtbToSliceRightBd are derived as follows:
[0079] for(ctbAddrX=0;ctbAddrX <PicWidthInCtbsY;ctbAddrX++){
[0080] if (pps_rect_slice_flag = 1) {
[0081] ctbAddrX0=(CtbAddrInCurrSlice[0]%PicWidthInCtbsY)
[0082] ctbAddrXLz=(CtbAddrInCurrSlice[NumCtusInCurrSlice-1]%PicWidthInCtbsY)
[0083] CtbToSliceLeftBd[ctbAddrX]=ctbAddrX0
[0084] CtbToSliceRightBd[ctbAddrX]=ctbAddrXz
[0085] }else{ / / (pps_rect_slice_flag=0)
[0086] CtbToSliceLeftBd[ctbAddrX]=0
[0087] CtbToSliceRightBd[ctbAddrX]=PicWidthInCtbsY
[0088] }
[0089] }
[0090] IV. Example Video Encoder
[0091] Figure 6 An example video encoder 600 that can implement raster scanning and rectangular slices is shown. As shown, the video encoder 600 receives an input video signal from a video source 605 and encodes the signal into a bitstream 695. The video encoder 600 has multiple components or modules for encoding the signal from the video source 605, including at least some components selected from a transform module 610, a quantization module 611, an inverse quantization module 614, an inverse transform module 615, an intra-frame estimation module 620, an intra-frame prediction module 625, a motion compensation module 630, a motion estimation module 635, a loop filter 645, a reconstructed picture buffer 650, an MV buffer 665, an MV prediction module 675, and an entropy encoder 690. The motion compensation module 630 and the motion estimation module 635 are part of the inter-frame prediction module 640.
[0092] In some embodiments, modules 610-690 are software instruction modules executed by one or more processing units (e.g., processors) of a computing device or electronic device. In some embodiments, modules 610-690 are hardware circuit modules implemented by one or more integrated circuits (ICs) of an electronic device. Although modules 610-690 are shown as separate modules, some of them can be combined into a single module.
[0093] A video source 605 provides an uncompressed raw video signal representing pixel data for each video frame. A subtractor 608 calculates the difference between the raw video pixel data from the video source 605 and the predicted pixel data 613 from the motion compensation module 630 or the intra-frame prediction module 625 as a prediction residual 609. A transform module 610 converts the difference (or residual pixel data or residual signal 608) into transform coefficients (e.g., by performing a discrete cosine transform, or DCT). A quantization module 611 quantizes the transform coefficients into quantized data (or quantized coefficients) 612, which are encoded into a bitstream 695 by an entropy encoder 690.
[0094] The inverse quantization module 614 inversely quantizes the quantized data (or quantized coefficients) 612 to obtain transform coefficients, and the inverse transform module 615 inversely transforms the transform coefficients to generate a reconstructed residual 619. The reconstructed residual 619 is added to the predicted pixel data 613 to generate reconstructed pixel data 617. In some embodiments, the reconstructed pixel data 617 is temporarily stored in a line buffer (not shown) for intra-frame prediction and spatial MV prediction. The reconstructed pixels are filtered by a loop filter 645 and stored in a reconstructed picture buffer 650. In some embodiments, the reconstructed picture buffer 650 is a memory external to the video encoder 600. In some embodiments, the reconstructed picture buffer 650 is a memory internal to the video encoder 600.
[0095] The intra estimation module 620 performs intra prediction based on the reconstructed pixel data 617 to generate intra prediction data. The intra prediction data is provided to the entropy encoder 690 for encoding into a bitstream 695. The intra prediction data is also used by the intra prediction module 625 to generate predicted pixel data.
[0096] The motion estimation module 635 performs inter-frame prediction by generating motion vectors (MVs) on reference pixel data of previously decoded frames stored in the reconstructed picture buffer 650. These MVs are provided to the motion compensation module 630 to generate predicted pixel data.
[0097] The video encoder 600 generates predicted MVs using MV prediction instead of encoding the complete actual MVs in the bitstream, and encodes the difference between the MVs used for motion compensation and the predicted MVs as residual motion data and stores it in the bitstream 695.
[0098] The MV prediction module 675 generates predicted MVs based on reference MVs generated for encoding of the previous video frame, i.e., motion compensated MVs used to perform motion compensation. The MV prediction module 675 retrieves the reference MVs of the previous video frame from the MV buffer 665. The video encoder 600 stores the MVs generated for the current video frame in the MV buffer 665 as reference MVs for generating the predicted MVs.
[0099] The MV prediction module 675 creates predicted MVs using reference MVs. The predicted MVs can be calculated using spatial MV prediction or temporal MV prediction. The difference (residual motion data) between the predicted MVs and the motion-compensated MVs (MC MVs) for the current frame is encoded into the bitstream 695 by the entropy encoder 690.
[0100] The entropy encoder 690 encodes various parameters and data into a bitstream 695 using entropy coding techniques, such as context-adaptive binary arithmetic coding (CABAC) or Huffman coding. The entropy encoder 690 encodes various header elements, flags, quantized transform coefficients 612, and residual motion data as syntax elements into the bitstream 695. The bitstream 695 is then stored in a storage device or transmitted to a decoder via a communication medium, such as a network.
[0101] The loop filter 645 performs filtering or smoothing operations on the reconstructed pixel data 617 to reduce encoding and decoding artifacts, particularly at pixel block boundaries. In some embodiments, the filtering or smoothing operations performed by the loop filter 645 include a deblocking filter (DBF), sample adaptive offset (SAO), and / or an adaptive loop filter (ALF).
[0102] Figure 7 A portion of a video encoder 600 that implements parallel processing of rectangular and raster scan slices is shown. The figure conceptually illustrates a pixel processing unit 750, which may include processing units (e.g., modules 610, 611, 614, 615, 645, 650, 640, 625, 620) that implement the prediction, transform, quantization, and filtering operations of the encoding loop. Pixel processing unit 750 is capable of parallel processing within a plurality of computational threads 751-759 (labeled 1 through N). Threads 751-759 perform encoding operations on pixel data provided by video source 605, converting it into encoded data for transmission. Encoding operations may also be based on pixel data provided by reconstructed picture buffer 650.
[0103] The entropy encoder 690 may signal slice parameters to indicate whether the current picture is divided into rectangular or raster scan slices. The slice parameters may also include parameters that define each slice (e.g., starting and ending CTUs, left and right boundaries, height and width, etc.). The slice parameters are inserted into the bitstream 695 as syntax elements and provided to the pixel processing unit 750. Based on these slice parameters, the computing threads 751-759 are assigned different sets of CTUs of the current picture. CTUs of different slices can be assigned to different threads 751-759. In some embodiments, when WPP is enabled, different CTU rows of the same slice or different slices can be assigned to different threads 751-759.
[0104] The data generated by threads 751-759 is provided to the entropy encoder 690 for transmission or storage. The data for each slice is packaged into a NAL unit for transmission, regardless of whether the slice is rectangular or raster scan. When using WPP, CTUs from different rows of the same slice can be collected and transmitted as a single NAL unit.
[0105] Figure 8 The process 800 of encoding a video picture into rectangular or raster scan slices is conceptually illustrated. In some embodiments, one or more processing units (e.g., processors) of a computing device implementing the encoder 600 perform the process 800 by executing instructions stored in a computer-readable medium. In some embodiments, the electronic device implementing the encoder 600 performs the process 800.
[0106] The encoder receives (at block 805) data to be encoded as a current picture having one or more slices. Each slice includes one or more data blocks (eg, CTUs).
[0107] The encoder signals (at block 810) parameters for a slice of the current picture. The encoder determines (at block 815) whether the slice is a rectangular slice or a raster scan slice. In some embodiments, the encoder may signal a flag as a syntax element (e.g., pps_rect_slice_flag) in the bitstream to indicate whether the slice is a rectangular slice or a raster scan slice. If the slice is (to be encoded as) a rectangular slice, the encoder signals (at block 820) parameters to indicate the top left CTU of the slice, the width of the slice, and the height of the slice. If the slice is (to be encoded as) a raster scan slice, the encoder signals (at block 825) parameters to indicate the starting CTU of the slice and the ending CTU of the slice (but not including the width or height of the slice).
[0108] The encoder then encodes the CTUs of the slice based on the received data (at block 830). The encoder may process CTUs in two or more different slices of the current picture in parallel (simultaneously). The encoder may perform wavefront parallel processing (WPP) by processing different groups (e.g., rows) of CTUs of a slice in parallel with different delays. In some embodiments, the encoder may derive the right and left boundaries of the slice based on the slice parameters. In some embodiments, when the slice is a raster scan slice, the right boundary of the slice is derived based on the width of the current picture, particularly when the raster scan slice is rectangular in shape. The current picture may also include adjacent but not aligned rectangular slices (e.g., slices 411 and 414).
[0109] The encoder packages the encoded blocks into a network abstraction layer (NAL) unit for storage or transmission (at block 840). In some embodiments, the size of the NAL unit is defined based on the size of the slice and does not include other slices. The NAL unit transmits the data of the slice without including other slices, and different slices of the current picture are transmitted by different NAL units.
[0110] V. Example Video Decoder
[0111] In some embodiments, an encoder may signal (or generate) one or more syntax elements in a bitstream so that a decoder may parse the one or more syntax elements from the bitstream.
[0112] Figure 9 An example video decoder 900 that implements raster scanning and rectangular slices is shown. As shown, video decoder 900 is an image decoding or video decoding circuit that receives a bitstream 995 and decodes the contents of the bitstream into pixel data for displaying a video frame. Video decoder 900 has multiple components or modules for decoding bitstream 995, including a selection of components from an inverse quantization module 911, an inverse transform module 910, an intra-frame prediction module 925, a motion compensation module 930, a loop filter 945, a decoded picture buffer 950, an MV buffer 965, an MV prediction module 975, and a parser 990. Motion compensation module 930 is part of inter-frame prediction module 940.
[0113] In some embodiments, modules 910-990 are software instruction modules executed by one or more processing units (e.g., processors) of a computing device. In some embodiments, modules 910-990 are hardware circuit modules implemented by one or more integrated circuits (ICs) of an electronic device. Although modules 910-990 are shown as separate modules, some of them may be combined into a single module.
[0114] The parser 990 (or entropy decoder) receives the bitstream 995 and performs initial parsing according to the syntax defined by the video codec or image codec standard. The parsed syntax elements include various header elements, flags, and quantization data (or quantization coefficients) 912. The parser 990 parses the various syntax elements using entropy coding techniques such as context-adaptive binary arithmetic coding (CABAC) or Huffman coding.
[0115] The inverse quantization module 911 inversely quantizes the quantized data (or quantized coefficients) 912 to obtain transform coefficients, and the inverse transform module 910 inversely transforms the transform coefficients 916 to generate a reconstructed residual signal 919. The reconstructed residual signal 919 is added to the predicted pixel data 913 from the intra-frame prediction module 925 or the motion compensation module 930 to generate decoded pixel data 917. The decoded pixel data is filtered by the loop filter 945 and stored in the decoded picture buffer 950. In some embodiments, the decoded picture buffer 950 is a memory external to the video decoder 900. In some embodiments, the decoded picture buffer 950 is a memory internal to the video decoder 900.
[0116] The intra prediction module 925 receives intra prediction data from the bitstream 995 and, based on the data, generates predicted pixel data 913 from decoded pixel data 917 stored in the decoded picture buffer 950. In some embodiments, the decoded pixel data 917 is also stored in a line buffer (not shown) for intra picture prediction and spatial MV prediction.
[0117] In some embodiments, the contents of the decoded picture buffer 950 are used for display. The display device 905 can directly retrieve the contents of the decoded picture buffer 950 for display, or retrieve the contents of the decoded picture buffer into a display buffer. In some embodiments, the display device receives pixel values from the decoded picture buffer 950 via pixel transfer.
[0118] The motion compensation module 930 generates predicted pixel data 913 from decoded pixel data 917 stored in the decoded picture buffer 950 according to motion compensated MVs (MC MVs). These motion compensated MVs are decoded by adding residual motion data received from the bitstream 995 to predicted MVs received from the MV prediction module 975.
[0119] The MV prediction module 975 generates a predicted MV based on a reference MV generated for decoding a previous video frame, for example, a motion-compensated MV for performing motion compensation. The MV prediction module 975 retrieves the reference MV of the previous video frame from the MV buffer 965. The video decoder 900 stores the motion-compensated MV generated for decoding the current video frame in the MV buffer 965 as a reference MV for generating the predicted MV.
[0120] The loop filter 945 performs filtering or smoothing operations on the decoded pixel data 917 to reduce encoding and decoding artifacts, particularly at pixel block boundaries. In some embodiments, the filtering or smoothing operations performed by the loop filter 945 include a deblocking filter (DBF), sample adaptive offset (SAO), and / or an adaptive loop filter (ALF).
[0121] Figure 10A portion of video decoder 900 that implements parallel processing of rectangular and raster scan slices is shown. The figure conceptually illustrates pixel processing unit 1050, which may include processing units (e.g., modules 911, 910, 945, 950, 940, 925) that implement the prediction, (inverse) transform, (inverse) quantization, and filtering operations of the decoding loop. Pixel processing unit 1050 is capable of parallel processing in multiple computational threads 1051-1059 (labeled 1 through N). Threads 1051-1059 perform decoding operations on encoded video data provided by entropy decoder 990, converting it into pixel data for display device 905. Decoding operations can also be based on pixel data provided by decoded picture buffer 950.
[0122] The entropy decoder 990 may receive slice parameters to indicate whether the current picture is divided into rectangular or raster scan slices. The slice parameters may also include parameters that define each slice (e.g., starting and ending CTUs, left and right boundaries, height and width, etc.). The entropy decoder 990 receives the slice parameters as syntax elements from the bitstream 995 and provides them to the pixel processing unit 1050. Based on these slice parameters, the computing threads 1051-1059 are assigned different sets of CTUs of the current picture. CTUs of different slices can be assigned to different threads 1051-1059. In some embodiments, when wavefront parallel processing (WPP) is enabled, CTUs of different rows of the same slice or different slices can be assigned to different threads 1051-1059.
[0123] The coded video data provided to threads 1051-1059 comes from NAL units in the bitstream 995. The data for each slice is packaged into a NAL unit for transmission, regardless of whether the slice is rectangular or raster scan. When using WPP, CTUs from different rows of the same slice can be collected and transmitted as a single NAL unit.
[0124] Figure 11 The process 1100 of decoding a video picture into rectangular or raster scan slices is conceptually illustrated. In some embodiments, one or more processing units (e.g., processors) of a computing device implementing decoder 900 perform process 1100 by executing instructions stored in a computer-readable medium. In some embodiments, an electronic device implementing decoder 900 performs process 1100.
[0125] The decoder receives (at block 1105) data (eg, a bitstream of encoded video) to be decoded as a current picture having one or more slices. Each slice includes one or more data blocks (eg, CTUs).
[0126] The decoder receives (at block 1110) parameters for a slice of the current picture. The decoder determines (at block 1115) whether the slice is a rectangular slice or a raster scan slice. In some embodiments, the decoder may receive a flag from a syntax element in the bitstream (e.g., pps_rect_slice_flag) indicating whether the slice is a rectangular slice or a raster scan slice. If the slice is (to be decoded as) a rectangular slice, the decoder receives (at block 1120) parameters indicating the top left CTU of the slice, the width of the slice, and the height of the slice. If the slice is (to be decoded as) a raster scan slice, the decoder signals (at block 1125) parameters indicating the starting CTU and ending CTU of the slice (but not including the width or height of the slice).
[0127] The decoder extracts (at block 1130) a network abstraction layer (NAL) unit corresponding to the slice from the received data (e.g., a bitstream of encoded video). In some embodiments, the size of the NAL unit is defined based on the size of the slice and does not include other slices. The NAL unit conveys data for the slice and does not include other slices, and different slices of the current picture are conveyed by different NAL units.
[0128] The decoder then decodes the CTUs of the slice based on the NAL units (at block 1140). The decoder may process CTUs in two or more different slices of the current picture in parallel or simultaneously. The decoder may perform wavefront parallel processing (WPP) by processing different groups (e.g., rows) of CTUs of a slice in parallel with different delays. In some embodiments, the decoder may derive the right and left boundaries of the slice based on the slice parameters. In some embodiments, when the slice is a raster scan slice, the right boundary of the slice is derived based on the width of the current picture, particularly when the raster scan slice is rectangular in shape. In some embodiments, the current picture may include adjacent but misaligned rectangular slices (e.g., slices 411 and 414). The decoded CTUs may be provided for display as part of a reconstructed current picture.
[0129] VI. Example Electronic System
[0130] Many of the features and applications described above are implemented as software processes that are specified as a set of instructions recorded on a computer-readable storage medium (also referred to as a computer-readable medium). When these instructions are executed by one or more computing or processing units (e.g., one or more processors, processor cores, or other processing units), they cause the processing units to perform the operations indicated in the instructions. Examples of computer-readable media include, but are not limited to, CD-ROMs, flash drives, random access memory (RAM) chips, hard disks, erasable programmable read-only memories (EPROMs), electrically erasable programmable read-only memories (EEPROMs), and the like. Computer-readable media do not include carrier waves and electronic signals transmitted over wireless or wired connections.
[0131] In this specification, the term "software" is intended to include firmware residing in read-only memory or application programs stored in magnetic storage that can be read into memory by a processor for processing. In addition, in some embodiments, multiple software inventions can be implemented as sub-parts of a larger program while still maintaining independent software inventions. In some embodiments, multiple software inventions can also be implemented as independent programs. Finally, any combination of independent programs that jointly implement the software inventions described herein are within the scope of this disclosure. In some embodiments, when the software program is installed and run on one or more electronic systems, one or more specific machine implementations are defined to execute and perform the operations of the software program.
[0132] Figure 12 The electronic system 1200 according to some embodiments of the present disclosure is conceptually illustrated. The electronic system 1200 can be a computer (e.g., a desktop computer, a personal computer, a tablet computer, etc.), a phone, a PDA, or any other type of electronic device. Such an electronic system includes various types of computer-readable media and interfaces for various other types of computer-readable media. The electronic system 1200 includes a bus 1205, a processing unit 1210, a graphics processing unit (GPU) 1215, system memory 1220, a network 1225, a read-only memory 1230, a permanent storage device 1235, an input device 1240, and an output device 1245.
[0133] Bus 1205 collectively represents all system, peripheral, and chipset buses that communicatively connect the numerous internal devices of electronic system 1200. For example, bus 1205 communicatively connects processing unit 1210 with GPU 1215, read-only memory 1230, system memory 1220, and permanent storage device 1235.
[0134] From these various memory units, processing unit 1210 retrieves instructions to execute and processes data to perform the processes of the present disclosure. In various embodiments, the processing unit may be a single processor or a multi-core processor. Certain instructions are passed to and executed by GPU 1215. GPU 1215 may offload various computations or supplement the image processing provided by processing unit 1210.
[0135] Read-only memory (ROM) 1230 stores static data and instructions used by processing unit 1210 and other modules of the electronic system. Persistent storage device 1235, on the other hand, is a read-write storage device. This device is a non-volatile memory unit that stores instructions and data even when electronic system 1200 is turned off. Some embodiments of the present disclosure use a mass storage device (e.g., a magnetic or optical disk and its corresponding disk drive) as permanent storage device 1235.
[0136] Other embodiments use removable storage devices (e.g., floppy disks, flash memory devices, etc. and their corresponding disk drives) as permanent storage devices. Similar to permanent storage device 1235, system memory 1220 is a read-write storage device. However, unlike storage device 1235, system memory 1220 is a volatile read-write memory, such as random access memory. System memory 1220 stores some instructions and data used by the processor at runtime. In some embodiments, processes according to the present disclosure are stored in system memory 1220, permanent storage device 1235 and / or read-only memory 1230. For example, various memory units include instructions for processing multimedia clips according to certain embodiments. From these various memory units, processing unit 1210 retrieves instructions to execute and processes data to perform the processes of certain embodiments.
[0137] Bus 1205 is also connected to input and output devices 1240 and 1245. Input device 1240 enables a user to convey information and select commands to the electronic system. Input device 1240 includes, for example, an alphanumeric keyboard and pointing device (also known as a "cursor control device"), a camera (e.g., a webcam), a microphone, or similar device for receiving voice commands. Output device 1245 displays images generated by the electronic system or outputs data in other ways. Output device 1245 includes a printer and a display device, such as a cathode ray tube (CRT) or a liquid crystal display (LCD), as well as a speaker or similar audio output device. Some embodiments include devices such as a touch screen that can serve as both an input device and an output device.
[0138] Finally, if Figure 12 As shown, bus 1205 also couples electronic system 1200 to a network 1225 via a network adapter (not shown). In this manner, the computer can become part of a computer network, such as a local area network ("LAN"), a wide area network ("WAN"), or an intranet, or a network of networks, such as the Internet. Any or all components of electronic system 1200 may be used in conjunction with the present disclosure.
[0139] Certain embodiments include electronic components, such as microprocessors, storage, and memory, which store computer program instructions in machine-readable or computer-readable media (also referred to as computer-readable storage media, machine-readable media, or machine-readable storage media). Some examples of such computer-readable media include RAM, ROM, compact disc read-only memory (CD-ROM), compact disc recordable memory (CD-R), compact disc rewritable memory (CD-RW), read-only digital versatile discs (e.g., DVD-ROM, dual-layer DVD-ROM), various recordable / rewritable DVDs (e.g., DVD-RAM, DVD-RW, DVD+RW, etc.), flash memory (e.g., SD cards, mini-SD cards, micro-SD cards, etc.), magnetic and / or solid-state drives, read-only and recordable Blu-ray discs, ultra-density optical discs, any other optical or magnetic media, and floppy disks. Computer-readable media can store a computer program executed by at least one processing unit and include an instruction set for performing various operations. Examples of computer programs or computer code include machine code, such as code generated by a compiler, and files containing high-level code that are executed by a computer, electronic component, or microprocessor using an interpreter.
[0140] While the above discussion primarily involves microprocessors or multi-core processors executing software, many of the features and applications described above are performed by one or more integrated circuits, such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs). In some embodiments, these integrated circuits execute instructions stored on the circuits themselves. Additionally, some embodiments execute software stored in programmable logic devices (PLDs), ROM, or RAM devices.
[0141] As used in this specification and any claims herein, the terms "computer," "server," "processor," and "memory" refer to electronic or other technological devices. These terms do not include individuals or groups of individuals. For purposes of this specification, the terms display or showing mean displaying on an electronic device. The terms "computer-readable medium," "computer-readable media," and "machine-readable medium," as used in this specification and any claims herein, are strictly limited to tangible, physical objects that store information in a computer-readable form. These terms do not include any wireless signals, wired download signals, or any other transient signals.
[0142] Although the present disclosure has been described with many specific details, those skilled in the art will recognize that the present disclosure may be embodied in other specific forms without departing from the spirit of the present disclosure. Figure 8 and Figure 11 ) conceptually illustrates the processes. The specific operations of these processes may not be performed in the exact order shown and described. The specific operations may not be performed in a continuous series of operations, and different specific operations may be performed in different embodiments. In addition, the process may be implemented using several sub-processes or as part of a larger macro-process. Therefore, it will be understood by those skilled in the art that the present disclosure should not be limited to the foregoing illustrative details, but should be defined by the appended claims.
[0143] Additional Notes
[0144] The subject matter described herein sometimes illustrates different components contained within different components or connected to different other components. It should be understood that these illustrated architectures are merely examples and that many other architectures can be implemented to achieve the same functionality. Conceptually, any arrangement of components to achieve the same functionality is effectively "associated" to achieve the desired functionality. Therefore, any two components combined herein to achieve a particular functionality can be considered to be "associated" together to achieve the desired functionality, regardless of the architecture or intermediate components. Similarly, any two components so associated can also be considered to be "operably connected" or "operably coupled" together to achieve the desired functionality, and any two components that can be so associated can also be considered to be "operably coupled" to achieve the desired functionality. Specific examples of operable coupling include, but are not limited to, physically matable and / or physically interactive components and / or wirelessly interactive and / or wirelessly interactive components and / or logically interactive and / or logically interactive components.
[0145] Furthermore, with respect to the use of almost any plural and / or singular terms herein, those skilled in the art can translate from the plural to the singular and / or from the singular to the plural as appropriate, depending on the context and / or application. Various singular / plural permutations may be explicitly listed herein for clarity.
[0146] Furthermore, those skilled in the art will understand that, in general, the terms used herein, particularly in the appended claims, such as the bodies of the appended claims, are generally to be considered "open" terms. For example, the term "comprising" should be interpreted as "including, but not limited to," the term "having" should be interpreted as "having at least," and the term "includes" should be interpreted as "including, but not limited to," etc. Those skilled in the art will also understand that if a specific number is intended in an introduced claim recitation, such intent will be explicitly stated in the claim; if no such statement is made, no such intent is present. For example, to aid understanding, the following appended claims may contain the use of the introductory phrases "at least one" and "one or more" to introduce claim recitations. However, the use of these phrases should not be construed to imply that a claim recitation introduced by the indefinite article "a" or "an" limits any particular claim containing such introduced claim recitation to containing only one such recitation, even if the same claim includes the introductory phrases "one or more" or "at least one" and an indefinite article such as "a" or "an." For example, "a" and / or "an" should be interpreted as "at least one" or "one or more." The same applies to the use of definite articles used to introduce claim recitations. Furthermore, even if a specific number of claim recitations is explicitly recited, those skilled in the art will recognize that such recitation should be interpreted as at least the recited number, e.g., the single recitation "two recitations," without other modifiers, means at least two recitations, or two or more recitations. Furthermore, where conventions similar to "at least one of A, B, and C, etc.," are used, generally, such constructions are to be interpreted in a manner that is understood by those skilled in the art to be conventional, e.g., "a system having at least one of A, B, and C" would include, but are not limited to, systems having only A, only B, only C, A and B together, A and C together, B and C together, and / or A, B, and C together, etc. Where conventions similar to "at least one of A, B, or C, etc.," generally, such constructions are to be interpreted in a manner that is understood by those skilled in the art to be conventional, e.g., "a system having at least one of A, B, or C" would include, but are not limited to, systems having only A, only B, only C, A and B together, A and C together, B and C together, and / or A, B, and C together, etc. Those skilled in the art will also understand that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to include the possibility of one, either, or both terms. For example, the phrase "A or B" will be understood to include the possibilities of "A" or "B" or "A and B."
[0147] In summary, it will be appreciated that various embodiments of the present disclosure have been described herein for illustrative purposes and that various modifications may be made without departing from the scope and spirit of the present disclosure. Therefore, the various embodiments disclosed herein are not intended to be limiting, with the true scope and spirit being indicated by the following claims.
Claims
1. A video decoding method, comprising: Receive data to be decoded into a current picture, including one or more slices, each slice including one or more coding tree units (CTUs); receiving a plurality of parameters of a slice of the current picture, the plurality of parameters comprising an upper left CTU of the slice, a width of the slice, and a height of the slice; Extracting a network abstraction layer (NAL) unit of the slice from the received data, wherein the size of the NAL unit is defined based on the size of the slice and does not include other slices; and The one or more CTUs of the slice are decoded based on the NAL unit.
2. The video decoding method of claim 1, further comprising processing CTUs in two or more different slices of the current picture in parallel.
3. The video decoding method of claim 1 , wherein different groups of CTUs in a slice are processed in parallel with different delays. The video decoding method of claim 1 , wherein the specified parameters are parameters of rectangular slices. The video decoding method as claimed in claim 4 , wherein the current picture comprises at least two adjacent but non-aligned rectangular slices.
6. The video decoding method of claim 1 , further comprising receiving a flag to indicate whether the slice is a rectangular slice or a raster scan slice, wherein: When the flag indicates that the slice is a raster scan slice, the parameters of the slice indicate a start CTU of the slice and an end CTU of the slice; and When the flag indicates that the slice is a rectangular slice, the parameters of the slice indicate the upper left CTU of the slice, the width of the slice, and the height of the slice. 7 . The video decoding method of claim 1 , further comprising deriving a right boundary and a left boundary of the slice based on the slice parameters. 8 . The video decoding method of claim 7 , wherein when the slice is a raster scan slice, the right boundary of the slice is derived based on the width of the current picture. 9 . The video decoding method of claim 1 , wherein the NAL unit transmits data of the slice but does not include other slices.
10. The video decoding method of claim 1, wherein different slices of the current picture are transmitted by different NAL units.
11. A video encoding method, comprising: Receive data to be encoded as a current picture, including one or more slices, each slice including one or more coding tree units (CTUs); Signaling multiple parameters of a slice of the current picture, the multiple parameters including the top left CTU of the slice, the width of the slice, and the height of the slice; encoding the one or more CTUs of the slice based on the received data; and The coded CTUs are packaged into a network abstraction layer (NAL) unit for transmission or storage. The size of the NAL unit is defined based on the size of the slice and does not include other slices.
12. An electronic device comprising: A video encoder circuit configured to perform operations including: Receive data to be encoded as a current picture, including one or more slices, each slice including one or more coding tree units (CTUs); Signaling multiple parameters of a slice of the current picture, the multiple parameters including the top left CTU of the slice, the width of the slice, and the height of the slice; encoding the one or more CTUs of the slice based on the received data; and The coded one or more CTUs are packaged into a network abstraction layer (NAL) unit for transmission or storage, where the size of the NAL unit is defined based on the size of the slice and does not include other slices.
13. An electronic device comprising: A video decoding circuit is configured to perform operations including: Receive data to be decoded into a current picture, including one or more slices, each slice including one or more coding tree units (CTUs); receiving a plurality of parameters of a slice of the current picture, the plurality of parameters comprising an upper left CTU of the slice, a width of the slice, and a height of the slice; Extracting a network abstraction layer (NAL) unit of the slice from the received data, the size of the NAL unit being defined based on the size of the slice and not based on other slices; and The one or more CTUs of the slice are decoded based on the NAL unit.