Method and apparatus for motion vector prediction for scalable video coding
By using interlayer motion mapping and temporal motion vector prediction techniques, the reference image and motion vector of the enhancement layer video block are predicted using information from the base layer, which solves the problem of low coding efficiency in existing technologies and achieves more efficient video coding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2013-08-29
- Publication Date
- 2026-03-31
AI Technical Summary
Existing video coding techniques struggle to effectively utilize inter-layer relationships for temporal motion vector prediction in heterogeneous networks, resulting in low coding efficiency.
By employing interlayer motion mapping information and temporal motion vector prediction (TMVP) technology, the reference image and motion vector of the enhancement layer video block are determined by juxtaposing the relationship between the base layer video block and the enhancement layer video block, and the motion information of the base layer is used for prediction.
It improves the coding efficiency of enhancement layer video blocks, reduces the number of bits in the bitstream, and enhances video service quality and coding efficiency.
Smart Images

Figure CN115243046B_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese invention patent application 201810073911.5, filed on August 29, 2013, entitled "Method and apparatus for motion vector prediction of scalable video coding," which in turn is a divisional application of Chinese invention patent application 201380045712.9, filed on August 29, 2013, entitled "Method and apparatus for motion vector prediction of scalable video coding," the contents of which are incorporated herein by reference.
[0002] Cross-references to related applications
[0003] This application claims the benefit of U.S. Provisional Patent Application No. 61 / 694,555, filed August 29, 2012; U.S. Provisional Patent Application No. 61 / 734,650, filed December 7, 2012; and U.S. Provisional Patent Application No. 61 / 866,822, filed August 16, 2013, the contents of which are incorporated herein by reference. Background Technology
[0004] Over the past two decades, digital video compression technology has developed significantly and been standardized, enabling efficient digital video communication, distribution, and consumption. Most commercially widely used standards were developed by ISO / IEC and ITU-T, such as MPEC-2 and H.263 (MPEG-4 Part 10). The emergence and maturation of video compression technology led to the development of High Efficiency Video Coding (HEVC).
[0005] Compared to traditional digital video services transmitted via satellite, cable, and terrestrial channels, heterogeneous networks offer a growing number of video applications available on both the client and network sides, including but not limited to video chat, mobile video, and streaming media. Smartphones, tablets, and TVs dominate the client side, where video can be transmitted via the Internet, mobile networks, and / or a combination of both. To improve user experience and video service quality, scalable video coding (SVC) can be implemented. In SVC, once the signal is encoded at the highest resolution, it can be decoded from a subset of the data stream according to the specific rate and resolution required by the application and supported by the client device. International video standards MPEG-2 Video, H.263, MPEG4 Visual, and H.264 have tools and / or configuration files that support scalability modes. Summary of the Invention
[0006] Interlayer motion mapping information is used to enable Temporal Motion Vector Prediction (TMVP) of enhancement layers in the bitstream. For example, a reference image for an enhancement layer video block can be determined based on a collocated base layer video block. The enhancement layer video block is associated with an enhancement layer in the bitstream, while the collocated base layer video block is associated with a base layer in the bitstream. For example, the enhancement layer video block is associated with an enhancement layer image, while the collocated base layer video block is associated with a base layer image. The collocated base layer video block can be determined by selecting a video block in the collocated base layer image characterized by the largest overlap area with the enhancement layer video block. The video block can be a computational unit at any level in the bitstream. The video block can be of any size (e.g., block size (e.g., 16x16), PU, SPU, etc.).
[0007] The reference image for the enhancement layer video block can be determined by determining the reference image of the juxtaposed base layer video block. The reference image for the enhancement layer video block can be a juxtaposed enhancement layer image of the reference image of the juxtaposed base layer video block. The reference image for the juxtaposed base layer video block can be determined, the reference image for the interlayer video block can be determined using the reference image of the juxtaposed base layer video block, and the reference image for the enhancement layer video block can be determined using the reference image of the interlayer video block. The interlayer video block can be collated with the enhancement layer video block and / or the base layer video block.
[0008] The MV of the enhancement layer video block can be determined based on the motion vector (MV) of the juxtaposed base layer video block. The MV of the juxtaposed base layer video block can be determined by scaling the MV of the juxtaposed base layer video block according to the spatial ratio between the base and enhancement layers, thereby determining the MV of the enhancement layer video block.
[0009] The MV of the juxtaposed base layer video block can be determined, the MV of the juxtaposed base layer video block can be scaled according to the spatial ratio between the base and enhancement layers to determine the MV of the interlayer video block, and the MV of the enhancement layer video block can be predicted based on the MV of the interlayer video block, thereby determining the MV of the enhancement layer video block. For example, the MV of the enhancement layer video block can be predicted based on the MV of the interlayer video block by performing time scaling on the MV of the interlayer video block. The interlayer video block can be juxtaposed with the enhancement layer video block and / or the base layer video block.
[0010] TMVP can be performed using the MV and / or reference image of the inter-layer video block on the enhancement layer video block. The enhancement layer video block can be decoded based on the reference image and / or MV of the enhancement layer video block and / or the reference image and / or MV of the inter-layer video block.
[0011] The method includes receiving a bitstream containing a base layer and an enhancement layer, and decoding the enhancement layer in the encoded bitstream using Temporal Motion Vector Prediction (TMVP). An interlayer reference image can be used as a juxtaposed reference image for the enhancement layer TMVP.
[0012] Decoding enhancement layers in an encoded bitstream using TMVP may include decoding the enhancement layer image using TMVP. Decoding the enhancement layer image using TMVP may include determining the motion vector (MV) field of the interlayer reference image and decoding the enhancement layer image based on the MV field of the interlayer reference image. The MV field of the interlayer reference image may be determined based on the MV field of the juxtaposed base layer image. The MV field includes the reference image index and MV of the video block of the interlayer reference image. For example, the MV field may include one or more index values (e.g., depending on whether it is a P-slice or a B-slice) and MV of one or more video blocks of the interlayer reference image. Determining the MV field of the interlayer reference image may include determining the compressed MV field of the juxtaposed base layer image and determining the MV field of the interlayer reference image based on the compressed MV field of the juxtaposed base layer image.
[0013] Determining the MV domain of an interlayer reference image may include determining the MV of a video block in the interlayer reference image and a reference image. Determining the MV of a video block in the interlayer reference image and a reference image may include determining the reference image of the interlayer video block based on the reference image of the juxtaposed base layer video block and determining the MV of the interlayer video block based on the MV of the juxtaposed base layer video block. The juxtaposed base layer video block can be determined by selecting a video block in the juxtaposed base layer image characterized by the largest overlap area with the interlayer reference image.
[0014] Determining the reference image for interlayer video blocks may include determining a reference image for juxtaposed base layer video blocks and determining a reference image for interlayer video blocks. The reference image for interlayer video blocks may be an interlayer reference image of the juxtaposed base layer video block reference image. Determining the MV of interlayer video blocks may include determining the MV of the juxtaposed base layer video blocks and scaling the MV of the juxtaposed base layer video blocks according to the spatial ratio between the base layer and enhancement layers to determine the MV of the interlayer video blocks.
[0015] The MV field of an enhancement layer video block can be determined based on the MV field of the inter-layer video block. The enhancement layer video block can be concatenated in parallel with intermediate layer video blocks and / or base layer video blocks. For example, the reference image of the enhancement layer video block can be determined based on a reference image of the inter-layer video block (e.g., possibly a juxtaposed enhancement layer image). The MV of the enhancement layer video block can be determined based on the MV of the inter-layer video block. For example, the MV of the enhancement layer video block can be scaled (e.g., time-scaled) to determine the MV of the inter-layer video block. The enhancement layer video block can be decoded based on its MV field.
[0016] The method includes receiving a bitstream containing a base layer and enhancement layers, and inter-layer motion mapping information, and performing inter-layer motion prediction for the enhancement layers. It can be determined whether inter-layer motion prediction for the enhancement layers is enabled based on the inter-layer mapping information.
[0017] Interlayer mapping information can be signaled at the sequence level of the bitstream. For example, interlayer mapping information can be a variable (e.g., a flag), which can be signaled at the sequence level of the bitstream. Interlayer mapping information can be inferred at the sequence level of the bitstream. Interlayer mapping information can be signaled using variables (e.g., flags) in the video parameter set (VPS) of the bitstream (e.g., the interlayer mapping information can be a flag in the VPS of the bitstream). For example, interlayer mapping information can be signaled using variables (flags) in the sequence parameter set (SPS) of the bitstream (e.g., the interlayer mapping information can be a flag in the SPS of the bitstream). For example, interlayer mapping information can be signaled using variables (flags) in the picture parameter set (PPS) of the bitstream (e.g., the interlayer mapping information can be a flag in the PPS of the bitstream). Attached Figure Description
[0018] Figure 1 This is an example diagram illustrating a hierarchical structure with additional inter-layer predictions for SVC spatial hierarchical coding.
[0019] Figure 2 This is an example inter-layer prediction structure diagram considered for HEVC scalable coding.
[0020] Figure 3 This is an example diagram illustrating Spatial Motion Vector (MV) Prediction (SMVP).
[0021] Figure 4 This is an example diagram illustrating Time MV Prediction (TMVP).
[0022] Figure 5 This is an example diagram illustrating the replication of the predicted structure from the base layer to the upsampled base layer.
[0023] Figure 6 This is an example diagram showing the relationship between the SPU of the upsampled base layer and the SPU of the original base layer.
[0024] Figures 7A-7C This is an example diagram showing the relationship between slices of the base layer image and slices of the processed base layer image.
[0025] Figure 8A This is a graph showing the MV prediction over a short period of time.
[0026] Figure 8B This is a graph showing the MV prediction of the time short-term MV based on the mapped short-term MV.
[0027] Figure 9A This is an example diagram showing MV prediction over a long period of time.
[0028] Figure 9B This is an example diagram showing the prediction of the long-term MV based on the mapped long-term MV.
[0029] Figure 10A This is an example diagram showing the prediction of the MV of the short-term MV based on the long-term MV.
[0030] Figure 10B This is an example diagram showing the MV prediction of the short-term MV based on the mapped long-term MV.
[0031] Figure 10C This is an example diagram showing the prediction of the long-term MV based on the short-term MV.
[0032] Figure 10D This is an example diagram showing the prediction of the long-term MV based on the mapped short-term MV.
[0033] Figure 11A This is an example diagram showing the disabled MV prediction based on the interlayer MV to the short-term MV over time.
[0034] Figure 11B This is an example diagram showing the disabled MV prediction of inter-layer MV based on short-term MV over time.
[0035] Figure 11C This is an example diagram showing the disabled MV prediction of inter-layer MV based on the mapped short-term MV.
[0036] Figure 12A This is an example diagram showing the disabled MV prediction based on the interlayer MV to the long-term MV over time.
[0037] Figure 12B This is an example diagram showing the disabled MV prediction of inter-layer MV based on long-term MV over time.
[0038] Figure 12C This is an example diagram showing the disabled MV prediction of inter-layer MV based on the long-term MV of the mapping.
[0039] Figure 13A This is an example diagram showing the MV prediction between two layers when Te = Tp.
[0040] Figure 13B This is an example diagram illustrating the disabled MV prediction between two layers when Te≠Tp.
[0041] Figure 14AThis is a system diagram of an example communication system that can implement one or more of the disclosed implementation methods.
[0042] Figure 14B It is possible Figure 14A The system diagram shown is of an example wireless transmit / receive unit (WTRU) used in the communication system.
[0043] Figure 14C It is possible Figure 14A The system diagram shows an example radio access network and an example core network used in the communication system shown.
[0044] Figure 14D It is possible Figure 14A The system diagram shows another example of a radio access network and another example of a core network used in the communication system shown.
[0045] Figure 14E It is possible Figure 14A The system diagram shows another example of a radio access network and another example of a core network used in a communication system.
[0046] Figure 15 This is an example block diagram illustrating a block-based video encoder.
[0047] Figure 16 This is an example block diagram illustrating a block-based video decoder.
[0048] Figure 17 This is a diagram illustrating an example communication system. Detailed Implementation
[0049] Encoding and / or decoding (e.g., transmission and / or reception) of bitstreams (e.g., partial bitstreams) can be provided to deliver video services with lower temporal resolution, lower spatial resolution, and / or reduced fidelity while preserving reconstruction quality closely related to the rate of the partial bitstream, for example, through the scalability extension of H.264. Figure 1 This is a diagram illustrating an example of a scalable structure with additional inter-layer prediction for SVC spatial scalability coding. Figure 100 illustrates an example of a two-layer SVC inter-layer prediction mechanism that can improve the efficiency of scalable coding. Similar mechanisms can be used in multi-layer SVC coding structures. In Figure 100, the base layer and enhancement layer can represent two adjacent spatially scalable layers with different resolutions. Within a single layer (e.g., the base layer and / or enhancement layer), an H.264 encoder, for example, can employ motion-compensated prediction and / or intra-layer prediction. Inter-layer prediction can use base layer information (e.g., spatial texture, motion vectors, reference image index values, and residual signals, etc.) to improve the coding efficiency of the enhancement layer. When decoding the enhancement layer, the SVC can be fully reconstructed without using reference images from lower layers (e.g., dependent layers of the current layer).
[0050] Interlayer prediction can be employed in scalable coding systems (e.g., HEVC scalable coding extension) to, for example, determine the relationships between multiple layers and / or improve the efficiency of scalable coding. Figure 2 This is a diagram illustrating an example inter-layer prediction structure considered for HEVC scalable coding. For example, Figure 200 may illustrate an example of a scalable structure with additional inter-layer predictions for HEVC spatial scalable coding. The prediction for the enhancement layer can be generated from motion-compensated predictions based on the reconstructed base layer signal (e.g., after upsampling if the spatial resolution between the two layers differs), temporal predictions in the current enhancement layer, and / or the average of the base layer reconstructed signal and the temporal prediction signal. Full reconstruction of the lower-layer image can be performed. Similar implementations can be used in scalable coding systems with more than two layers (e.g., HEVC scalable coding systems with more than two layers).
[0051] HEVC utilizes advanced motion-compensated prediction techniques to determine inherent inter-layer image redundancy in video signals, for example, by predicting pixels in the current video image using pixels from an encoded video image. The displacement of the current prediction unit (PU) to be encoded and its displacement between one or more matching blocks in a reference image (e.g., neighboring PUs) can be represented by motion vectors in motion-compensated prediction. MV can comprise two parts, MVx and MVy. MVx and MVy can represent horizontal and vertical displacements, respectively. MVx and MVy can be encoded directly or not.
[0052] Advanced Motion Vector Prediction (AMVP) can be used to predict motion vectors (MVs) based on one or more MVs of neighboring physical units (PUs). The difference between the true MV and the MV prediction (predictor) can be encoded. By encoding the difference in MVs (e.g., encoding only), the number of bits used for MV encoding can be reduced. MVs for prediction can be obtained from spatial and / or temporal neighborhoods. A spatial neighborhood can refer to those spatial PUs surrounding the current encoded PU. A temporal neighborhood can refer to those juxtaposed PUs in nearby images. In HEVC, to obtain accurate MV predictions, prediction candidate values from spatial and / or temporal neighborhoods can be grouped together to form a candidate value list, and the best prediction value can be selected to predict the MV of the current PU. For example, the best MV prediction value can be selected based on Lagrange rate distortion (RD) consumption, etc. MV differences can be encoded into a bitstream.
[0053] Figure 3This is a diagram illustrating an example of Spatial MV Prediction (SMVP). Graph 300 may show examples of a neighboring reference image 310, a current reference image 320, and a current image 330. In the current image (CurrPic) 330 to be encoded, the hashed square may be the current PU (CurrPU) 332. CurrPU 332 may have the best-matching block of the current reference PU (CurrRefPU) 322 located in the reference image (CurrRefPic) 320. The MV (MV2 340) of the CurrPU can be predicted. For example, in HEVC, the spatial neighborhood of the current PU can be the PU above, to the left, to the upper left, to the lower left, or to the upper right of the current PU 332. For example, the neighboring PU 334 shown is the upper neighbor of CurrPU 332. For example, the reference images (NeighbRefPic) 310, PU 314 and MV (MV1 350) of the neighboring PU (NeighbPU) are known because NeighbPU 334 is encoded before CurrPU 332.
[0054] Figure 4 This is a diagram illustrating an example of Temporal MV Prediction (TMVP). Graph 400 may include four images, such as a juxtaposed reference image (ColRefPic) 410, CurrRefPic 420, a juxtaposed image (ColPic) 430, and CurrPic 440. In the current image (CurrPic 440) to be encoded, a hashed square (CurrPU 442) may be the current PU. The hashed square (CurrPU 442) may have a best-matching block (CurrRefPU 422) located in the reference image (CurrRefPic 420). The MV (MV2 460) of the CurrPU can be predicted. For example, in HEVC, the temporal neighborhood of the current PU may be the juxtaposed PU (ColPU) 432, for example, ColPU 432 is part of the neighboring image (ColPic) 430. For example, the reference image of ColPU (ColRefPic 410), PU 412 and MV (MV1 450) are known because ColPic 430 is encoded before CurrPic 440.
[0055] The motion between PUs is a uniform translation. The MV between two PUs is proportional to the temporal interval between the moments when the two associated images were captured. The motion vector prediction value can be scaled before predicting the MV of the current PU (e.g., in AMVP). For example, the temporal interval between CurrrPic and CurrrRefPic can be referred to as TB. For example, the time interval between CurrrPic and NeighbRefPic (e.g., ...) Figure 3(in Chinese) or ColPic and ColRefPic (e.g., Figure 4 The time interval (in the middle) can be referred to as TD. Given TB and TD, the scaled prediction of MV2 (e.g., MV) can be equal to:
[0056]
[0057] Both short-term and long-term reference images are supported. For example, a reference image stored in the decoded image buffer (DPB) can be labeled as a short-term or long-term reference image. For example, in equation (1), if one or more of the reference images are long-term reference images, scaling of motion vectors may be disabled.
[0058] This paper describes the use of MV prediction for multi-layer video coding. The examples described herein can be used with the HEVC standard as the underlying single-layer coding standard and with scalable systems that include two spatial layers (e.g., an enhancement layer and a base layer). The examples described herein can be applied to other scalable coding systems using other types of underlying single-layer codecs, more than two layers, and / or supporting other types of scalability.
[0059] When decoding a video slice (e.g., a P-slice or a B-slice) begins, one or more reference images from the DPB can be added to the reference image list for the P-slice (e.g., list 0) and / or the two reference image lists for the B-slice (e.g., list 0 and list 1) used for motion compensation prediction. Scalable coding systems can utilize temporal reference images of enhancement layers and / or processed reference images from base layers (e.g., upsampled base layer images if the two layers have different spatial resolutions) for motion compensation prediction. When predicting the MV of the current image in the enhancement layer, the inter-layer MV pointing to the processed reference image from the base layer can be used to predict the temporal MV pointing to the temporal reference image of the enhancement layer. Temporal MVs can also be used to predict inter-layer MVs. Because these two types of MVs are almost uncorrelated, this leads to a loss of efficiency in MV prediction for enhancement layers. Single-layer codecs do not support predicting inter-enhancement layer temporal MVs based on inter-base layer temporal MVs, which are highly correlated and can be used to improve MV prediction performance.
[0060] The MV prediction process can be simplified and / or compression efficiency can be improved for multi-layer video coding. MV prediction in enhancement layers can be backward compatible with the MV prediction process of single-layer encoders. For example, there may be MV prediction implementations that do not require any changes to block-level operations in the enhancement layer, allowing single-layer encoder and decoder logic to be reused in the enhancement layer. This reduces the implementation complexity of scalable systems. MV prediction in enhancement layers can distinguish between temporal MVs pointing to temporal reference images in the enhancement layer and inter-layer MVs pointing to processed (e.g., upsampled) reference images from the base layers. This improves coding efficiency. MV prediction in enhancement layers can support MV prediction between temporal MVs of enhancement layer images and between temporal MVs of base layer images, which improves coding efficiency. When the spatial resolution between two layers is different, the temporal MVs between base layer images can be scaled according to the ratio of the spatial resolutions between the two layers.
[0061] The implementation described herein relates to an inter-layer motion information mapping algorithm for the base layer MV, for example, enabling the base layer MV mapped in the AMVP process to be used to predict the enhancement layer MV (e.g., Figure 4 (In the TMVP mode). Block-level operations may remain unchanged. Single-layer encoders and decoders can be used for MV prediction of enhancement layers without modification. This article will describe an MV prediction tool with block-level modifications for the encoding and decoding process of enhancement layers.
[0062] Interlayers may include processed base layers and / or upsampled base layers. For example, interlayers, processed base layers, and / or upsampled base layers can be used interchangeably. Interlayer reference images, processed base layer reference images, and / or upsampled base layer reference images can be used interchangeably. Interlayer video blocks, processed base layer video blocks, and / or upsampled base layer video blocks can be used interchangeably. Timing relationships may exist between enhancement layers, interlayers, and base layers. For example, video blocks and / or images of enhancement layers may be temporally associated with corresponding video blocks and / or images of interlayers and / or base layers.
[0063] A video block can be a unit of operation at any level and / or bitstream. For example, a video block can be a unit of operation at the image level, block level, and slice level, etc. A video block can be of any size. For example, a video block can refer to a video block of any size, such as a 4x4 video block, an 8x8 video block, and a 16x16 video block, etc. For example, a video block can refer to a prediction unit (PU) and a minimum PU (SPU), etc. A PU can be a video block unit used to carry information related to motion prediction, such as including a reference image index and a motion signature (MV). A PU can include one or more minimum PUs (SPUs). Although an SPU within the same PU refers to the same reference image with the same MV, storing motion information in units of SPUs in some implementations facilitates motion information retrieval. Motion information (e.g., MV domains) can be stored in units of video blocks, such as PUs and SPUs, etc. Although the examples described herein are illustrated with reference to an image, video blocks, PUs, and / or SPUs, and any unit of operation of any size (e.g., image, video block, PU, SPU, etc.) can also be used.
[0064] The texture of the reconstructed base layer signal may be processed for interlayer prediction of enhancement layers. For example, interlayer reference image processing may include upsampling of one or more base layer images when spatial scalability between two layers is enabled. It may not be possible to correctly generate motion-related information (e.g., motion velocity m, reference image list, reference image index value, etc.) for the processed reference image based on the base layers. This is especially problematic when the temporal motion velocity m prediction is derived from the processed base layer reference image (e.g., as shown in the image). Figure 4 As shown, missing motion information can affect the prediction of the MV of the enhancement layer (e.g., via TMVP). For example, when the selected processed base layer reference image is used as a temporally neighboring image (ColPic) including the temporally juxtaposed PU (ColPU), TMPV will not function correctly if the MV prediction value (MV1) and the reference image (ColRefPic) for the processed base layer reference image are not correctly generated. To enable TMVP for MV prediction of the enhancement layer, inter-layer motion information mapping can be implemented, for example, as described herein. For example, an MV domain (including MV and reference image) for the processed base layer reference image can be generated.
[0065] One or more variables can be used to describe the reference image of the current video slice, such as a list of reference images ListX (e.g., X is 0 or 1), the index of the reference image in ListX refIdx, etc. Utilizing Figure 4For example, to obtain a reference image (ColRefPic) of juxtaposed PUs (ColPUs), reference images of the PUs (e.g., each PU) (ColPU) in the processed reference image (ColPic) can be generated. This can be broken down into generating a list of reference images for the ColPic and / or a reference image index for each ColPU (e.g., each ColPU) in the ColPic. Given a list of reference images, the generation of reference image indices for the PUs in the processed base layer reference image is described herein. Implementations relating to the composition of the reference image list of the processed base layer reference image are described herein.
[0066] Because the base layer and the processed base layer are related, it can be assumed that the base layer and the processed base layer have the same or substantially the same predictive dependencies. The predictive dependencies of the base layer images can be replicated to form a list of reference images for the processed base layer images. For example, if a base layer image BL1 is a temporal reference image of another base layer image BL2 with a reference image index refIdx in a list of reference images ListX (e.g., X is 0 or 1), then the processed base layer image pBL1 of BL1 can be added to the same list of reference images ListX (e.g., X is 0 or 1) with the same index refIdx as the processed base layer image pBL2 of BL2. Figure 5 This is a diagram illustrating an example of the replication of the prediction structure from the base layer to the upsampled base layer. Figure 500 shows an example of spatial scalability, where the same Class B structure used for motion prediction of the base layer is replicated (shown as solid lines in the figure), just like the motion information of the sampled base layer (shown as dashed lines in the figure).
[0067] A reference image for the processed base layer PU can be determined based on the juxtaposed base layer prediction unit (PU). For example, the juxtaposed base layer PUs within the processed base layer PUs can be determined. The juxtaposed base layer PUs can be determined by selecting the PU in the juxtaposed base layer image characterized by the largest overlap area with the processed base layer PU, for example, as described herein. A reference image for the juxtaposed base layer PUs can be determined. The reference image for the processed base layer PUs can be determined as a juxtaposed processed base layer reference image of the reference image for the juxtaposed base layer PUs. The reference image for the processed base layer PUs can be used for the TMVP of the enhancement layer and / or for decoding the enhancement layer (e.g., the juxtaposed enhancement layer PU).
[0068] The processed base layer PU is associated with the processed base layer image. The MV domain of the processed base layer image may include a reference image of the processed base layer PU, for example, a TMVP (e.g., a juxtaposed enhancement layer PU) for an enhancement layer image. A list of reference images is associated with the processed base layer image. The list of reference images of the processed base layer image may include one or more reference images of the processed base layer PU. Images in the processed base layer (e.g., each image) may inherit the same Image Order Count (POC) and / or Short-Term / Long-Term Image Tag from the corresponding images in the base layer.
[0069] Spatial scalability with a 1.5x upsampling rate is used as an example. Figure 6 This is a diagram illustrating an example relationship between the upsampled base layer SPU and the original base layer SPU. Graph 600 may show the upsampled base layer SPU (e.g., labeled u). i The blocks) and the SPU of the original base layer (e.g., labeled b) j The example relationships between blocks. For example, given various upsampling rates and coordinate values in an image, the SPUs in the upsampled base layer image can correspond to various numbers and / or proportions of SPUs from the original base layer image. For example, SPU u4 can cover four SPU regions in the base layer (e.g., b0, b1, b2, b3). SPU u1 can cover two base layer SPUs (e.g., b0 and b1). SPU u0 can cover one base layer SPU (e.g., b0). The MV domain mapping implementation can be used to estimate the reference image index and MV for the SPUs in the processed base layer image, for example, by utilizing the motion information of the corresponding SPUs from the original base layer image.
[0070] The MV of the processed base layer PU can be determined based on the MV of the juxtaposed base layer PU. For example, the juxtaposed base layer PU of the processed base layer PU can be determined. The MV of the juxtaposed base layer PU can be determined. The MV of the base layer PU can be scaled to determine the MV of the processed base layer PU. For example, the MV of the base layer PU can be scaled based on the space ratio between the base layer and the enhancement layer to determine the MV of the processed base layer PU. The MV of the processed base layer PU can be used for the TMVP of the enhancement layer (e.g., the juxtaposed enhancement layer PU) and / or for decoding the enhancement layer (e.g., the juxtaposed enhancement layer PU).
[0071] The processed base layer PU is associated (e.g., temporally associated) with the enhancement layer image (e.g., the PU of the enhancement layer image). The MV domain of the juxtaposed enhancement layer image is based on the MV of the processed base layer PU, for example, the TMVP of the enhancement layer image (e.g., the juxtaposed enhancement layer PU). The MV of the enhancement layer PU (e.g., the juxtaposed enhancement layer PU) can be determined based on the MV of the processed base layer PU. For example, the MV of the enhancement layer PU (e.g., the juxtaposed enhancement layer PU) can be predicted (e.g., spatially predicted) using the MV of the processed base layer PU.
[0072] The reference image for each SPU in the processed base layer image can be selected based on the reference image index value of the corresponding SPU in the base layer. For example, for an SPU in the processed base layer image, the main rule for determining the reference image index is the reference image index of the SPU from the base layer image that is used most frequently. For example, suppose an SPU u in the processed base layer image... h Corresponding to K SPUs b from the base layer i (i = 0, 1, ..., K-1), there are M reference images with index values {0, 1, ..., M-1} in the reference image list of the processed base layer image. Assume that the reference images are based on index values {r0, r1, ..., r...}. k-1 The set of reference images is used to predict the corresponding K SPUs in the base layer, where for i = 0, 1, ..., K-1, r i If u ∈{0,1,...,M-1}, then u h The reference image index value can be determined by equation (2):
[0073] r(u h ) = r l l = argmax i∈{0,1,...K-1} c(r i Equation (2)
[0074] Where C(r) i ), i = 0, 1, ..., K-1 are used to represent the reference image r i A counter indicating how many times it has been used. For example, if the base layer image has two reference images labeled {0,1} (M=2), and given a processed base layer image with u... h For the four (K=4) base layer SPUs predicted according to {0,1,1,1} (e.g., {r0,r1,…,r3} equals {0,1,1,1}), then according to equation (2), r(u h The value is set to 1. A reference image r with the minimum POC (Proof of Concept) is selected for the currently processed image. iFor example, because two images with smaller time intervals have better correlation (e.g., to break the application of equation (2) when C(r) is used). i (constraints).
[0075] Different SPUs in the processed base layer image can correspond to various numbers and / or proportions of SPUs from the original base layer (e.g., such as...). Figure 6 (As shown). The reference image index of the base layer SPU with the largest coverage area can be selected to determine the reference image of the corresponding SPU in the processed base layer. For a given SPU in the processed base layer... h Its reference image index can be determined by equation (3):
[0076] r(u h ) = r l l = argmax i∈{0,1,...K-1} S i Equation (3)
[0077] Among them, S i It is the i-th corresponding SPU b in the base layer i The coverage area. Select a reference image r with the minimum POC distance to the currently processed image. i For example, when two or more corresponding SPUs have the same coverage area, to break equation (3)S i Constraints.
[0078] The corresponding base layer SPU b can be configured through internal patterns. j Encoding. Reference image index (e.g., corresponding base layer SPU b) j The reference image index can be set to -1 and is ignored when applying equations (2) and / or (3). If the corresponding base layer SPU b j It is internally encoded, SPU u h The reference image index can be set to -1 or marked as invalid for TMVP.
[0079] For the given SPU u in the processed base layer h Its corresponding SPU b i The areas may be different. The area-based implementation described herein can be used to estimate the MV of each SPU in the processed base layer image.
[0080] To estimate a SPU u in the processed base layer image h The MV can be found in the base layer SPU candidate b. i Selected with SPU u hThe MV of the base layer SPU b1 with the largest coverage area (e.g., the largest overlap area). For example, equation (4) can be used:
[0081] MV' = N·MV l l = argmax i∈{0,1,...K-1} S i Equation (4)
[0082] Where MV' represents the SPU u as the result. h MV, MV i This represents the i-th corresponding SPU b in the base layer. i The MV is the sum of the base layer's PU and the upsampling factor (e.g., N can be equal to 2 or 1.5), depending on the spatial ratio (e.g., spatial resolution) between the two layers (e.g., the base layer and the enhancement layer). For example, the upsampling factor (e.g., N) can be used to scale the MV of the PU in the processed base layer image, which is determined by the PU of the base layer as a result.
[0083] A weighted average can be used to determine the MV of the SPUs in the processed base layer. For example, a weighted average can be used to determine the MV of the SPUs in the processed base layer by utilizing the MV associated with the corresponding SPUs in the base layer. Using a weighted average can improve the accuracy of the MV of the processed base layer. For the SPU u in the processed base layer... h This can be achieved by determining for one or more (e.g., each) with u h Overlapping basic base layer SPU b i The weighted average is used to obtain the SPU u h MV. For example, it is shown by equation (5):
[0084]
[0085] Where B is the reference image index in the base layer, equal to r(u h ) of SPU b i A subset thereof, for example, is determined by equation (2) and / or equation (3).
[0086] One or more filters (e.g., median filter, low-pass Gaussian filter, etc.) can be applied to the MV group represented by B in equation (5), for example, to obtain the MV represented by MV'. A reliable mean can be used to improve the accuracy of the evaluated MV, as shown in equation (6):
[0087]
[0088] Where the parameter w i Is it an estimate of SPU uh Time base layer SPU b i (For example, each base layer SPU b) i A reliable measurement of MV. Different metrics can be used to obtain w. i The value of w can be determined, for example, based on the amount of predicted residual during motion compensation prediction. i It can also be based on the MV i w is determined by its coherence with neighboring MVs. i .
[0089] Motion information in the processed base layer image can be mapped from the original motion domain of the base layer; for example, it can be used to perform temporal motion compensation prediction in the base layer. Motion domain compensation algorithms (e.g., those supported in HEVC) can be applied to the motion domain of the base layer, for example, to produce a compressed motion domain of the base layer. Motion information in one or more of the processed base layer images can be mapped from the compressed motion domain of the base layer.
[0090] It can generate motion information lost in the processed base layer image, for example, as described herein. No additional modifications to block-level operations are required; TMVP supported by a single-layer codec (e.g., HEVC codec) can be used for the enhancement layer.
[0091] When the corresponding base layer reference image includes one or more slices, a reference image list generation process and / or MV mapping process can be used, for example, as shown here. If multiple slices exist in the base layer reference image, slice partitioning can be mapped from the base layer image to the processed base layer image. For slices in the processed base layer, a reference image list generation step can be performed to obtain the appropriate slice type and / or reference image list.
[0092] Figure 7A -C is a graph illustrating an example relationship between slices of the base layer image and slices of the processed base layer image, for example, for 1.5x spatial scalability. Figure 7A Figure 701 shows an example of slice division in the base layer. Figure 7B Figure 702 shows an example of mapped slice partitioning in the processed base layer. Figure 7C Figure 703 shows an example of the adjusted slice division in the processed base layer.
[0093] The base layer image may include multiple slices, for example, two slices as shown in Figure 701. The slice division mapped in the processed base layer may cross the boundaries of adjacent coding tree blocks (CTBs) in the enhancement layer, for example, when the base layer is upsampled (e.g., as shown in Figure 702). This is because the spatial ratio between the base layer image and the enhancement layer image is different. The slice division (e.g., in HEVC) may be aligned with the CTB boundaries. The slice division in the processed base layer may be adjusted so that the slice boundaries are aligned with the CTB boundaries, for example, as shown in Figure 703.
[0094] The enhancement layer TMVP export process may include constraints. For example, if there is one slice in the corresponding base layer image, the processed base layer image can be used as the juxtaposed image. When there is more than one slice in the corresponding base layer image, interlayer motion information mapping (e.g., reference image list generation and / or MV mapping as described herein) may not be performed on the processed base layer reference image. If there is more than one slice in the corresponding base layer image, a temporal reference image can be used as the juxtaposed image for the enhancement layer TMVP export process. The number of slices in the base layer image can be used to determine whether to use an interlayer reference image and / or a temporal reference image as the juxtaposed image for the enhancement layer TMVP.
[0095] If a slice and / or slice information (e.g., the list of reference images, slice type, etc. in the corresponding base layer image) are identical in the corresponding base layer image, the processed base layer image can be used as the juxtaposed image. When two or more slices in the corresponding base layer image have different slice information, inter-layer motion information mapping (e.g., reference image list generation and / or MV mapping as described herein) is not performed on the processed base layer reference image. If two or more slices in the corresponding base layer image have different slice information, the temporal reference image can be used as the juxtaposed image for the TMVP export process of the enhancement layer.
[0096] Motion information mapping enables the application of various single-layer MV prediction techniques in scalable coding systems. Block-level MV prediction operations can be used to improve the performance of enhancement layer coding. MV prediction for enhancement layers is described here. The MV prediction method for the base layer remains unchanged.
[0097] A temporal MV refers to an MV pointing to a reference image from the same enhancement layer. An interlayer MV refers to an MV pointing to another layer, such as a processed base layer reference image. A mapped MV refers to an MV generated for a processed base layer image. Mapped MVs can include mapped temporal MVs and / or mapped interlayer MVs. A mapped temporal MV refers to a mapped MV derived from the temporal prediction of the last coding layer. A mapped interlayer MV refers to an MV generated based on the interlayer prediction of the last coding layer. Mapped interlayer MVs may exist for scalable coding systems with more than two layers. Temporal MVs and / or mapped temporal MVs can be short-term or long-term MVs, for example, depending on whether the MV points to a short-term or long-term reference image. Temporal short-term MVs and mapped short-term MVs refer to temporal MVs and mapped temporal MVs using short-term temporal references in their respective coding layers. Temporal long-term MVs and mapped long-term MVs refer to temporal MVs and mapped temporal MVs using long-term temporal references in their respective coding layers. Temporal MVs, mapped temporal MVs, mapped interlayer MVs, and interlayer MVs can be considered as different types of MVs.
[0098] Enhanced layer MV prediction may include one or more of the following: MV prediction of temporal MV based on inter-layer MV and / or mapped inter-layer MV may be enabled or disabled. MV prediction of inter-layer MV based on temporal MV and / or mapped temporal MV may be enabled or disabled. MV prediction of temporal MV based on mapped temporal MV may be enabled. MV prediction of inter-layer MV based on inter-layer MV and / or mapped inter-layer MV may be enabled or disabled. For long-term MVs included in the MV prediction, such as long-term temporal MVs and mapped long-term MVs, MV prediction without MV scaling may be used.
[0099] Predictions between short-term MVs using MV scaling (e.g., similar to single-layer MV prediction) can be enabled. Figure 8A This is a graph showing the MV prediction over a short period of time. Figure 8B This is a graph illustrating the prediction of the time short-term MV based on the mapped short-term MV. In graph 800, the time short-term MV 802 can be predicted based on the time short-term MV 804. In graph 810, the time short-term MV 812 can be predicted based on the mapped short-term MV 814.
[0100] For example, due to the large POC spacing, predictions can be provided over long periods of time without using MV scaling. This is similar to MV prediction in single-layer encoding and decoding. Figure 9A This is a graph illustrating an example of MV prediction over a long period of time. Figure 9BThis is a diagram illustrating an example of MV prediction for a time-long MV based on a mapped time-long MV. In Figure 900, a time-long MV 902 can be predicted based on a time-long MV 904. In Figure 910, a time-long MV 912 can be predicted based on a mapped time-long MV 914.
[0101] For example, because the spacing between the two reference images is large, predictions can be provided between the short-term and long-term MVs without using MV scaling. This is similar to MV prediction with single-layer encoding and decoding. Figure 10A This is a diagram illustrating an example of MV prediction for short-term MV based on long-term MV. Figure 10B This is a diagram illustrating an example of MV prediction for the short-term MV over time based on the long-term MV of the mapping. Figure 10C This is a diagram illustrating an example of MV prediction for the long-term MV based on the short-term MV. Figure 10D This is a diagram illustrating an example of MV prediction for the long-term MV based on the short-term MV of the mapping.
[0102] In Figure 1000, the short-term MV 1002 can be predicted based on the long-term MV 1004. In Figure 1010, the short-term MV 1012 can be predicted based on the mapped long-term MV 1014. In Figure 1020, the long-term MV 1024 can be predicted based on the short-term MV 1022. In Figure 1030, the long-term MV 1032 can be predicted based on the mapped short-term MV 1034.
[0103] Prediction of the short-term temporal MV based on the inter-layer MV and / or the mapped inter-layer MV may be disabled. Figure 11A This is a diagram illustrating an example of disabled MV prediction based on inter-layer MV for short-term MV over time. Figure 11B This is a diagram illustrating an example of disabled MV prediction for inter-layer MV based on short-term MV over time. Figure 11C This is a diagram illustrating an example of disabled MV prediction for inter-layer MV based on the short-term MV of the mapping.
[0104] Figure 1100 illustrates an example of disabled MV prediction for the short-term temporal MV 1102 based on the inter-layer MV 1104. For example, the short-term temporal MV 1102 may not be predictable based on the inter-layer MV 1104. Figure 1110 illustrates an example of disabled MV prediction for the inter-layer MV 1112 based on the short-term temporal MV 1114. For example, the inter-layer MV 1112 may not be predictable based on the short-term temporal MV 1114. Figure 1120 illustrates an example of disabled MV prediction for the inter-layer MV 1122 based on the mapped short-term MV 1124. For example, the inter-layer MV 1122 may not be predictable based on the mapped short-term MV 1124.
[0105] Prediction of the long-term MV based on the inter-layer MV and / or the mapped inter-layer MV may be disabled. Figure 12A This is a diagram illustrating an example of disabled MV prediction based on inter-layer MV to long-term MV. Figure 12B This is a diagram illustrating an example of disabled MV prediction for inter-layer MV based on long-term MV over time. Figure 12C This is a diagram illustrating an example of disabled MV prediction for inter-layer MV based on the long-term MV of the mapping.
[0106] Figure 1200 illustrates an example of disabled MV prediction for the long-term MV 1202 based on the inter-layer MV 1204. For example, the long-term MV 1202 may not be predictable based on the inter-layer MV 1204. Figure 1210 illustrates an example of disabled MV prediction for the inter-layer MV 1212 based on the long-term MV 1214. For example, the inter-layer MV 1212 may not be predictable based on the long-term MV 1214. Figure 1220 illustrates disabled MV mapping for the inter-layer MV 1222 based on the mapped long-term MV 1224. For example, the inter-layer MV 1222 may not be predictable based on the mapped long-term MV 1224.
[0107] For example, if two inter-layer predictive variables (MVs) have the same time interval in the enhancement layer and the processed base layer, predicting an inter-layer MV based on another inter-layer MV can be enabled. If the two inter-layer MVs have different time intervals in the enhancement layer and the processed base layer, prediction between the two inter-layer MVs may be disabled. This is because the lack of explicit correlation prevents prediction from producing good coding performance.
[0108] Figure 13A This is a diagram illustrating an example of MV prediction between two layers when Te = Tp. Figure 13B This is a diagram illustrating an example of MV prediction between two interlayer MVs disabled when Te ≠ Tp. TMVP can be used as an example (e.g., as shown in the diagram). Figure 13A-B in Figure 1300. In Figure 1300, the current interlayer MV (e.g., MV2) 1302 can be predicted based on another interlayer MV (e.g., MV1) 1304. The time interval between the current image CurrPic and its temporally adjacent image ColPic (e.g., including the juxtaposed PU ColPU) is denoted as T. e The time interval between their respective reference images (e.g., CurrRefPic and ColRefPic) is denoted as T. p CurrPic and ColPic are in the enhancement layer, while CurrRefPic and ColRefPic are in the processed base layer. If T e =T p Then MV1 can be used to predict MV2.
[0109] For example, because POC-based MV scaling may fail, MV scaling for the predicted MV between two interlayer MVs may be disabled. In Figure 1310, the current interlayer MV (e.g., MV2) 1312 may not be predictable based on another interlayer MV (e.g., MV1) 1314, for example, because of the time interval (e.g., T) between the current image CurrPic and its neighboring image ColPic. e The time intervals between the images and their respective reference images are not equal (e.g., T). p ).
[0110] For example, if the inter-layer MV and the mapped inter-layer MV have the same time interval, prediction of the inter-layer MV without scaling based on the mapped inter-layer MV can be enabled. If their time intervals are different, prediction of the inter-layer MV based on the mapped inter-layer MV may be disabled.
[0111] Table 1 summarizes examples of different conditions for MV prediction used in enhancement layer coding for SVC.
[0112] Table 1: Example conditions for MV prediction in the enhancement layer of SVC
[0113]
[0114]
[0115] For motion information mapping implementations between different coding layers, MV mapping between layers may be disabled, for example, as described herein. Mapped inter-layer MVs are unavailable for MV prediction in enhancement layers.
[0116] MV prediction, including inter-layer MVs, may be disabled. For enhancement layers, temporal MVs can be predicted based on other temporal MVs (e.g., temporal MVs only). This is equivalent to MV prediction for a single-layer codec.
[0117] Devices (e.g., processors, encoders, decoders, WTRUs, etc.) can receive bitstreams (e.g., scalable bitstreams). For example, the bitstream may include a base layer and one or more enhancement layers. A TMVP can be used to decode the base layer (e.g., base layer video blocks) and / or enhancement layers (e.g., enhancement layer video blocks) in the bitstream. A TMVP can be performed on both the base layer and enhancement layers in the bitstream. For example, a TMVP can be performed on the base layer (e.g., base layer video blocks) in the bitstream without any modifications, for example, refer to... Figure 4 The description states that TMVP can be performed on enhancement layers (e.g., enhancement layer video blocks) in a bitstream using an interlayer reference image, as described herein. For example, the interlayer reference image can be used as a juxtaposition reference image for the TMVP of the enhancement layer (e.g., enhancement layer video blocks). For example, the compressed MV domain of the juxtaposition base layer image can be determined. The MV domain of the interlayer reference image can be determined based on the compressed MV domain of the juxtaposition base layer image. The MV domain of the interlayer reference image can be used to perform TMVP on the enhancement layer (e.g., enhancement layer video blocks). For example, the MV domain of the interlayer reference image can be used to predict the MV domain for the enhancement layer video block (e.g., and enhancement layer video blocks).
[0118] The MV domain of the interlayer reference layer image can be determined. For example, the MV domain of the interlayer reference layer image can be determined based on the MV domain of the juxtaposed base layer image. The MV domain may include one or more MVs and / or reference image indices. For example, the MV domain may include the reference image index and MV of PUs in the interlayer reference layer image (e.g., for each PU in the interlayer reference layer image). The enhancement layer image (e.g., the juxtaposed enhancement layer image) can be decoded based on the MV domain. TMVP can be performed on the enhancement layer image based on the MV domain.
[0119] Syntax signaling (e.g., high-level syntax signaling) for inter-layer motion prediction can be provided. Inter-layer runtime information mapping and MV prediction may be enabled or disabled at the sequence level. Inter-layer runtime information mapping and MV prediction may also be enabled or disabled at the image / slice level. For example, decisions on whether to enable and / or disable certain inter-layer motion prediction techniques may be made based on considerations of improving coding efficiency and / or reducing system complexity. Sequence-level signaling has less overhead than image / slice-level signaling, for example, because the added syntax can be applied to images in the sequence (e.g., all images). Image / slice-level signaling provides greater flexibility, for example, because images in the sequence (e.g., each image) can receive their own motion prediction implementation and / or MV prediction implementation.
[0120] Sequence-level signaling can be provided. Interlayer motion information mapping and / or MV prediction can be sent at the sequence level. If sequence-level signaling is used, the same motion information mapping and / or MV prediction can be used for all images in the sequence (e.g., all images). For example, the syntax shown in Table 2 indicates whether interlayer motion information mapping and / or MV prediction at the sequence level is allowed. The syntax in Table 2 can be used on parameter sets, for example, such as video parameter sets (VPS) (e.g., in HEVC), sequence parameter sets (SPS) (e.g., in H.264 and HEVC), and image parameter sets (PPS) (e.g., in H.264 and HEVC), etc., but is not limited to these.
[0121] Table 2: Examples of syntax for increasing sequence-level signaling
[0122]
[0123]
[0124] The `inter_layer_mvp_present_flag` indicates whether inter-layer motion prediction is used at the sequence level or at the image / slice level. For example, if this flag is set to 0, the signaling is at the image / slice level. If this flag is set to 1, motion mapping and / or MV prediction signaling is at the sequence level. The inter-layer motion mapping sequence enable flag (`inter_layer_motion_mapping_seq_enabled_flag`) indicates whether inter-layer motion mapping (e.g., inter-layer motion prediction) is used at the sequence level. The inter-layer add-mvp_seq_enabled_flag` indicates whether block MV prediction (e.g., additional block MV prediction) is used at the sequence level.
[0125] Image / slice-level signaling can be provided. Interlayer motion information mapping can be sent at the image / slice level. If image / slice-level signaling is used, each image in the sequence (e.g., each image) can receive its own signaling. For example, images in the same sequence can use different motion information mappings and / or MV predictions (e.g., based on the signaling they receive). For example, the syntax in Table 3 can be used in the slice header to indicate whether interlayer motion information mapping and / or MV prediction are used in the current image / slice in the enhancement layer.
[0126] Table 3: Examples of the modified slice header syntax
[0127]
[0128]
[0129] The `inter_layer_motion_mapping_slice_enabled_flag` indicates whether inter-layer motion mapping is applied to the current slice. The `inter_layer_add_mvp_slice_enabled_flag` indicates whether additional block MV prediction is applied to the current slice.
[0130] It is recommended to use MV predictive coding in multi-layer video coding systems. The inter-layer motion information mapping algorithm described herein is used to generate runtime-related information for the processed base layer, for example, enabling the utilization of the temporal correlation between the base and enhancement layers' MVs during the TMVP process in the enhancement layer. Because block-level operations can be maintained without modification, single-layer encoders and decoders can be used for MV prediction in the enhancement layer without alteration. MV prediction can be based on the characteristic analysis of different types of MVs in a scalable system (e.g., to improve MV prediction efficiency).
[0131] Although a two-layer SVC system with spatial scalability is described herein, this disclosure is extensible to SVC systems with more layers and other scalability patterns.
[0132] Interlayer motion prediction can be performed on enhancement layers in a bitstream. Interlayer motion prediction can be signaled, for example, as described herein. Interlayer motion prediction can be signaled at the sequence level of the bitstream (e.g., using inter_layer_motion_mapping_seq_enabled_flag, etc.). For example, interlayer motion prediction can be signaled using variables in the video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), etc., within the bitstream.
[0133] Devices (e.g., processors, encoders, decoders, WTRUs, etc.) may perform any of the functions described herein. For example, an encoder may include a processor configured to receive a bitstream (e.g., a scalable bitstream). The bitstream may include a base layer and an enhancement layer. A decoder may use temporal motion vector prediction (TMVP) in the bitstream, which uses an interlayer reference image as a juxtaposed reference image for the TMVP of the enhancement layer. Enhancement layer video blocks, interlayer video blocks, and / or base layer video blocks may be juxtaposed (e.g., temporally juxtaposed).
[0134] The decoder can employ a TMVP to decode the enhancement layer image. For example, the decoder can determine the MV field of the interlayer reference image based on the MV field of the juxtaposed base layer image. The interlayer reference image and the enhancement layer reference image are juxtaposed. The MV field of the interlayer reference image may include the reference image index and MV of the video block of the interlayer reference image. The decoder can decode the enhancement layer image based on the MV field of the interlayer reference image. For example, the decoder can determine the MV field of the enhancement layer image based on the MV field of the interlayer reference image and decode the enhancement layer image based on the MV field of the enhancement layer image.
[0135] The MV domain of the interlayer reference image can be determined based on the compressed MV domain. For example, the decoder can determine the compressed MV domain of the juxtaposed base layer image and determine the MV domain of the interlayer reference image based on the compressed MV domain of the juxtaposed base layer image.
[0136] The decoder can determine the video block's MV and reference image of the interlayer reference image. For example, the decoder can determine the reference image of the interlayer video block based on the reference image of the juxtaposed base layer video block. The decoder can determine the MV of the interlayer video block based on the MV of the juxtaposed base layer video block. The decoder determines the juxtaposed base layer video block by selecting the video block in the juxtaposed base layer image characterized by the largest overlap area with the interlayer video block. The decoder can determine the MV and / or reference image of the enhancement layer's video block (e.g., the juxtaposed video block of the enhancement layer image) based on the MV and / or reference image of the video block of the interlayer reference image.
[0137] The decoder can determine a reference image for juxtaposed base layer video blocks and, based on the reference image of the juxtaposed base layer video blocks, determine a reference image for interlayer video blocks. For example, the reference image for interlayer video blocks could be a juxtaposed interlayer reference image of the reference image of the juxtaposed base layer video blocks. The decoder can determine a reference image for video blocks of the enhancement layer image based on the reference image of the interlayer video blocks. For example, the reference image for the enhancement layer could be a juxtaposed enhancement layer reference image of the reference image of the interlayer video blocks. Enhancement layer video blocks, interlayer video blocks, and / or base layer video blocks may be juxtaposed (e.g., temporally juxtaposed).
[0138] The decoder can determine the MV of inter-layer video blocks. For example, the decoder can determine the MV of the juxtaposed base layer video blocks and scale the MV of the juxtaposed base layer video blocks according to the spatial ratio between the base and enhancement layers to determine the MV of the inter-layer video blocks. The decoder can determine the MV of the enhancement layer video blocks based on the MV of the inter-layer video blocks. For example, the decoder can use the MV of the inter-layer video blocks to predict the MV of the enhancement layer video blocks, for example, by temporally scaling the MV of the inter-layer video blocks.
[0139] The decoder can be configured to determine a reference image for the enhancement layer video block based on the juxtaposed base layer video block, determine the MV of the enhancement layer video block based on the MV of the juxtaposed base layer video block, and / or decode the enhancement layer video block based on both the reference image and the MV of the enhancement layer video block. For example, the decoder can determine the juxtaposed base layer video block by selecting a video block in the juxtaposed base layer image characterized by the largest overlap area with the enhancement layer video block.
[0140] The decoder can determine a reference image for juxtaposed base layer video blocks. The decoder can use the reference image of the juxtaposed base layer video blocks to determine a reference image for interlayer video blocks. The decoder can determine a reference image for enhancement layer video blocks. For example, the reference image for an enhancement layer video block can be a juxtaposed enhancement layer image of the reference images of the juxtaposed base layer video blocks and the reference images of the juxtaposed interlayer video blocks. Enhancement layer video blocks, interlayer video blocks, and / or base layer video blocks can be juxtaposed (e.g., temporally juxtaposed).
[0141] The decoder determines the MV of the juxtaposed base layer video blocks. The decoder can scale the MV of the juxtaposed base layer video blocks according to the spatial ratio between the base and enhancement layers to determine the MV of the interlayer video blocks. The decoder can predict the MV of the enhancement layer video blocks based on the MV of the interlayer video blocks, for example, by temporal scaling of the MV of the interlayer video blocks.
[0142] The decoder may include a processor capable of receiving a bitstream. The bitstream may include a base layer and an enhancement layer. The bitstream may include inter-layer motion mapping information. The decoder may determine whether inter-layer motion prediction for the enhancement layer can be enabled based on the inter-layer mapping information. The decoder may perform inter-layer motion prediction for the enhancement layer based on the inter-layer mapping information. The inter-layer mapping information may be signaled at the sequence level of the bitstream. For example, the inter-layer mapping information may be signaled via variables (e.g., flags) in the VPS, SPS, and / or PPS of the bitstream.
[0143] Although described from the decoder's perspective, the functions described herein (e.g., the reverse functions of the functions described herein) can be performed by other devices, such as encoders.
[0144] Figure 14AThis is a system diagram of an example communication system 1400 in which one or more implementation methods can be carried out. Communication system 1400 can be a multi-access system that provides content, such as voice, data, video, messaging, broadcasting, etc., to multiple users. Communication system 1400 allows multiple wireless users to access this content through system resource sharing (including wireless bandwidth). For example, communication system 1400 can use one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FMDMA (SC-FDMA), etc.
[0145] like Figure 14A As shown, the communication system 1400 may include wireless transmit / receive units (WTRUs) 1402a, 1402b, 1402c and / or 1402d (commonly referred to collectively as WTRU 1402), radio access networks (RANs) 1403 / 1404 / 1405, core networks 1406 / 1407 / 1407, public switched telephone network (PSTN) 1408, the Internet 1410, and other networks 1412. However, it will be understood that the disclosed implementations accommodate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 1402a, 1402b, 1402c, and 1402d can be any type of device configured to operate and / or communicate in a wireless environment. As an example, WTRUs 1402a, 1402b, 1402c, and 1402d can be configured to transmit and / or receive wireless signals and may include user equipment (UE), base stations, fixed or mobile user units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, consumer electronics, and so on.
[0146] The communication system 1400 may also include base stations 1414a and 1414b. Each of base stations 1414a and 1414b may be any type of device configured to wirelessly interface with at least one of WTRUs 1402a, 1402b, 1402c, and 1402d to facilitate access to one or more communication networks, such as core networks 1406 / 1407, 1409, the Internet 1410, and / or network 1412. As an example, base stations 1414a and 1414b may be base transceiver stations (BTS), Node Bs, evolved Node Bs (eNode Bs), home Node Bs, home eNBs, site controllers, access points (APs), wireless routers, etc. Although each of base stations 1414a and 1414b is described as a separate element, it will be understood that base stations 1414a and 1414b may include any number of interconnected base station and / or network elements.
[0147] Base station 1414a may be part of RAN 1403 / 1404 / 1405, which may also include other base stations and / or network elements (not shown), such as Base Station Controller (BSC), Radio Network Controller (RNC), relay nodes, etc. Base station 1414a and / or base station 1414b may be configured to transmit and / or receive radio signals within a specific geographical area, which may be referred to as a cell (not shown). The cell may also be divided into cell sectors. For example, the cell associated with base station 1414a may be divided into three sectors. Therefore, in one embodiment, base station 1414a may include three transceivers, each for one sector of the cell. In another embodiment, base station 1414a may use Multiple-Input Multiple-Output (MIMO) technology, thus allowing multiple transceivers to be used for each sector of the cell.
[0148] Base stations 1414a and 1414b can communicate with one or more of WTRUs 1402a, 1402b, 1402c, and 1402d via air interfaces 1415 / 1416 / 1417, which can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, infrared (IR), ultraviolet (UV), visible light, etc.). Any suitable radio access technology (RAT) can be used to establish air interfaces 1415 / 1416 / 1417.
[0149] More specifically, as described above, the communication system 1400 can be a multi-access system and can use one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base stations 1414a and WTRUs 1402a, 1402b, and 1402c in RAN 1403 / 1404 / 1405 can use radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which can use Wideband CDMA (WCDMA) to establish air interfaces 1415 / 1416 / 1417. WCDMA can include communication protocols such as High-Speed Packet Access (HSPA) and / or evolved HSPA (HSPA+). HSPA can include High-Speed Downlink Packet Access (HSDPA) and / or High-Speed Uplink Packet Access (HSUPA).
[0150] In another implementation, base stations 1414a and WTRUs 1402a, 1402b, 1402c may use, for example, evolved UMTS terrestrial radio access (E-UTRA) radio technology, which may use Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) to establish air interfaces 1415 / 1416 / 1417.
[0151] In other implementations, base station 1414a and WTRUs 1402a, 1402b, 1402c may use radio technologies such as IEEE 802.16 (i.e., Global Microwave Access Interoperability (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate Evolution of GSM (EDGE), GSM EDGE (GERAN), etc.
[0152] Figure 14A Base station 1414b can be a wireless router, home node B, home e node B, or access point, and can use any suitable RAT to facilitate wireless connectivity in a local area, such as a commercial location, residence, vehicle, campus, etc. In one implementation, base station 1414b and WTRUs 1402c, 1402d can implement radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In another implementation, base station 1414b and WTRUs 1402c, 1402d can use radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another implementation, base station 1414b and WTRUs 1402c, 1402d can use cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, etc.) to establish picocells or femtocells. Figure 14A As shown, base station 1414b can have a direct connection to the Internet 1410. Therefore, base station 1414b can access the Internet 1410 without going through core networks 1406 / 1407 / 1408.
[0153] RAN 1403 / 1404 / 1405 can communicate with core networks 1406 / 1407 / 1409, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more WTRUs 1402a, 1402b, 1402c, and 1402d. For example, core networks 1406 / 1407 / 1409 can provide call control, billing services, location-based services, prepaid calling, internet connectivity, video distribution, and / or perform advanced security functions such as user authentication. Although Figure 14A As not shown, it will be understood that RAN 1403 / 1404 / 1405 and / or core network 1406 / 1407 / 1409 can communicate directly or indirectly with other RANs using the same RAT as or a different RAT than RAN 1403 / 1404 / 1405. For example, in addition to being connected to RAN 1403 / 1404 / 1405 which is using E-UTRA radio technology, core network 1406 / 1407 / 1409 can also communicate with another RAN (not shown) using GSM radio technology.
[0154] Core networks 1406 / 1407 / 1409 can also act as gateways for WTRUs 1402a, 1402b, 1402c, and 1402d to access PSTN 1408, the Internet 1410, and / or other networks 1412. PSTN 1408 may include a circuit-switched telephone network providing Common Old-Style Telephone Service (POTS). The Internet 1410 may include a global system of interconnected computer networks and devices using common communication protocols, such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and Internet Protocol (IP) from the TCP / IP Internet Protocol suite. Network 1412 may include wired or wireless communication networks owned and / or operated by other service providers. For example, network 1412 may include another core network connected to one or more RANs, which may use the same RAT as RAN 1403 / 1404 / 1405 or a different RAT.
[0155] Some or all of the WTRUs 1402a, 1402b, 1402c, and 1402d in the communication system 1400 may include multi-mode capability, meaning that the WTRUs 1402a, 1402b, 1402c, and 1402d may include multiple transceivers for communicating with different wireless networks on different wireless links. For example, Figure 14AThe WTRU 1402c shown can be configured to communicate with base station 1414a, which can use cellular-based radio technology, and with base station 1414b, which can use IEEE 802 radio technology.
[0156] Figure 14B This is a system diagram of the WTRU 1402 example. (Example:) Figure 14B As shown, WTRU 1402 may include a processor 1418, a transceiver 1420, a transmit / receive element 1422, a speaker / microphone 1424, a keyboard 1426, a display / touchpad 1428, non-removable memory 1430, removable memory 1432, a power supply 1434, a Global Positioning System (GPS) chipset 1436, and other peripheral devices 1438. It will be understood that WTRU 1402 may include any sub-combination of the foregoing elements while remaining consistent with the implementation. Similarly, the nodes represented by base stations 1414a and 1414b and / or base stations 1414a and 1414b, which are of interest in the implementation, such as but not limited to base transceiver stations (BTS), node B, site controllers, access points (APs), home node B, evolved home node B (eNodeB), home evolved node B (HeNB), gateways and proxy nodes of home evolved node B, may include... Figure 14B And some or all of the elements described herein.
[0157] Processor 1418 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. Processor 1418 may perform signal encoding, data processing, power control, input / output processing, and / or any other function that enables WTRU 1402 to operate in a wireless environment. Processor 1418 may be coupled to transceiver 1420, which may be coupled to transmitting / receiving element 1422. Although Figure 14B The processor 1418 and transceiver 1420 are described as separate components, but it will be understood that the processor 1418 and transceiver 1420 can be integrated together in an electronic package or chip.
[0158] Transmitting / receiving element 1422 can be configured to transmit signals to or receive signals from a base station (e.g., base station 1414a) via air interfaces 1415 / 1416 / 1417. For example, in one embodiment, transmitting / receiving element 1422 can be an antenna configured to transmit and / or receive RF signals. In another embodiment, transmitting / receiving element 1422 can be a transmitter / detector configured to transmit and / or receive signals such as IR, UV, or visible light. In yet another embodiment, transmitting / receiving element 1422 can be configured to transmit and receive both RF and optical signals. It should be understood that transmitting / receiving element 1422 can be configured to transmit and / or receive any combination of wireless signals.
[0159] Additionally, although the transmitting / receiving element 1422 is in Figure 14B While described as separate components, WTRU 1402 may include any number of transmit / receive elements 1422. More specifically, WTRU 1402 may use, for example, MIMO technology. Thus, in one embodiment, WTRU 1402 may include two or more transmit / receive elements 1422 (e.g., multiple antennas) for transmitting and receiving wireless signals via air interfaces 1415 / 1416 / 1417.
[0160] Transceiver 1420 can be configured to modulate signals to be transmitted by transmitting / receiving element 1422 and / or demodulate signals received by transmitting / receiving element 1422. As described above, WTRU 1402 can have multi-mode capability. Therefore, transceiver 1420 can include multiple transceivers enabling WTRU 1402 to communicate via multiple RATs, such as UTRA and IEEE 802.11.
[0161] The processor 1418 of WTRU 1402 can be coupled to and receive user input data from devices such as a speaker / microphone 1424, a keyboard 1426, and / or a display / touchpad 1428 (e.g., a liquid crystal display (LCD) unit or an organic light-emitting diode (OLED) display unit). The processor 1418 can also output user data to the speaker / microphone 1424, keyboard 1426, and / or display / touchpad 1428. Additionally, the processor 1418 can access information from any type of suitable memory and can store data in any type of suitable memory, such as non-removable memory 1430 and / or removable memory 1432. Non-removable memory 1430 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory device. Removable memory 1432 may include a Subscriber Identity Module (SIM) card, a Memory Stick, a Secure Digital (SD) memory card, etc. In other embodiments, the processor 1418 may access information from memory that is not physically located on the WTRU 1402, such as on a server or home computer (not shown), and may store the data in that memory.
[0162] The processor 1418 can receive electrical energy from the power supply 1434 and can be configured to distribute and / or control the electrical energy to other components in the WTRU 1402. The power supply 1434 can be any suitable device that powers the WTRU 1402. For example, the power supply 1434 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.
[0163] The processor 1418 may also be coupled to a GPS chipset 1436, which may be configured to provide location information (e.g., longitude and latitude) about the current location of the WTRU 1402. Additionally, in addition to information from the GPS chipset 1436 or as an alternative, the WTRU 1402 may receive location information from base stations (e.g., base stations 1414a, 1414b) via air interfaces 1415 / 1416 / 1417 and / or determine its location based on the timing of signals received from two or more neighboring base stations. It will be understood that the WTRU 1402 may obtain location information using any suitable location determination method while maintaining consistency in implementation.
[0164] The processor 1418 can be coupled to other peripheral devices 1438, which may include one or more software and / or hardware modules providing additional features, functions, and / or wired or wireless connectivity. For example, peripheral devices 1438 may include accelerometers, electronic compasses, satellite transceivers, digital cameras (for photos or videos), Universal Serial Bus (USB) ports, vibration devices, television transceivers, hands-free headsets, and Bluetooth devices. Modules, FM radio units, digital music players, media players, video game console modules, internet browsers, etc.
[0165] Figure 14C This is a structural diagram of RAN 1403 and core network 1406 according to an implementation method. As described above, for example, RAN 1403 can communicate with WTRUs 1402a, 1402b, and 1402c via air interface 1415 using UTRA radio technology. RAN 1403 can also communicate with core network 1406. Figure 14C As shown, RAN 1403 may include Node Bs 1440a, 1440b, and 1440c, each of which includes one or more transceivers for communicating with WTRUs 1402a, 1402b, and 1402c via air interface 1415. Each of Node Bs 1440a, 1440b, and 1440c may be associated with a specific cell (not shown) within RAN 1403. RAN 1403 may also include RNCs 1442a and 1442b. It will be understood that RAN 1403 may include any number of Node Bs and RNCs while maintaining consistency in implementation.
[0166] like Figure 14C As shown, nodes B 1440a and 1440b can communicate with RNC 1442a. Additionally, node B 1440c can communicate with RNC 1442b. Nodes B 1440a, 1440b, and 1440c can communicate with RNCs 1442a and 1442b respectively via the Iub interface. RNCs 1442a and 1442b can communicate with each other via the Iur interface. Each of RNCs 1442a and 1442b can be configured to control the individual nodes B 1440a, 1440b, and 1440c connected to it. Furthermore, each of RNCs 1442a and 1442b can be configured to perform or support other functions, such as outer-loop power control, load control, admission control, packet scheduling, handover control, macro diversity, security functions, data encryption, etc.
[0167] Figure 14CThe core network 1406 shown may include a media gateway (MGW) 1444, a mobile switching center (MSC) 1446, a serving GPRS support node (SGSN) 1448, and / or a gateway GPRS support node (GGSN) 1450. Although each of the foregoing elements is described as part of the core network 1406, it will be understood that any of these elements may be owned or operated by an entity that is not a core network operator.
[0168] RNC 1442a in RAN 1403 can connect to MSC 1446 in core network 1406 via the IuCS interface. MSC 1446 can connect to MGW 1444. MSC 1446 and MGW 1444 can provide WTRU 1402a, 1402b, and 1402c with access to circuit-switched networks such as PSTN 1408, facilitating communication between WTRU 1402a, 1402b, and 1402c and traditional terrestrial line communication equipment.
[0169] In RAN 1403, RNC 1442a can also connect to SGSN 1448 in core network 1406 via the IuPS interface. SGSN 1448 can connect to GGSN 1450. SGSN 1448 and GGSN 1450 can provide WTRUs 1402a, 1402b, and 1402c with access to packet-switched networks such as the Internet 1410, facilitating communication between WTRUs 1402a, 1402b, 1402c and IP-enabled devices.
[0170] As described above, core network 1406 can also be connected to network 1412, which may include other wired or wireless networks owned or operated by other service providers.
[0171] Figure 14D This is a structural diagram of RAN 1404 and core network 1407 according to an implementation method. As described above, for example, RAN 1404 can communicate with WTRUs 1402a, 1402b, and 1402c via air interface 1416 using E-UTRA radio technology. RAN 1404 can also communicate with core network 1407.
[0172] RAN 1404 may include eNodeBs 1460a, 1460b, and 1460c, but it will be understood that RAN 1404 may include any number of eNodeBs while maintaining consistency with various implementations. Each of eNodeBs 1460a, 1460b, and 1460c may include one or more transceivers for communicating with WTRUs 1402a, 1402b, and 1402c via air interface 1416. In one implementation, eNodeBs 1460a, 1460b, and 1460c may implement MIMO technology. Thus, for example, eNodeB 1460a may use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 1402a.
[0173] Each of the eNodeB 1460a, 1460b, and 1460c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in the uplink and / or downlink, etc. Figure 14D As shown, nodes B1460a, 1460b, and 1460c can communicate with each other via the X2 interface.
[0174] Figure 14D The core network 1407 shown may include a Mobility Management Entity (MME) 1462, a Serving Gateway 1464, and / or a Packet Data Network (PDN) Gateway 1466. While each of the foregoing units is described as part of the core network 1407, it will be understood that any one of these units may be owned and / or operated by an entity other than the core network operator.
[0175] The MME 1462 can connect to each of the eNodeBs 1460a, 1460b, and 1460c in RAN 1404 via the S1 interface and can act as a control node. For example, the MME 1462 can handle user authentication, bearer activation / deactivation, and selection of a specific serving gateway during the initial attachment of WTRUs 1402a, 1402b, and 1402c, etc. The MME 1462 can also provide control plane functions for handover between RAN 1404 and other RANs (not shown) using other radio technologies such as GSM or WCDMA.
[0176] Service gateway 1464 can connect to each of eNodeBs 1460a, 1460b, and 1460c in RAN 104b via the S1 interface. Service gateway 1464 can typically route and forward user data packets to / from WTRUs 1402a, 1402b, and 1402c. Service gateway 1464 can also perform other functions, such as anchoring the user plane during eNB handover, triggering paging when downlink data is available for WTRUs 1402a, 1402b, and 1402c, managing and storing the context of WTRUs 1402a, 1402b, and 1402c, etc.
[0177] Service gateway 1464 can also be connected to PDN gateway 1466, which can provide WTRUs 1402a, 1402b, and 1402c with access to a packet-switched network (e.g., the Internet 1410) to facilitate communication between WTRUs 1402a, 1402b, and 1402c and IP-enabled devices.
[0178] Core network 1407 facilitates communication with other networks. For example, core network 1406 can provide WTRUs 1402a, 1402b, and 1402c with access to a circuit-switched network (e.g., PSTN 1408) to facilitate communication between WTRUs 1402a, 1402b, and 1402c and traditional terrestrial line communication equipment. Core network 1407 may include, or communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server), which acts as an interface between core network 1407 and PSTN 1408. Additionally, core network 1407 can provide WTRUs 1402a, 1402b, and 1402c with access to network 1412, which may include other wired or wireless networks owned and / or operated by other service providers.
[0179] Figure 14E This is a structural diagram of RAN 1405 and core network 1409 according to an implementation method. RAN 1405 may be an access service network (ASN) that communicates with WTRUs 1402a, 1402b, and 1402c via air interface 1417 using IEEE 802.16 radio technology. As discussed further below, links between different functional entities of WTRUs 1402a, 1402b, 1402c, RAN 1405, and core network 1409 can be defined as reference points.
[0180] like Figure 14EAs shown, RAN 1405 may include base stations 1480a, 1480b, 1480c and ASN gateway 1482, but it will be understood that RAN 1405 may include any number of base stations and ASN gateways as consistent with the implementation. Each of base stations 1480a, 1480b, and 1480c may be associated with a specific cell (not shown) in RAN 1405 and may include one or more transceivers communicating with WTRUs 1402a, 1402b, and 1402c via air interface 11417. In one example, base stations 1480a, 1480b, and 1480c may implement MIMO technology. Therefore, for example, base station 1480a may use multiple antennas to transmit or receive radio signals from WTRU 1402a. Base stations 1480a, 1480b, and 1480c can provide mobility management functions, such as call handoff triggering, tunnel establishment, radio resource management, service classification, and quality of service policy enforcement. ASN gateway 1482 can act as a service aggregation point and is responsible for paging, caching user profiles, routing to core network 1409, etc.
[0181] The air interface 1417 between WTRUs 1402a, 1402b, 1402c and RAN 1405 can be defined as an R1 reference point implementing the IEEE 802.16 specification. Additionally, each of WTRUs 1402a, 1402b, and 1402c can establish a logical interface (not shown) with the core network 1409. The logical interface between WTRUs 1402a, 1402b, 1402c and the core network 1409 can be defined as an R2 reference point, which can be used for authentication, authorization, IP host configuration management, and / or mobility management.
[0182] The communication link between each of base stations 1480a, 1480b, and 1480c can be defined as an R8 reference point, including protocols to facilitate WTRU handover and inter-base station data transfer. The communication link between base stations 1480a, 1480b, and 1480c and ASN gateway 1482 can be defined as an R6 reference point. The R6 reference point may include protocols to facilitate mobility management based on mobility events associated with each of WTRUs 1402a, 1402b, and 1402c.
[0183] like Figure 14EAs shown, RAN 1405 can connect to core network 1409. The communication link between RAN 1405 and core network 1409 can be defined as an R3 reference point including protocols such as those facilitating data transfer and mobility management capabilities. Core network 1409 may include a Mobile IP Local Agent (MIP-HA) 1484, an Authentication, Authorization, Accounting (AAA) server 1486, and a gateway 1488. Although each of the foregoing elements is described as part of core network 1409, it will be understood that any of these elements may be owned or operated by an entity that is not the core network operator.
[0184] MIP-HA manages IP addresses and enables WTRU 1402a, 1402b, and 1402c to roam between different ASNs and / or core networks. MIP-HA 1484 provides WTRU 1402a, 1402b, and 1402c with access to packet-switched networks (e.g., Internet 1410) to facilitate communication between WTRU 1402a, 1402b, 1402c and IP-enabled devices. AAA server 1486 handles user authentication and user service support. Gateway 1488 facilitates interoperability with other networks. For example, gateway 1488 provides WTRU 1402a, 1402b, and 1402c with access to circuit-switched networks (e.g., PSTN 1408) to facilitate communication between WTRU 1402a, 1402b, 1402c and legacy terrestrial communication equipment. In addition, gateway 1488 can provide network 1412 to WTRUs 1402a, 1402b, and 1402c, which may include other wired or wireless networks owned or operated by other service providers.
[0185] Although not in Figure 14E As shown, it will be understood that RAN 1405 can connect to other ASNs, and core network 1409 can connect to other core networks. The communication link between RAN 1405 and other ASNs can be defined as an R4 reference point, which may include protocols coordinating the mobility of WTRUs 1402a, 1402b, and 1402c between RAN 1405 and other ASNs. The communication link between core network 1409 and other core networks can be defined as an R5 reference, which may include protocols facilitating interoperability between the local core network and the accessed core network.
[0186] Figure 15This is a block diagram illustrating a block-based video encoder (e.g., a hybrid video encoder). The input video signal 1502 can be processed block by block. A video block unit comprises 16x16 pixels. Such a block unit can be referred to as a macroblock (MB). In High-Efficiency Video Coding (HEVC), extended block sizes (e.g., referred to as "coding units" or CUs) can be used to efficiently compress high-resolution (e.g., greater than or equal to 1080p) video signals. In HEVC, CUs can be up to 64x64 pixels. CUs can be divided into prediction units (PUs), and individual prediction methods can be employed for each prediction unit.
[0187] For an input video block (e.g., MB or CU), spatial prediction 1560 and / or temporal prediction 1562 can be performed. Spatial prediction (e.g., "internal prediction") uses pixels from encoded neighboring blocks in the same video image / slice to predict the current video block. Spatial prediction reduces inherent spatial redundancy in the video signal. Temporal prediction (e.g., "inter-prediction" or "motion-compensated prediction") uses pixels from an encoded video image (e.g., which may be referred to as a "reference image") to predict the current video block. Temporal prediction reduces inherent temporal redundancy in the signal. The temporal prediction for a video block can be signaled via one or more motion vectors, which can be used to indicate the amount and / or direction of motion between its predicted block and the current block in the reference image. If multiple reference images are supported (e.g., in the case of H.264 / AVC and / or HEVC), an additional reference image index is sent for each video block. The reference image index can be used to identify which reference image in the reference image memory 1564 (e.g., which may be referred to as a "decoded image buffer" or DPB) the temporal prediction signal originates from.
[0188] Following spatial and / or temporal prediction, mode decision block 1580 in the encoder selects a prediction mode. The prediction block is subtracted from the current video block 1516. The prediction residual is transformed 1504 and / or quantized by 1506. The quantized residual coefficients are inversely quantized 1510 and / or inverse transformed 1512 to form the reconstructed residual, which is then added back to the prediction block 1526 to form the reconstructed video block.
[0189] Before the reconstructed video block is placed in the reference image memory 1564 and / or used for encoding subsequent video blocks, closed-loop filtering 1566, such as but not limited to deblocking filters, sampling adaptive offsets, and / or adaptive loop filters, can be applied to the reconstructed video block. To form the output video bitstream 1520, the encoding mode (e.g., inter-prediction mode or internal prediction mode), prediction mode information, motion information, and / or quantized residual coefficients are sent to the entropy coding unit 1508 for compression and / or packing to form the bitstream.
[0190] Figure 16 This is a diagram illustrating an example of a block-based video decoder. The video bitstream 1602 is broken down and / or entropy-decoded in the entropy decoding unit 1608. To form prediction blocks, the coding mode and / or prediction information may be sent to the spatial prediction unit 1660 (e.g., if it is inter-coded) and / or the temporal prediction unit 1662 (e.g., if it is inter-coded). If it is inter-coded, the prediction information may include the prediction block size, one or more motion vectors (e.g., which may be used to indicate the direction and amount of motion), and / or one or more reference indices (e.g., which may be used to indicate from which reference image the prediction signal is derived).
[0191] The timing prediction unit 1662 can apply motion-compensated prediction to form a timing prediction block. Residual conversion coefficients can be sent to the inverse quantization unit 1610 and the inverse conversion unit 1612 to reconstruct the residual block. The predicted block and the residual block are added together in 1626. The reconstructed block undergoes closed-loop filtering before being stored in the reference image storage 1664. The reconstructed video in the reference image storage 1664 can be used to drive a display device and / or to predict subsequent video blocks.
[0192] A single-layer video encoder can take a single video sequence as input and generate a single compressed bitstream that is passed to a single-layer decoder. Video codecs can be designed for digital video services (e.g., but not limited to, transmitting TV signals via satellite, cable, and terrestrial transmission channels). With the development of video-centric applications in heterogeneous environments, multi-layer video coding techniques can be developed as extensions to video standards to enable a variety of applications. For example, scalable video coding techniques are designed to handle cases with more than one video layer, where each layer can be decoded to reconstruct a video signal with specific spatial resolution, temporal resolution, fidelity, and / or view. (See reference...) Figure 15 and Figure 16 Single-layer encoders and decoders are described, and the concepts described herein also utilize multi-layer encoders and decoders, for example, for multi-layer or scalable coding techniques. Figure 15 encoder and / or Figure 16 The decoder can perform any of the functions described herein. For example, Figure 15 encoder and / or Figure 16 The decoder can use the MV of the enhancement layer PU to perform TMVP on the enhancement layer (e.g., the enhancement layer image).
[0193] Figure 17 This diagram illustrates an example of a communication system. The communication system 1700 may include an encoder 1702, a communication network 1704, and a decoder 1706. The encoder 1702 can communicate with the communication network 1704 via a connection 1708. The connection 1708 can be a wired or wireless connection. The encoder 1702 is similar to... Figure 15 A block-based video encoder. Encoder 1702 may include a single-layer codec (e.g., such as...). Figure 15 (as shown) or a multi-layer codec.
[0194] Decoder 1706 can communicate with communication network 1704 via connection 1710. Connection 1710 can be a wired or wireless connection. Decoder 1706 is similar to... Figure 16 The video decoder is a block-based decoder. Decoder 1706 may include a single-layer codec (e.g., such as...). Figure 16 (as shown) or a multi-layer codec. Encoder 1702 and / or decoder 1706 can be incorporated into any of a wide variety of wired communication devices and / or wireless transmit / receive units (WTRUs), such as, but not limited to, digital televisions, wireless broadcasting systems, network elements / terminals, servers (e.g., content or website servers, such as Hypertext Transfer Protocol (HTTP) servers), personal digital assistants (PDAs), laptops or desktop computers, tablets, digital cameras, digital recording devices, video game devices, video game consoles, cellular or satellite wireless phones, and digital media players, etc.
[0195] Communication network 1704 is applicable to communication systems. For example, communication network 1704 can be a multi-access system that provides content (e.g., voice, data, video, messaging, broadcasting, etc.) to multiple wireless users. Communication network 1704 enables multiple wireless users to access this content through system resource sharing (including wireless bandwidth). For example, communication network 1704 can use one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FMDMA (SC-FDMA), etc.
[0196] The methods described herein can be implemented using computer programs, software, or firmware, and can be contained in a computer-readable medium executed by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, buffer memory, semiconductor memory devices, magnetic media (e.g., internal hard disks and removable disks), magneto-optical media, and optical media, such as optical discs (CDs) or digital multipurpose discs (DVDs). The processor associated with the software is used to implement an radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. A method of video decoding, comprising: obtaining an inter-layer picture for temporal motion vector prediction (TMVP) of an enhancement layer picture; generating a reference picture list for the inter-layer picture based on a reference picture list of a corresponding base layer picture; determining a reference picture index referring to the reference picture list of the inter-layer picture; and performing TMVP of the enhancement layer picture using the reference picture index and the reference picture list of the inter-layer picture.
2. The method of video decoding of claim 1, wherein the reference picture list of the inter-layer picture is generated by: adding reference pictures in the reference picture list of the corresponding base layer picture to the reference picture list of the inter-layer picture, and wherein the reference pictures are added to the reference picture list of the inter-layer picture with the same index as corresponding reference picture indices referring to the reference picture list of the corresponding base layer picture.
3. The method of video decoding of claim 2, wherein the reference pictures in the reference picture list of the inter-layer picture have the same picture order count (POC) value as corresponding reference pictures in the reference picture list of the corresponding base layer picture.
4. The method of video decoding of claim 1, wherein performing TMVP of the enhancement layer picture comprises: temporally scaling MVs of the inter-layer picture using the reference picture list; and determining MVs of the enhancement layer picture using the temporally scaled MVs of the inter-layer picture, wherein the TMVP of the enhancement layer picture is further based on the MVs of the enhancement layer picture.
5. An apparatus for video decoding, comprising: one or more processors configured to: obtain an inter-layer picture for temporal motion vector prediction (TMVP) of an enhancement layer picture; generate a reference picture list for the inter-layer picture based on a reference picture list of a corresponding base layer picture; determine a reference picture index referring to the reference picture list of the inter-layer picture; and perform TMVP of the enhancement layer picture using the reference picture index and the reference picture list of the inter-layer picture.
6. The apparatus for video decoding of claim 5, wherein the one or more processors are further configured to determine MVs of the inter-layer picture based on motion vectors (MVs) of the corresponding base layer picture, wherein the TMVP of the enhancement layer picture is further based on the MVs of the inter-layer picture.
7. The apparatus for video decoding of claim 5, wherein the one or more processors are further configured to determine temporally scaled MVs using the reference picture list, wherein the TMVP of the enhancement layer picture is further based on the temporally scaled MVs. the reference picture index refers to a reference picture associated with a video block collocated with a current video block in the enhancement layer picture.
9. A method of video decoding, comprising: identifying an enhancement layer video block of an enhancement layer picture; 8. The video decoding method of claim 1 or the video decoding apparatus of claim 5, wherein, determining a motion vector (MV) of a collocated base layer video block; and performing temporal motion vector prediction (TMVP) of the enhancement layer video block using the MV of the collocated base layer video block. spatially scaling the MV of the collocated base layer video block based on a spatial ratio between the base layer and the enhancement layer to generate a MV of a processed base layer video block; generating a reference picture list for the processed base layer video block based on a reference picture list of the collocated base layer video block; determining a reference picture index associated with the reference picture list of the processed base layer video block based on a reference picture index associated with the reference picture list of the collocated base layer video block; performing temporal motion vector prediction (TMVP) using the MV of the processed base layer video block and the reference picture index to generate a MV for the enhancement layer video block by: temporally scaling the MV of the processed base layer video block based on the reference picture index of the processed base layer video block, and generating the MV of the enhancement layer video block using the temporally scaled MV of the processed base layer video block; and decoding the enhancement layer picture using the MV of the enhancement layer video block.
10. The video decoding method of claim 9, further comprising: determining a temporal distance between the processed base layer video block and a reference picture of the processed base layer video block based on the reference picture index of the processed base layer video block; and temporally scaling the MV of the processed base layer video block based on both a temporal distance between the enhancement layer picture and a reference picture of the enhancement layer video block and the temporal distance between the processed base layer video block and the reference picture of the processed base layer video block.
11. A computer-readable medium containing instructions for causing one or more processors to: obtain an inter-layer picture for temporal motion vector prediction (TMVP) of an enhancement layer picture; generate a reference picture list of the inter-layer picture based on a reference picture list of a corresponding base layer picture; determine a reference picture index referring to the reference picture list of the inter-layer picture; and decode the enhancement layer picture using TMVP of the enhancement layer picture, wherein TMVP of the enhancement layer picture is based on the reference picture index and the reference picture list of the inter-layer picture.
12. The computer-readable medium of claim 11, wherein the inter-layer picture is used as a collocated picture for TMVP of the enhancement layer picture.
13. A video encoding method, comprising: obtaining an inter-layer picture for temporal motion vector prediction (TMVP) of an enhancement layer picture; generating a reference picture list of the inter-layer picture based on a reference picture list of a corresponding base layer picture; determining a reference picture index referring to the reference picture list of the inter-layer picture; and performing TMVP of the enhancement layer picture using the reference picture index and the reference picture list of the inter-layer picture.
14. The video encoding method of claim 13, wherein the reference picture list of the inter-layer picture is generated by: adding a reference picture in the reference picture list of the corresponding base layer picture to the reference picture list of the inter-layer picture, and wherein the reference picture is added to the reference picture list of the inter-layer picture with a same index as a corresponding reference picture index referring to the reference picture list of the corresponding base layer picture.
15. The video coding method of claim 14, wherein the reference picture in the reference picture list of the inter-layer picture has a same picture order count (POC) value as a corresponding reference picture in the reference picture list of the corresponding base layer picture.
16. A video coding apparatus comprising: one or more processors configured to: obtain an inter-layer picture for temporal motion vector prediction (TMVP) of an enhancement layer picture; generate a reference picture list of the inter-layer picture based on a reference picture list of a corresponding base layer picture; determine a reference picture index referring to the reference picture list of the inter-layer picture; and use the reference picture index and the reference picture list of the inter-layer picture to perform TMVP of the enhancement layer picture.
17. The video coding apparatus of claim 16, wherein the one or more processors are further configured to determine a motion vector (MV) of the inter-layer picture based on a MV of the corresponding base layer picture, wherein TMVP of the enhancement layer picture is further based on the MV of the inter-layer picture.
18. The video coding apparatus of claim 16, wherein the one or more processors are further configured to determine a temporally scaled MV using the reference picture list, wherein TMVP of the enhancement layer picture is further based on the temporally scaled MV.
19. A video coding method comprising: identifying an enhancement layer video block of an enhancement layer picture; determining a motion vector (MV) of a collocated base layer video block; spatially scaling the MV of the collocated base layer video block according to a spatial ratio between a base layer and an enhancement layer to generate a MV of a processed base layer video block; generating a reference picture list for the processed base layer video block based on a reference picture list of the collocated base layer video block; determining a reference picture index associated with the reference picture list of the processed base layer video block based on a reference picture index associated with the reference picture list of the collocated base layer video block; performing temporal motion vector prediction (TMVP) using the MV of the processed base layer video block and the reference picture index to generate a MV for the enhancement layer video block by: temporally scaling the MV of the processed base layer video block based on the reference picture index of the processed base layer video block, and generating the MV of the enhancement layer video block using the temporally scaled MV of the processed base layer video block; and encoding the enhancement layer picture using the MV of the enhancement layer video block.
20. The video coding method of claim 19, further comprising: determining a temporal distance between the processed base layer video block and a reference picture of the processed base layer video block based on the reference picture index of the processed base layer video block; and temporally scaling the MV of the processed base layer video block based on both: a temporal distance between the enhancement layer picture and a reference picture of the enhancement layer video block, and the temporal distance between the processed base layer video block and the reference picture of the processed base layer video block.
21. A computer readable medium containing instructions for causing one or more processors to: obtain an inter-layer picture for temporal motion vector prediction (TMVP) of an enhancement layer picture; generate a reference picture list of the inter-layer picture based on a reference picture list of a corresponding base layer picture; determine a reference picture index of a reference picture of the reference picture list referring to the inter-layer picture; and encode the enhancement layer picture using TMVP of the enhancement layer picture, wherein the TMVP of the enhancement layer picture is based on the reference picture index and the reference picture list of the inter-layer picture.
22. The computer readable medium of claim 21, wherein the inter-layer picture is used as a collocated picture for TMVP of the enhancement layer picture.
Citation Information
Patent Citations
Method and apparatus for motion vector prediction in scalable video coding
CN108156463B
Interlayer forcasting method for video signal
RU2384970C1
System and method for videoconferencing using scalable video coding and compositing scalable video conferencing servers
US20070200923A1