Scalability of Multi-Directional Video Streams

By breaking down multidirectional video into multiple tiles and providing high-quality and low-quality encoding for each tiles, the delay and inefficiency problems caused by viewport movement in multidirectional video communication is solved, and higher image quality and lower latency are achieved.

CN112703737BActive Publication Date: 2025-06-17APPLE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980059922.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-09-14
Filing Date
2019-09-05
Publication Date
2025-06-17
Estimated Expiration
2039-09-05

AI Technical Summary

Technical Problem

In multi-directional video communication, existing encoding technology leads to inefficient delay and coding efficiency when viewports move, and cannot effectively reduce delay and improve image quality.

Method used

By breaking down multi-directional video into multiple tiles and providing high-quality encoding and low-quality encoding for each tiles, the receiving terminal can decode and render data of the current viewport. When the viewport moves, only high-quality encodings for tiles containing the new viewport location are updated, and other tiles with low-quality encodings continue to be used.

Benefits of technology

This technology significantly reduces the delay when viewport is moved and improves the quality of viewport images in multi-directional video communication. By locally buffering and prefetching of unviewed viewport data, it reduces communication delay to the source terminal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112703737B_ABST
    Figure CN112703737B_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide techniques for reducing latency and improving the image quality of a viewport extracted from a multi-way video communication. A first stream of encoded video data is received from a source. The first stream includes encoded data for each of a plurality of tiles representing a multi-way video, where each tile corresponds to a predetermined spatial region of the multi-way video, and at least one of the plurality of tiles in the first stream contains the current viewport position at a receiver. The techniques include decoding the first stream and displaying the tile containing the current viewport position. When the viewport position at the receiver changes to include a new tile among the plurality of tiles, retrieving and decoding a first stream of the new tile, displaying the decoded content of the changed viewport position, and transmitting the changed viewport position to the source.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE DISCLOSURE

[0001] The present disclosure relates to encoding techniques for multi - view imaging applications.

[0002] Some modern imaging applications capture image data from multiple directions of a camera. Some cameras pivot during the image capture process, which allows the camera to scan across angles to capture image data, thereby extending the effective field of view of the camera. Some other cameras have multiple imaging systems that capture image data in several different fields of view. In either case, an aggregated image can be created that combines the image data captured from these multiple views.

[0003] Multiple rendering applications are available for multi - view content. One rendering application involves extracting and displaying a subset of the content contained in a multi - view image. For example, an observer may employ a head - mounted display and change the orientation of the display to identify a portion of the multi - view image that the observer is interested in. Alternatively, the observer may employ a stationary display and identify a portion of the multi - view image that the observer is interested in through user interface controls. In these rendering applications, the display device extracts a portion of the image content from the multi - view image (referred to as a "viewport" for convenience) and displays it. The display device will not display other portions of the multi - view image that are located outside the region occupied by the viewport. BRIEF DESCRIPTION OF THE DRAWINGS

[0004] Figure 1 A system is shown in accordance with one aspect of the present disclosure.

[0005] Figure 2 A rendering application for a receiving terminal is schematically shown in accordance with one aspect of the present disclosure.

[0006] Figure 3 An exemplary partitioning scheme is shown, where a frame is partitioned into non - overlapping tiles.

[0007] Figure 4 An encoded data stream that can be developed from the encoding of a single tile 410 is shown in accordance with one aspect of the present disclosure.

[0008] Figure 5 A method is shown in accordance with one aspect of the present disclosure.

[0009] Figure 6 A method is shown in accordance with one aspect of the present disclosure.

[0010] Figure 7 is shown Figure 6 exemplary data stream of

[0011] Figure 8 A frame of an omnidirectional video that can be encoded by a source terminal is shown.

[0012] Figure 9 Shows a frame of an omnidirectional video that can be encoded by a source terminal.

[0013] Figure 10 Is a simplified block diagram of an example video distribution system.

[0014] Figure 11 Shows a frame 1100 of a multi - directional video with a moving viewport.

[0015] Figure 12 Is a functional block diagram of an encoding system according to one aspect of the present disclosure.

[0016] Figure 13 Is a functional block diagram of a decoding system according to one aspect of the present disclosure.

[0017] Figure 14 Shows an exemplary multi - directional image projection format according to one aspect.

[0018] Figure 15 Shows an exemplary multi - directional image projection format according to another aspect.

[0019] Figure 16 Shows another exemplary multi - directional projection image format 1630.

[0020] Figure 17 Shows an exemplary prediction reference pattern.

[0021] Figure 18 Shows two exemplary multi - directional projections for combination.

[0022] Figure 19 Shows an exemplary system for generating a residual from two different multi - directional projections. Detailed Description

[0023] In a communication application, the aggregated source image data at the transmitter exceeds the data required for rendering the viewport at the receiver. Encoding techniques for transmitting the source data can consider the current viewport of the receiving rendering device. However, when considering a moving viewport, these encoding techniques result in encoding and transmission delays and inefficient encoding.

[0024] Aspects of the present disclosure provide techniques for reducing latency and improving the image quality of a viewport extracted from a multi-way video communication. According to such techniques, a first stream of encoded video data is received from a source. The first stream includes encoded data representing each of a plurality of tiles of the multi-way video, where each tile corresponds to a predetermined spatial region of the multi-way video, and at least one of the plurality of tiles in the first stream includes the current viewport position at the receiver. The techniques include decoding the first stream corresponding to at least one tile that includes the current viewport position, and displaying the decoded content of the current viewport position. When the viewport position at the receiver changes to include a new tile of the plurality of tiles, retrieving the first stream of the new tile, decoding the retrieved first stream, displaying the decoded content of the changed viewport position, and transmitting information representing the changed viewport position to the source.

[0025] Figure 1 System 100 is shown in accordance with one aspect of the present disclosure. As shown, system 100 is shown to include a source terminal 110 and a receiving terminal 120 interconnected by a network 130. The source terminal 110 may transmit an encoded representation of an omnidirectional video to the receiving terminal 120. The receiving terminal 120 may receive the encoded video, decode it, and display a selected portion of the decoded video.

[0026] Figure 1 The source terminal 110 is shown as a multi-way camera that captures image data of the local environment before encoding it. In another aspect, the source terminal 110 may receive an omnidirectional video from an external source (not shown) such as a streaming service or a storage device.

[0027] The receiving terminal 120 may determine a viewport position in a three-dimensional space represented by the multi-way image. The receiving terminal 120 may, for example, select a portion of the decoded video to display based on the orientation of the terminal in free space. Figure 1 The receiving terminal 120 is shown as a head-mounted display, but in other aspects, the receiving terminal 120 may be another type of display device, such as a fixed flat panel display, a smart phone, a tablet, a gaming device, or a portable media player. Each such display type may be provided with different types of user controls by which an observer identifies the viewport. Unless otherwise specified herein, the device type of the receiving terminal is immaterial to this discussion.

[0028] Network 130 represents any number of computers and / or communication networks extending from source terminal 110 to receiving terminal 120. Network 130 can include one or a combination of a circuit-switched communication network and / or a packet-switched communication network. Network 130 can transfer data between source terminal 110 and receiving terminal 120 via any number of wired and / or wireless communication media. Unless otherwise indicated herein, the architecture and operation of network 130 are not important for this discussion.

[0029] Figure 1 A communication configuration is shown in which encoded video data is transmitted in a single direction from source terminal 110 to receiving terminal 120. Aspects of the present disclosure can be applied to communication devices that exchange encoded video data in a two-way manner from terminal 110 to terminal 120 and from terminal 120 to terminal 110. The principles of the present disclosure can be applied to both unidirectional and bidirectional exchanges of video.

[0030] Figure 2 A rendering application for receiving terminal 200 is schematically shown in accordance with one aspect of the present disclosure. As shown, omnidirectional video is represented as if it exists along a spherical surface 210 provided around receiving terminal 200. Based on the orientation of receiving terminal 200, terminal 200 can select a portion of the video (referred to herein for convenience as a "viewport") and display the selected portion. As the orientation of receiving terminal 200 changes, terminal 200 can select different portions from the video. For example, Figure 2 A viewport changing from a first position 230 to a second position 240 along surface 210 is shown.

[0031] Aspects of the present disclosure can be applied to video compression techniques according to any one of a plurality of encoding protocols. For example, source terminal 110 ( Figure 1 ) can encode video data according to ITU-T / ISO MPEG encoding protocols such as H.265 (HEVC), H.264 (AVC), and the upcoming H.266 (VVC) standards, AOM encoding protocols such as AV1, or a precursor encoding protocol. Generally, such protocols parse the individual frames of the video into a spatial array of the video referred to herein as "pixel blocks" and can encode the pixel blocks in a regular encoding order such as raster scan order.

[0032] In one aspect, the individual frames of multi-directional content can be parsed into individual spatial regions, referred to herein as "tiles", and encoded as independent data streams. Figure 3An exemplary partitioning scheme is shown, in which frame 300 is partitioned into non-overlapping tiles 310.0 - 310.11. In the case where frame 300 represents omnidirectional content (e.g., it represents image content in a perfect 360° field of view, and the image content will be continuous across the opposite left edge 320 and right edge 322 of frame 300).

[0033] In one aspect, the tiles described herein can be a special case of the tiles used in some standards such as HEVC. In this aspect, the tiles used herein can be a "motion-constrained tile set", where all frames are segmented using exactly the same tile partitioning, and each tile in each frame is only allowed to use predictions from co-located tiles in other frames. Filtering in the decoder loop can also be not allowed on the tiles, thus providing decoding independence between the tiles.

[0034] Figure 4 An encoded data stream that can be developed from the encoding of a single tile 410 according to one aspect of the present disclosure is shown. The encoded tile 410 can be encoded into a number of representations 420 - 450 labeled "layer 0", "layer 1", "layer 2", and "layer 3" respectively, each representation corresponding to a predetermined bandwidth constraint. For example, layer 0 encoding can be generated for a 500 kbps representation, layer 1 encoding can be generated for a 2 Mbps representation, layer 2 encoding can be generated for a 4 Mbps representation, and layer 3 encoding can be generated for an 8 Mbps representation. In implementation, the number of layers and the selection of target bandwidths can be adjusted to suit individual application requirements.

[0035] The encoded tile 410 can also include multiple differential encodings 460 - 480, each differential encoding differentially encoded with respect to the encoded data of the layer 0 representation, and each differential encoding having a bandwidth bound to the bandwidth of another bandwidth layer. Thus, in the example where layer 0 encoding is generated for a 500 Kbps representation and layer 1 encoding is generated for a 2 Mbps representation, the layer 1 differential encoding 460 can be encoded at a 1.5 Mbps representation (1.5 Mbps = 2 Mbps - 500 Kbps). The other differential encodings 470, 480 can have data rates that match the differences between the data rates of their base layers 440, 450 and the data rate of the layer 0 encoding 420. In one aspect, the elements of the differential encodings 460, 470, 480 can be predictively encoded using the content of the corresponding blocks from the layer 0 encoding as a prediction reference; in such embodiments, the differential encodings 460, 470, 480 can be generated as enhancement layers according to a scalable coding protocol, where layer 0 serves as the base layer for those encodings.

[0036] The encoding 420 - 480 of the tile is shown as divided into individual blocks (e.g., for layer 0, 420 is block 420.1 - 420.N, for layer 1, 430 is block 430.1 - 430.N, etc.). Each block can be referenced by its own network identifier. During operation, the client device 120 ( Figure 1 ) can select individual blocks for download and request the blocks from the source terminal 120 ( Figure 1 ).

[0037] Figure 5 Method 500 according to one aspect of the present disclosure is shown. According to method 500, the terminal 110 can transmit a high - quality encoding of the tiles included in the current viewport (msg.510) and a low - quality encoding of other tiles (msg.520) from the source terminal 110 to the receiving terminal 120. Then, the receiving terminal 120 can decode and render the data of the current viewport (block 530). If the viewport does not move to include different tiles (block 540), the terminal 120 repeats decoding and rendering the current tiles (back to block 530). Alternatively, if the viewport moves such that the tiles included in the viewport change, the change in the viewport is reported back to the source terminal 110 (msg.550). Then, the source terminal 110 repeats by sending a high - quality encoding of the tiles at the new viewport position (back to msg.510) and low - quality tiles that do not include the new viewport position (msg.520).

[0038] Desired Figure 5 The operations shown in provide low - latency rendering of a new viewport of multi - view video in the presence of a communication delay between the source terminal 110 and the receiving terminal 120. By transmitting a low - quality encoding of tiles that do not belong to the current viewport, the receiving terminal 120 can locally buffer the data. If / when the viewport changes to a spatial location that coincides with one of the previously unviewed viewports, the locally buffered video can be decoded and displayed. Decoding and display can be performed without incurring the latency involved in a round - trip communication from the receiving terminal 120 to the source terminal 110, which would be required if the data for the unviewed viewports were not prefetched to the receiving device 120.

[0039] In one embodiment, the receiving terminal 120 can identify the position of the current viewport by identifying a spatial location within the multi - view image, where the viewport is located at the multi - view image, e.g., by identifying its position within the coordinate space defined for the image (see Figure 2 ). In another aspect, the receiving terminal 120 can identify the layer of the multi - view image ( Figure 3 ) in which its current viewport is located, and request blocks from that layer ( Figure 4 ) based on that identification.

[0040] Figure 6Illustrates an exemplary method 600 for tile downloading according to an aspect of the present disclosure. Figure 6 Illustrates the download operations that may occur for tiles that were not initially viewed but to which the viewport moves during operation. Thus, the receiving terminal 120 may issue requests for tiles at the layer 0 service level, and these requests are downloaded from the source terminal 110 to the terminal 120. Figure 6 Illustrates a request 610 for block Y of a tile from the layer 0 service level. The terminal 110 may provide the content of block Y in a response message 630. The request message 610 for block Y and the response message 630 may be interleaved with other requests and responses related to blocks of other tiles exchanged by the source terminal 110 and the receiving terminal 120 (shown in dashed lines), where the other tiles include both the tile where the viewport is located and other tiles that are not viewed.

[0041] In Figure 6 the example, the viewport changes (block 620) from a previous tile to the tile requested in msg.610. The viewport may change while the request for block Y (msg.610) is pending, or after the content of block Y has been received (msg.630). Figure 6 the example shows a viewport change (block 620) that occurs while msg.610 is pending. In response to the viewport change, the terminal 120 may determine, based on the history of previous requests, that block Y at the layer 0 service level has been requested or has been received and locally stored at the terminal 120. The terminal 120 may estimate whether there is time to request additional data for block Y (the differential layer) before block Y must be rendered. If so, the terminal 120 may issue a request (msg.640) for block Y of the new tile using the differential layer.

[0042] If the source terminal 110 provides the media content of the differential layer (msg.650) before block Y must be rendered, the receiving terminal 120 may render block Y (block 660) using the content developed from the content provided in messages 630 and 650. If not, the receiving terminal 120 may render block Y (block 660) using the content developed from the layer 0 service level (msg.630).

[0043] Figure 7 Illustrates the rendering timeline of the blocks that may occur according to the foregoing aspect of the present disclosure. Figure 7 Includes a data stream for a previous tile 710, such as a tile for the viewport position before the change of the viewport position in block 620 as in Figure 6 and Figure 7 includes a data stream for a new tile 720, such as for including Figure 6tiles at the new viewport position after the frame 620. Data for the previous tile 710 includes tiles Y-3 to Y+1, and data for the new tile includes tiles Y-3 to Y+4. In this example, tiles Y-3 to Y-1 of the previous tile are shown to have been retrieved at a relatively high service level or quality (shown as layer 3) and rendered before the viewport switch. When a viewport switch occurs from the previous tile 710 to the new tile 720 in the middle of tile Y-1, a layer 0 service level can be rendered for the tile 720 at tile Y-1. This may occur, for example, if the receiving device 120 estimates that there is not enough time to download the differential layer of the new tile 720 at tile Y-1, or if the receiving device 120 requests the differential layer of the tile but does not receive the differential layer to be rendered in time.

[0044] Figure 7 The example of shows the rendering of tile 720 at tiles Y to Y+2 using data from both layer 0 and the differential layer. For example, if the receiving device 120 has requested layer 0 level service for tiles Y to Y+2 before the viewport switch and (e.g., see Figure 6 request 610 in ) after the switch, the receiving device retrieves the differential layer for those tiles Y to Y+2 (e.g., see Figure 6 response 650 in ), then this may occur.

[0045] Figure 7 The example of shows the rendering of tile 720 from layer 3 starting from tile Y+3. For tiles for which a download request is made after the viewport switch occurs, a switch from the differential layer to a higher quality layer (e.g., layer 3) can occur. Thus, when the viewport changes from one tile to another, the receiving terminal 120 can determine the layer to request for the new tile based on its operating state and the transmission delay in the system. In some cases, there will be a transition period (such as Figure 7 layer 3 of tile Y+3 and later layers in ) after the viewport moves and before the receiving terminal can render the new viewport position at a high quality of service. The transition period can include rendering the new viewport position from a lower quality of service (e.g., Figure 7 layer 0 of tile Y-1 in ). The transition period can also include rendering the new viewport position from an enhanced lower quality of service (such as Figure 7 the enhanced layer 0 of the differential layer for tiles Y to Y+2 in ).

[0046] Figure 8Frame 800 of the omnidirectional video that can be encoded by source terminal 110 is shown. As shown, frame 800 is shown to have been parsed into a plurality of tiles 810.0 - 810.n. Each tile can be encoded in raster scan order. Thus, the content of tile 810.0 can be encoded separately from the content of tile 810.1, and the content of tile 810.1 can be encoded separately from the content of tile 810.2. In addition, tiles 810.1 - 810.n can be encoded in multiple layers, thereby generating discrete encoded data that can be segmented by both layers and tiles. In one aspect, the encoded data can also be segmented into time blocks. Thus, the encoded data can be segmented into discrete segments for each time block, tile, and layer.

[0047] As discussed, after frame 800 is encoded by source terminal 110 ( Figure 1 ) and transmitted to receiving terminal 120 ( Figure 1 ) and decoded, receiving terminal 120 can extract viewport 830 from frame 800. Receiving terminal 120 can locally display viewport 800. Receiving terminal 120 can send viewport information to source terminal 110, such as data identifying the position of viewport 830 within the region of frame 800. For example, receiving terminal 120 can send offset data, which is shown as an offset x and an offset y from origin 820, to identify the position of viewport 830 within the region of frame 800. In one aspect, the size and / or shape of viewport 830 can be included in the viewport information sent to source terminal 110. Source terminal 120 can then use the received viewport information to select which discrete portions of the encoded data to transmit to receiving terminal 120. In Figure 8 the example, viewport 830 spans tiles 810.5 and 810.6. Thus, the first layer can be sent for tiles 810.5 and 810.6, while the second layer can be sent for the remaining tiles that do not include any part of the viewport. For example, when the first layer provides higher quality video and the second layer provides more efficient encoding (high compression), the first layer can be sent to receiving terminal 120 for tiles 810.5 and 810.6, while the second layer providing lower quality video can be sent for some or all of the other tiles.

[0048] In one aspect, a lower quality layer can be provided for all tiles. In another aspect, a lower quality layer can be provided only for a part of frame 800. For example, a lower quality layer can be provided only for a 180 - degree perspective (instead of 360 degrees) centered on the current viewport, or a lower quality layer can be provided only in the region of frame 800 where the viewport may subsequently move.

[0049] In one aspect, frame 800 may be encoded according to a hierarchical coding protocol, where one layer is encoded as a base layer and other layers are encoded as enhancement layers of the base layer. The enhancement layers may be predicted from one or more lower layers. For example, a first enhancement layer may be predicted from the base layer, and a second higher enhancement layer may be predicted from the base layer or from the first lower enhancement layer.

[0050] The enhancement layers may be encoded differentially or predictively from one or more lower layers. Non-enhancement layers (such as the base layer) may be encoded independently of other layers. Reconstruction at the decoder of a differentially encoded layer will require both the encoded data segment of the differentially encoded layer and the segment from which the differentially encoded layer is predicted. In the case of a prediction-encoded layer, transmitting the layer may include transmitting the discrete encoded data segment of the prediction-encoded layer and also transmitting the discrete encoded data segment of the layer used as a prediction reference. In one example, for all tiles, the differentially hierarchical encoded, lower base layer of frame 800 may be sent to receiving terminal 120, while the discrete data segments of the higher differential layers, which are encoded using predictions from the base layer, may be sent only for tiles 810.5 and 810.6, since viewport 830 is included in those tiles.

[0051] Figure 9 Frame 900 of an omnidirectional video that may be encoded by source terminal 110 is shown. As shown, as in Figure 8 frame 800, frame 900 is shown as having been parsed into a plurality of tiles 810.0 - 810.n. Frame 900 may represent a different video time than frame 800, e.g., frame 900 may be a later time on the video timeline. At a later time, the viewport of receiving terminal 120 may have moved to the position of viewport 930, which may be identified by an offset x' and an offset y' relative to origin 820. When the viewport of receiving terminal 120 moves from the Figure 8 position of viewport 830 in Figure 9 to the position of viewport 930 in Figure 9 , the receiving terminal sends the new viewport information to source terminal 110. In response, receiving terminal 120 may change which discrete segments of the encoded video are sent to the receiving terminal such that the first layer may be sent for tiles that include a portion of the viewport, while the second layer may be sent for tiles that do not include a portion of the viewport. In the

[0052] Figure 10FIG. 0 is a simplified block diagram of an example video distribution system 100 of the present invention, including when multi-directional video is pre-encoded and stored on a server. System 1000 may include a distribution server system 1010 and client devices 1020 connected via a communication network 1030. The distribution system 1000 may provide encoded multi-directional video data to the client 1020 in response to a client request. The client 1020 may decode the encoded video data and present it on a display.

[0053] The distribution server 1010 may include a storage system 1040 on which pre-encoded multi-directional video is stored in various layers for download by the client device 1020. The distribution server 1010 may store several encoded representations of a video content item, shown as layers 1, 2, and 3, which have been encoded with different encoding parameters. The video content item includes a manifest file that contains pointers to the encoded video data blocks of each layer.

[0054] In Figure 10 the example, the average bitrates of layer 1 and layer 2 are different, and layer 2 is capable of reconstructing the video content item with higher quality at a higher average bitrate compared to the bitrate provided by layer 1. The differences in bitrate and quality may be caused by differences in encoding parameters (e.g., encoding complexity, frame rate, frame size, etc.). Layer 3 may be an enhancement layer of layer 1, which, when decoded in combination with layer 1, may improve the quality of the representation of layer 1 when decoded by itself. Each video layer 1-3 may be parsed into multiple blocks CH1.1-CH1.N, CH2.1-CH2.N, and CH3.1-CH3.N. The manifest file 1050 may include pointers to each encoded video data block of each layer. Different blocks may be obtained from the storage device and delivered to the client 1020 via channels defined in the network 1030. The channel stream 1040 represents the aggregation of transmission blocks from multiple layers. Additionally, as described above in connection with Figure 4 and Figure 5 multi-directional video may be spatially segmented into tiles. Figure 10 FIG. shows the blocks of the respective layers available for one tile. The manifest 1050 may additionally include other tiles (not shown in Figure 10 ), for example, by providing metadata and pointers to multiple layers that include the storage locations of the encoded data blocks for each layer in the respective layers.

[0055] Figure 10 The example of N)。The block boundaries can provide preferred points for flow switching between layers. For example, flow switching can be facilitated by resetting the motion prediction coding state at the switching point.

[0056] Times A, B, C, and D are partially shown in Figure 10 to assist in showing a moving viewport in one aspect of the present disclosure. Times A, B, C, and D are positioned along the streaming timeline of the media block referenced by Listing 1050. Specifically, times A, B, and D may respectively correspond to the start of time periods t1, t2, and t3, while time C may correspond to some time in the middle of time period t2, between the start of t2 and the start of t3.

[0057] In one aspect, the multi-directional image data can include a depth map and / or occlusion information. The depth map and / or occlusion information can be included as separate channels, and Listing 1050 can include references to these separate channels for the depth map and / or occlusion information.

[0058] Figure 11 Frame 1100 of a multi-directional video with a moving viewport is shown. As shown, frame 1100 is shown as having been parsed into a plurality of tiles 1110.0 - 1110.n. Superimposed on frame 1100 are a viewport position 1130 that can correspond to a first position of the viewport in the client 1020 at a first time, and a viewport position 1140 that can correspond to a second position of the same viewport at a second time.

[0059] In one aspect, in the steady state where the viewport does not move, the client 1020 can extract the viewport image from the high reconstruction quality of layer 2. During the transition period, when the viewport moves into a new spatial tile, the client 1020 can extract the viewport image from the combined reconstruction of layer 1 and enhancement layer 3, and then once the client 1020 can use layer 2 again when the viewport moves to a new spatial tile, it returns to the steady state by extracting the viewport image from layer 2. An example of this is shown in Tables 1 and 2 for the viewport of the client 1020 that jumps from viewport position 1130 to viewport position 1140 at time C. The requests of the client 1020 for the tile layers are listed in Table 1, and the layers from which the viewport image is extracted are listed in Table 2.

[0060]

[0061] Table 1 - Requests for Tiles

[0062] Time A Time B Time C Time D Viewport position Tile 1110.0 Tile 1110.0 Tile 1110.5 Tile 1110.5 Extract for viewport Layer 2 Layer 2 Layer 1; then Layer 1 + Layer 3 Layer 2

[0063] Table 2 - Viewport Extraction

[0064] Under the initial steady-state condition during time period t1, the viewport does not move, and the viewport position 1130 is completely contained within tile 1110.0. Layer 2, which is the higher-quality layer, can be requested by client 1020 from server 1010 for tile 1110.0 at time A, as shown in Table 1. For tiles that are not included in the viewport at position 1130 (tiles 1110.1 - 1110.n), the lower-quality and more highly compressed layer 1 is requested instead. Thus, for all tiles other than tile 1110.0, layer 1 tiles are requested at time A for the duration of time period t1. Then, client 1020 extracts the viewport from the reconstruction of layer 2 starting from time A.

[0065] At time B, the viewport has not moved, so client 1020 requests the same layer for the same tiles as at time A, but the request is for a specific tile corresponding to time period t2. At time C, the viewport of client 1020 can jump from viewport position 1130 to position 1140. At time point C, sometime between the start and end of t2, the lower-quality layer 1 has been requested for the new viewport position (tile 1110.5). Thus, when the viewport moves, the viewport can be immediately extracted from layer 1 at time C. At time C, layer 3 can be requested, and once it is available, the combination of layer 1 and enhanced layer 3 can be used to extract the viewport image at client 1020. At time D, client 1020 can return to the steady state by requesting layer 2 for the tiles that contain the viewport position and layer 0 for the tiles that do not contain the viewport position.

[0066] Figure 12is a functional block diagram of an encoding system 1200 according to one aspect of the present disclosure. The system 1200 may include an image source 1210, an image processing system 1220, a video encoder 1230, a video decoder 1240, a reference picture repository 1250, and a predictor 1260. The image source 1210 may generate image data as a multi-directional image that includes image data of a field of view extending around a reference point in multiple directions. The image processing system 1220 may perform image processing operations to condition the image for encoding. In one aspect, the image processing system 1220 may generate different versions of the source data to facilitate encoding the source data into multi-layer encoded data. For example, the image processing system 1220 may generate multiple different projections of a source video aggregated from multiple cameras. In another example, the image processing system 1220 may generate resolutions of the source video for a higher layer with a higher spatial resolution and a lower layer with a lower spatial resolution. The video encoder 1230 generally may generate a multi-layer encoded representation of its input image data by exploiting spatial and / or temporal redundancy in the image data. The video encoder 1230 may output an encoded representation of the input data that consumes less bandwidth when transmitted and / or stored than the original source video. The video encoder 1230 may output data in discrete time blocks corresponding to temporal portions of the source image data, and in some aspects, may decode data encoded for a separate time block independently of other time blocks. The video encoder 1230 may also output data in discrete layers, and in some aspects, may transmit separate layers independently of other layers.

[0067] The video decoder 1240 may reverse the encoding operations performed by the video encoder 1230 to obtain a reconstructed picture from the encoded video data. Generally, the encoding process applied by the video encoder 1230 is a lossy process, which results in the reconstructed picture having various errors when compared to the original picture. The video decoder 1240 may reconstruct a picture of a selected encoded picture designated as a “reference picture” and store the decoded reference picture in the reference picture repository 1250. In the absence of transmission errors, the decoded reference picture may replicate the decoded reference picture obtained by the decoder ( Figure 12 not shown).

[0068] Predictor 1260 can select a prediction reference for a newly encoded input picture. For each part of the input picture being encoded (referred to for convenience as a "pixel block"), predictor 1260 can select an encoding mode and identify the part of the reference picture that can serve as the prediction reference search for the pixel block being encoded. The encoding mode can be an intra encoding mode, in which case the prediction reference can be drawn from a previously encoded (and decoded) part of the picture being encoded. Alternatively, the encoding mode can be an inter encoding mode, in which case the prediction reference can be drawn from another previously encoded and decoded picture. In one aspect of hierarchical encoding, the prediction reference can be a pixel block previously decoded from another layer, typically a lower layer than the layer currently being encoded. In the case of encoding two layers in two different projection formats of multi-view video, a function such as an image warping function can be applied to the reference image in one projection format at the first layer to predict the pixel block in a different projection format at the second layer.

[0069] In another aspect of the hierarchical encoding system, the differential encoded enhancement layer can be encoded with a restricted prediction reference to enable seeking or layer / layer switching in the middle of encoding the enhancement layer blocks. In a first aspect, predictor 1260 can limit the prediction reference for only each frame in the enhancement layer to frames in the base layer or other lower layers. When predicting each frame of the enhancement layer without referring to other frames of the enhancement layer, the decoder can effectively switch to the enhancement layer at any frame because the previous enhancement layer frames will no longer be needed as prediction references. In a second aspect, predictor 1260 may need to predict only every Nth frame (such as every other frame) within a block from only the base layer or other lower layers to enable seeking every Nth frame within the encoded data block.

[0070] When a suitable prediction reference is identified, predictor 1260 can provide the prediction data to video encoder 1230. Video encoder 1230 can differentially encode the input video data for the prediction data provided by predictor 1260. Typically, the prediction operation and differential encoding operate on a per-pixel-block basis. The prediction residual (representing the pixel-wise difference between the input pixel block and the predicted pixel block) can undergo further encoding operations to further reduce the bandwidth.

[0071] As noted, the encoded video data output by video encoder 1230 should consume less bandwidth when transmitted and / or stored than the input data. Encoding system 1200 can output the encoded video data to an output device 1270, such as a transceiver, that can transmit the encoded video data across communication network 130 ( Figure 1 )). Alternatively, encoding system 1200 can output the encoded data to a storage device (not shown) such as an electronic storage medium, a magnetic storage medium, and / or an optical storage medium.

[0072] The transceiver 1270 may also receive viewport information from a decoding terminal ( Figure 7 ) and provide the viewport information to the controller 1280. The controller 1280 may control the image processor 1220, the overall video encoding process, including the video encoder 1230 and the transceiver 1270. The viewport information received by the transceiver 1270 may include the viewport position and / or the preferred projection format. In one aspect, the controller 1280 may control the transceiver 1270 based on the viewport information to send certain encoded layers for certain spatial tiles while sending different encoded layers for other tiles. In another aspect, the controller 1280 may control the allowable prediction references in certain frames of certain layers. In yet another aspect, the controller 1280 may control the projection format or the scaling layer generated by the image processor 1230 based on the received viewport information.

[0073] Figure 13 is a functional block diagram of a decoding system 1300 according to one aspect of the present disclosure. The decoding system 1300 may include a transceiver 1310, a buffer 1315, a video decoder 1320, an image processor 1330, a video receiver 1340, a reference picture repository 1350, a predictor 1360, and a controller 1370. The transceiver 1310 may receive encoded video data from a channel and route it to the buffer 1315 before sending it to the video decoder 1320. The encoded video data may be organized into temporal blocks and spatial tiles and may include different encoded layers for different tiles. The video data buffered in the buffer 1315 may span the video time of multiple blocks. The video decoder 1320 may decode the encoded video data with reference to the prediction data provided by the predictor 1360. The video decoder 1320 may output the decoded video data in a representation determined by the source image processor (such as Figure 12 the image processor 1220) of the encoding system that generated the encoded video. The image processor 1330 may extract video data from the decoded video according to the viewport orientation currently valid at the decoding system. The image processor 1330 may output the extracted viewport data to the video receiving device 1340. The controller 1370 may control the image processor 1330, including the video decoding process of the video decoder 1320 and the transceiver 1310.

[0074] As shown, the video receiver 1340 may consume the decoded video generated by the decoding system 1300. The video receiver 1340 may be implemented by, for example, a display device that renders the decoded video. In other applications, the video receiver 1340 may be implemented by a computer application (e.g., a gaming application, a virtual reality application, and / or a video editing application) that integrates the decoded video into its content. In some applications, the video receiver may process the entire multi-view field of the decoded video for its application, but in other applications, the video receiver 1340 may process a subset of the content selected from the decoded video. For example, when rendering the decoded video on a flat panel display, it may be sufficient to display only a selected subset of the multi-view video. In another application, the decoded video may be rendered in a multi-view format, such as in a planetarium.

[0075] The transceiver 1310 may also send viewport information (such as viewport position and / or preferred projection format) provided by the controller 1370 to the encoded video source, such as Figure 12 the terminal 1200. When the viewport position changes, the controller 1370 may provide new viewport information to the transceiver 1310 for sending to the encoded video source. In response to the new viewport information, the transceiver 1310 may receive lost layers of certain previously received but not yet decoded encoded video tiles and store them in the buffer 1315. Then, the decoder 1320 may use these replacement layers (previously lost) instead of the previously received layers based on the old viewport position to decode these tiles.

[0076] The controller 1370 may determine the viewport information based on the viewport position. In one example, the viewport information may include only the viewport position, and the encoded video source may then use this position to identify which encoded layers are provided to the decoding system 1300 for a particular spatial tile. In another example, the viewport information sent from the decoding system may include a specific request for a particular layer of a particular tile, thus leaving most of the viewport position mapping in the decoding system. In yet another example, the viewport information may include a request for a particular projection format based on the viewport position.

[0077] The principles of the present disclosure may be used in applications having multiple multi-view image projection formats. In one aspect, appropriate projection conversion functions may be used to convert between Figures 14 - 16 the various projection formats.

[0078] Figure 14An exemplary multi-directional image projection format according to one aspect is shown. The multi-directional image 1430 can be generated by a camera 1410 pivoting along an axis. During operation, the camera 1410 can capture image content as it pivots through a predetermined angular distance 1420 (preferably, a full 360°) and combine the captured image content into a 360° image. The capture operation can result in a multi-directional image 1430 representing a multi-directional field of view that has been divided along a slice 1422 that divides the cylindrical field of view into a two-dimensional data array. In the multi-directional image 1430, pixels on either edge 1432, 1434 of the image 1430 represent adjacent image content, even if they appear on different edges of the multi-directional image 1430.

[0079] Figure 15 An exemplary multi-directional image projection format according to another aspect is shown. In Figure 15 this aspect, the camera 1510 can have image sensors 1512 - 1516 that capture image data in different fields of view from a common reference point. The camera 1510 can output a multi-directional image 1530 in which the image content is arranged according to a cube map capture operation 1520, where the sensors 1512 - 1516 capture image data in different fields of view 1521 - 1526 (typically six) with respect to the camera 1510. The image data of the different fields of view 1521 - 1526 can be stitched together according to a cube map layout 1530. In Figure 15 the example shown, six sub-images corresponding to a left view 1521, a front view 1522, a right view 1523, a rear view 1524, a top view 1525, and a bottom view 1526 can be captured, stitched, and arranged within the multi-directional picture 1530 according to the "seams" of the image content between the corresponding views 1521 - 1526. Thus, as Figure 15 shown, pixels from the front image 1532 that are adjacent to pixels from each of the left image 1531, the right image 1533, the top image 1535, and the bottom image 1536 represent image content that is adjacent to the content of the adjacent sub-images, respectively. Similarly, pixels from the right image 1533 and the rear image 1534 that are adjacent to each other represent adjacent image content. Additionally, the content of the terminal edge 1538 of the rear image 1534 is adjacent to the content of the opposite terminal edge 1539 of the left image. The image 1530 can also have regions 1537.1 - 1537.4 that do not belong to any image. Figure 15 The representation shown in

[0080] is generally referred to as a "cube map" image. Figure 3The encoding technique can be applied to the cube map image 1530.

[0081] In other encoding applications, the cube map image 1530 can be repacked before encoding to eliminate the empty regions 1537.1 - 1537.4, as shown in the image 1540. Figure 3 The techniques described in can also be applied to the packed image frame 1540. After decoding, the decoded image data can be unpacked before display.

[0082] Figure 16 Another exemplary multi - directional projection image format 1630 is shown. Figure 16 The frame format of can be generated by another type of omnidirectional camera 1600 called a panoramic camera. The panoramic camera typically consists of a pair of fish - eye lenses 1612, 1614 and an associated imaging device (not shown), with each imaging device arranged to capture image data in a hemispherical field of view. The images captured from the hemispherical field of view can be stitched together to represent the image data in a complete 360° field of view. For example, Figure 16 An image 1630 is shown that includes image content 1631, 1632 containing hemispherical views 1622, 1624 from the camera and joined at the seam 1635. The techniques described above can also be applied to the multi - directional image data of such format 1630.

[0083] In one aspect, cameras such as Figures 14 - 16 the cameras 1410, 1510, and 1610 in, in addition to visible light, can also capture depth or occlusion information. In some cases, the depth and occlusion information can be stored as separate data channels in a multi - projection format (such as images, such as 1430, 1530, 1540, and 1630). In other cases, the depth and occlusion information can be included in a manifest as separate data channels, such as Figure 10 the manifest 1050 of.

[0084] Figure 17 An exemplary prediction reference mode is shown. The video sequence 1700 includes a base layer 1720 and an enhancement layer 1710, with each layer including a series of corresponding frames. The base layer 1720 includes an intra - coded frame L0.I0, followed by predicted frames L0.P1 to L0.P7. The enhancement layer 1710 includes predicted frames L1.P0–L1.P7. The intra - coded frame L0.I0 can be encoded without prediction from any other frame. The predicted frames can be encoded by predicting the pixel blocks of the frame portions of the reference frames indicated by the solid - line arrows in Figure 17 where the arrows point to the reference frames that can be used as prediction references for the frames at the tails of the arrows. For example, the predicted frames in the base layer can be predicted using only the previous base - layer frame as a prediction reference. As Figure 17As shown, only frame L0.I0 is used as a reference to predict L0.P1. L0.P1 may be a reference for L0.P2, L0.P2 may be a reference for L0.P3, and so on, as indicated by the arrows within the base layer 1720. The frames of the enhancement layer 1710 can be predicted using only the corresponding base layer reference frames, such that L0.I0 can be a prediction reference for L1.P0, L0.P1 can be a prediction reference for L1.P1, and so on.

[0085] In one aspect, the enhancement layer 1710 frames can also be predicted from previous enhancement layer frames, as Figure 17 indicated by the optional dashed arrows in. For example, frame L1.P7 can be predicted from L0.P7 or L1.P6. The prediction references within the enhancement layer 1710 can be restricted such that only a subset of the enhancement layer frames can use other enhancement layer frames as prediction references, and this subset of enhancement layer frames can follow a pattern. In the Figure 17 example, every other frame (L1.P0, L1.P2, L1.P4, and L1.P6) of the enhancement layer 1710 is predicted only from the corresponding base layer frames, while the alternating frames (L1.P1, L1.P3, L1.P5, L1.P7) can be predicted from the base layer frames or previous enhancement layer frames. Layer switching to the enhancement layer 1710 can be facilitated at frames predicted only from the lower layer because previous frames of the enhancement layer do not need to be previously decoded to use the reference frames. Enhancement layer frames predicted only from lower layer frames can be considered safe switch frames, sometimes referred to as key frames, because previous frames from the enhancement layer are not needed for the correct decoding of these safe switch frames.

[0086] In one aspect, when some decoding quality drift is tolerable, the receiving terminal can switch to a new layer or new tier on a non-safe switch frame. A non-safe switch frame can be decoded without accessing the reference frames used for its prediction, and as errors from incorrect predictions accumulate into so-called quality drift, the quality gradually degrades. Error concealment techniques can be used to mitigate the quality drift caused by switching at non-safe switch enhancement layer frames. Exemplary error concealment techniques include predicting from frames similar to the lost reference frames, and a periodic internal refresh mechanism. By tolerating some quality shift caused by switching at non-safe switch frames, the latency between the moving viewport and the image presenting the new viewport position can be reduced.

[0087] Figure 18 Two exemplary multi-view projections for combination are shown. Images of the same scene can be encoded in multiple projection formats. In Figure 18In an example, a multi-directional scene is encoded as a first image having a first projection format, such as the equirectangular projection format image 1810, and the same scene is encoded as a second image having a second projection format, such as the cube map projection format image 1820. The region of interest 1812 projected onto the equirectangular image 1810 and the region of interest 1822 projected onto the cube map image 1820 can both correspond to the same region of interest in the scene projected onto images 1810 and 1820. The cube map image 1820 can include empty regions 1837.1 - 1837.4 and cube faces 1831 - 1836 for left, front, right, rear, top, and bottom.

[0088] In one aspect, multiple projection formats can be combined to form a better reconstruction of the region of interest (ROI) than can be produced from a single projection format. The reconstructed region of interest ROI can be produced from a weighted sum of the encoded projections or from a filtered sum of the encoded projections combo . For example, Figure 18 the region of interest in the scene of

[0089] ROI combo = f(ROI1, ROI2)

[0090] where f() is a function for combining two region of interest images, the first region of interest image ROI1 can be, for example, an equirectangular region of interest image from ROI 1812, and the second region of interest image ROI2 can be, for example, a cube map region of interest image from ROI 1822. If f() is a weighted sum,

[0091] ROI combo = alpha * ROI1 + beta * ROI2

[0092] where α and β are predetermined constants, and α + β = 1. In cases where the pixel positions do not exactly correspond in the combined projection format, a projection format conversion function can be used, such as:

[0093] ROI combo = alpha * PConv(ROI1) + beta * ROI2

[0094] where PConv() is a function that converts an image in the first projection format to the second projection format. For example, PConv() can be just an upsampling or downsampling function.

[0095] In another aspect, the best projection for encoding the entire multi-directional scene (such as for encoding the base layer) can be different from the best projection format for encoding only the region of interest (such as for encoding in the enhancement layer). Thus, Figure 18Multi-layer encoding of the scenario may include encoding the entire equirectangular image 1810 in a first layer and encoding only the ROI 1822 of the cube map image 1820 in a second layer. For example, the ROI 1822 may be encoded by encoding the entire front surface 1832 as a tile of the cube map image 1820. On the other hand, the second layer may be encoded as an enhancement layer above the first layer base layer, as Figure 19 shown.

[0096] Figure 19 An exemplary system for generating residuals from two different multi-directional projections is shown. The base layer ROI image 1910 in projection format P1 can be converted to projection format P2 through a conversion process 1902 to create a prediction of the ROI image 1920 in projection format P2. At the adder 1904, the prediction image from the conversion process 1902 is subtracted from the actual P2 ROI image 1920 to generate a P2 residual ROI, which can then be encoded as a P2 projection enhancement layer on the P1 base layer. In one aspect, the base layer may encode the entire scene in projection P1, while the enhancement layer may encode only the region of interest within the scene in projection P2. For example, this aspect may be beneficial when projection P1 is preferably used to encode the entire scene and projection P2 is preferably used to encode a specific region of interest. For example, referring to Figure 18 , the first layer may be encoded as a base layer including the entire equirectangular image 1810, and the second layer may be encoded as an enhancement layer including a subset of the cube map image 1820, such as a single tile or region of interest.

[0097] The foregoing discussion has described the operation of various aspects of the present disclosure in the context of video encoders and decoders. These components are often provided as electronic devices. The video decoder and / or controller may be embedded in an integrated circuit, such as an application specific integrated circuit, a field programmable gate array, and / or a digital signal processor. Alternatively, they may be embedded in a computer program executed on a camera device, a personal computer, a laptop, a tablet, a smart phone, or a computer server. Such computer programs include processor instructions and are typically stored in a physical storage medium such as an electronic-based, magnetic-based, and / or optical-based storage device, where they are read and executed by a processor. The decoder is often encapsulated in a consumer electronic device, such as a smart phone, a tablet, a gaming system, a DVD player, a portable media player, etc.; and, they may also be encapsulated in a consumer software application, such as a video game, a media player, a media editor, etc. And, of course, these components may be provided as a hybrid system that distributes functions as needed between dedicated hardware components and programmed general purpose processors.

Claims

1. A video receiving method, comprising: Receive a first stream of encoded data representing each tile of a multi - view video from a source, the first stream including a first tile encoded at a first quality layer and other tiles of the plurality of tiles that are encoded as a base layer at a second quality layer, where each tile corresponds to a predetermined spatial region of the multi - view video, and a current viewport position at a receiver includes at least the first tile, and the first quality layer is higher than the second quality layer; Decode the first quality layer of the first tile from the first stream; Display the decoded first layer of the current viewport position; When the viewport position at the receiver changes to include a second tile among the other tiles: Decode the base layer of the second quality layer of the second tile from the first stream, and Transmit an indication of the second tile to the source; In response to the transmission, receive a second stream of encoded video data from the source, the second stream including the second tile encoded as an enhancement layer of the base layer at the second quality layer, where when the enhancement layer is decoded and combined with the base layer of the second quality layer, the enhancement layer corresponds to a third quality layer higher than the second quality layer; Decode the enhancement layer of the second tile from the second stream using the base layer of the second tile from the first stream; And Display the decoded enhancement layer of the second quality layer of the changed viewport position, where the decoded first layer and the decoded enhancement layer of the second quality layer are not the same.

2. The video receiving method according to claim 1, wherein the base layer of the second tile is received and stored in a local buffer before the viewport position changes to include the second tile.

3. The video receiving method according to claim 1, further comprising: Display the decoded base layer of the second quality layer of the changed viewport position.

4. The video receiving method according to claim 3, wherein: The encoded data of the first stream and the second stream includes data encoded at a first, lower quality layer and a second, higher quality layer. The first stream includes the second layer for tiles of the current viewport position, a new first stream retrieved from a local storage device includes the first layer for tiles of the changed viewport position, and the second stream includes the second layer for tiles of the changed viewport position.

5. The video receiving method according to claim 3, wherein the encoded data of the first stream includes data encoded in a first projection format, and the data encoded in the second stream includes data encoded in a second projection format, and the video receiving method further comprises: Select the second projection format based on the changed viewport position.

6. The video receiving method according to claim 3, wherein: Encode the encoded data of the first stream and the second stream according to a hierarchical coding protocol. The first stream includes an enhancement layer for tiles of the current viewport position, the first stream retrieved from a local storage device includes a base layer, and the second stream includes an enhancement layer for tiles of the changed viewport position.

7. The video receiving method according to claim 6, wherein: A first subset of frames of the second - stream enhancement layer is predicted only from reconstructed base - layer frames; and A second subset of frames of the second - stream enhancement layer is predicted from reconstructed frames of both the base layer and the enhancement layer; and Decoding of the second stream starts from a frame time corresponding to frames in the first subset of frames.

8. The video receiving method according to claim 6, wherein decoding of the second stream starts on a frame with an unavailable prediction reference, and the video receiving method further comprises: Use error - concealment techniques to mitigate quality drift caused by lost prediction references.

9. The video receiving method according to claim 6, wherein: The base layer is encoded in a first projection format, the enhancement layer is encoded in a second projection format, and the video receiving method further includes: Enhanced layer data is predicted from the reconstructed base layer data by applying a function that converts from the first projection format to the second projection format.

10. A video receiving system, comprising: A receiver for receiving an encoded video stream from a source; A decoder for decoding the encoded video stream; A controller for causing: A first stream of encoded data representing each tile of a multi-view video is received by the receiver from the source, the first stream including a first tile encoded at a first quality layer and other tiles of a base layer encoded as a second quality layer, where each tile corresponds to a predetermined spatial region of the multi-view video, and the current viewport position at the receiver includes at least the first tile, the first quality layer being higher than the second quality layer; The decoder decodes the first quality layer of the first tile from the first stream; The decoded first layer of the current viewport position is displayed; And When the viewport position at the receiver changes to include a second tile among the other tiles: The decoder decodes the base layer of the second quality layer of the second tile from the first stream, and Transmits an indication of the second tile to the source; In response to the transmission, a second stream of encoded video data is received from the source, the second stream including the second tile of the enhancement layer of the base layer encoded as the second quality layer, where when the enhancement layer is decoded and combined with the base layer of the second quality layer, the enhancement layer corresponds to a third quality layer higher than the second quality layer; The enhancement layer of the second tile from the second stream is decoded using the base layer of the second tile from the first stream; And The decoded enhancement layer of the second quality layer of the changed viewport position is displayed, where the decoded first layer and the decoded enhancement layer of the second quality layer are not the same.

11. The video receiving system according to claim 10, further comprising: A local buffer for buffering the encoded video data received by the receiver; Wherein the base layer of the second tile is received and stored in the local buffer before the viewport position changes to include the second tile.

12. The video receiving system according to claim 10, wherein the controller further causes: To display the base layer of the second quality layer of the decoded changed viewport position.

13. The video receiving system according to claim 12, wherein: The encoded data of the first stream and the second stream includes data encoded at a first layer of lower quality and a second layer of higher quality, the first stream includes the second layer for the tiles of the current viewport position, a new first stream retrieved from a local storage device includes the first layer for the tiles of the changed viewport position, and the second stream includes the second layer for the tiles of the changed viewport position.

14. The video receiving system according to claim 12, wherein the encoded data of the first stream includes data encoded in a first projection format, and the data encoded in the second stream includes data encoded in a second projection format, and wherein the controller further causes: To select the second projection format based on the changed viewport position.

15. The video receiving system according to claim 12, wherein: The encoded data of the first stream and the second stream is encoded according to a hierarchical coding protocol, the first stream includes the enhancement layer for the tiles of the current viewport position, the first stream retrieved from a local storage device includes the base layer, and the second stream includes the enhancement layer for the tiles of the changed viewport position.

16. A non-transitory computer-readable medium, the non-transitory computer-readable medium comprising instructions that, when executed by a processor, cause: To receive, from a source, a first stream of encoded data representing each tile of a plurality of tiles of a multi-view video, the first stream including a first tile encoded at a first quality layer and other tiles of the plurality of tiles encoded as the base layer of a second quality layer, wherein each tile corresponds to a predetermined spatial region of the multi-view video, and the current viewport position at the receiver includes at least the first tile, and the first quality layer is higher than the second quality layer; The first quality layer of the first tile from the first stream is decoded; Decode the first layer of the current viewport position shown; When the viewport position at the receiver changes to include a second tile among the other tiles: Decode the base layer of the second quality layer of the second tile from the first stream, and Transmit an indication of the second tile to the source; In response to the transmission, receive a second stream of encoded video data from the source, the second stream including the second tile encoded as an enhancement layer of the base layer of the second quality layer, wherein when the enhancement layer is decoded and combined with the base layer of the second quality layer, the enhancement layer corresponds to a third quality layer higher than the second quality layer; Decode the enhancement layer of the second tile from the second stream using the base layer of the second tile from the first stream; And Display the enhancement layer of the decoded second quality layer of the changed viewport position, wherein the decoded first layer and the enhancement layer of the decoded second quality layer are different.

17. The computer-readable medium according to claim 16, wherein the base layer of the second tile is received and stored in a local buffer before the viewport position is changed to include the second tile.

18. The computer-readable medium according to claim 16, wherein the instructions further cause: To display the base layer of the decoded second quality layer of the changed viewport position.

19. The computer-readable medium according to claim 18, wherein: The encoded data of the first stream and the second stream includes data encoded at a first layer of lower quality and a second layer of higher quality. The first stream includes the second layer of tiles for the current viewport position. The new first stream retrieved from the local storage device includes the first layer of tiles for the changed viewport position. And the second stream includes the second layer of tiles for the changed viewport position.

20. The computer-readable medium according to claim 18, wherein the encoded data of the first stream includes data encoded in a first projection format, and the data encoded in the second stream includes data encoded in a second projection format, and the instructions further cause: To select the second projection format based on the changed viewport position.

21. The computer-readable medium according to claim 18, wherein: Encode the encoded data of the first stream and the second stream according to a hierarchical coding protocol. The first stream includes the enhancement layer of tiles for the current viewport position. The first stream retrieved from the local storage device includes the base layer. And the second stream includes the enhancement layer of tiles for the changed viewport position.

22. A video receiving method, comprising: Receive from a source a first stream of encoded data representing each tile of a plurality of tiles of multi-view video, the first stream including a steady-state stream representing a first tile encoded at a first quality layer and transition streams respectively representing other tiles among the plurality of tiles, the transition streams including a first transition stream encoded as a base layer of a second quality layer, wherein each tile corresponds to a predetermined spatial region of the multi-view video, the current viewport position at the receiver includes at least the first tile, the steady-state stream is a non-enhancement layer encoded independently of other layers, and the quality of the first quality layer is higher than the quality of the second quality layer; Decode the first quality layer of the first tile in the steady-state stream; Display the decoded first layer of the current viewport position; When the viewport position at the receiver changes to include a second tile among the other tiles: Decode the base layer of the first transition stream corresponding to the second tile, and Transmit an indication of the second tile to the source; In response to the transmission, a second transitional stream of encoded video data for the second tile is received from the source, the second transitional stream being encoded as an enhancement layer for the base layer of the first transitional stream corresponding to the second tile, wherein when the enhancement layer is decoded and combined with the base layer, the enhancement layer corresponds to a third quality layer higher than the second quality layer; decode the enhancement layer of the second tile from the second transitional stream corresponding to the second tile using the base layer of the second tile from the first transitional stream; and when the viewport position at the receiver changes to include the second tile, display all decoded first and second transitional streams for the second tile, and thereafter, retrieve, decode and display the steady-state stream for the second tile encoded at the first quality layer.

23. The video receiving method according to claim 22, wherein before the viewport position is changed to include the second tile, the base layer of the second tile is received and stored in a local buffer.

24. The video receiving method according to claim 22, wherein: The encoded data of the first stream and the second transitional stream includes data encoded at a first, lower quality layer and a second, higher quality layer, the first stream includes the second layer for the tiles of the current viewport position, the new first stream retrieved from the local storage device includes the first layer for the tiles of the changed viewport position, and the second transitional stream includes the second layer for the tiles of the changed viewport position.

25. The video receiving method according to claim 22, wherein the encoded data of the first stream includes data encoded in a first projection format, and the data encoded in the second transitional stream includes data encoded in a second projection format, and the video receiving method further comprises: Select the second projection format based on the changed viewport position.

26. The video receiving method according to claim 22, wherein: Encode the encoded data of the first stream and the second transitional stream according to a hierarchical coding protocol, the first stream includes an enhancement layer for the tiles of the current viewport position, the first stream retrieved from the local storage device includes a base layer, and the second transitional stream includes an enhancement layer for the tiles of the changed viewport position.

27. The video receiving method according to claim 26, wherein: A first subset of the frames of the enhancement layer of the second transitional stream is predicted only from the reconstructed base layer frames; and A second subset of the frames of the enhancement layer of the second transitional stream is predicted from the reconstructed frames of both the base layer and the enhancement layer; and Decoding of the second transitional stream starts at the frame time corresponding to the frames in the first subset of frames.

28. The video receiving method according to claim 26, wherein decoding the second transitional stream starts on a frame with an unavailable prediction reference, and the video receiving method further comprises: Mitigate quality drift caused by lost prediction references using error concealment techniques.

29. The video receiving method according to claim 26, wherein: The base layer is encoded in a first projection format, the enhancement layer is encoded in a second projection format, and the video receiving method further includes: Predict enhancement layer data from the reconstructed base layer data by applying a function that converts from the first projection format to the second projection format.

30. A computer program product, comprising a computer program which, when executed by a processor, causes the processor to execute the video receiving method according to any one of claims 1-9 or 22-29.

Citation Information

Patent Citations

  • Methods and apparatus for streaming content

    US20150249813A1

  • Systems and methods of generating and processing files for partial decoding and most interested regions

    US20180103199A1