Panoramic video streaming from multiple viewpoints

By streaming shared and viewpoint-specific video content and leveraging viewpoint redundancy, the high bandwidth issue of multi-view panoramic video is solved, enabling more efficient viewpoint switching and encoding while reducing storage requirements.

CN116325769BActive Publication Date: 2025-10-28KONINK KPN NV +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180067743.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-01
Filing Date
2021-09-28
Publication Date
2025-10-28
Estimated Expiration
2041-09-28

AI Technical Summary

Technical Problem

Existing technologies have excessively high bandwidth requirements when streaming multi-view panoramic video, which is detrimental to the encoding, decoding and storage of video data, and the efficiency of seamless switching between viewpoints is low.

Method used

By streaming shared video content and viewpoint-specific video content, and utilizing the redundancy between viewpoints, a portion of the first video data is temporarily continued to be streamed, while only viewpoint-specific video data is streamed, reducing bandwidth requirements. Metadata is used to instruct the shared video data to reconstruct the second panoramic video on the streaming client.

Benefits of technology

It reduces the bandwidth requirements of streaming, improves the switching speed and encoding efficiency between viewpoints, reduces storage requirements, and achieves faster viewpoint switching and higher encoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116325769B_ABST
    Figure CN116325769B_ABST
Patent Text Reader

Abstract

A streaming server and a streaming client are described. The streaming server is configured to stream video data representing a panoramic video scene to the streaming client. The panoramic video scene can be shown from one of multiple viewpoints within the scene. When streaming first video data of a first panoramic video and when streaming second video data of a second panoramic video begins, the streaming client can at least temporarily and simultaneously perform the following operations: continue streaming at least a portion of the first video data by continuing to stream a portion of shared video content representing the scene that is visible in both the first and second panoramic videos, while additionally streaming second viewpoint-specific video data representing viewpoint-specific video content of the scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a streaming server and a computer-implemented method for streaming video data to a streaming client, wherein the video data represents a panoramic video of a scene. The invention further relates to a streaming client and a computer-implemented method for receiving video data representing a panoramic video of a scene. The invention further relates to a creation system and a computer-implemented method for creating one or more video streams representing a panoramic video of a scene. The invention further relates to a computer-readable medium including data representing a computer program, and to a computer-readable medium including metadata. Background Technology

[0002] It is known to capture panoramic video of a scene and display it to a user. Here, the adjective "panoramic" can refer to a video that provides an immersive experience when displayed to a user. Generally, a video can be considered "panoramic" if it offers a wider field of view than the human eye (approximately 160° horizontally and 75° vertically). Panoramic videos can even provide a larger view of the scene, such as a full 360 degrees, thus providing a more immersive experience for the user. Such panoramic videos can be acquired by filming real-life scenes with cameras (such as 180° or 360° cameras) or by being synthesized as so-called computer-generated imagery (CGI) ("3D rendering"). Panoramic videos are also called (hemispherical) videos. Videos that provide at least 180° horizontal and / or 180° vertical views are also called "omnidirectional" videos. Therefore, omnidirectional video is a type of panoramic video.

[0003] Typically, panoramic videos can be two-dimensional (2D) videos, but they can also be three-dimensional (3D) videos, such as stereoscopic videos or volumetric videos.

[0004] Panoramic videos can be displayed in various ways, such as using various types of displays, including head-mounted displays (HMDs), holographic displays, and curved or other types of displays that provide an immersive experience, including but not limited to large-screen or multi-screen displays, such as CAVE or IMAX theater displays. Panoramic videos can also be rendered in virtual environments using virtual reality (VR) or augmented reality (AR) technologies. Panoramic videos can also be displayed using displays that are not typically known to the public to provide an immersive experience, such as on the displays of mobile devices or computer monitors. In such examples, users can still achieve a degree of immersion by being able to look around in the panoramic video, for example, by controlling the position of the viewport (through which a portion of the panoramic video is viewed on the display).

[0005] It is known to acquire different panoramic videos of a scene. For example, different panoramic videos can be captured at different spatial locations within a scene. Therefore, each spatial location can represent a different viewpoint within the scene. Examples of scenes are the interior of a building, or outdoor locations such as a beach or park. A scene can also consist of several locations, such as different rooms and / or different buildings, or a combination of interior and exterior locations.

[0006] It is known that users can choose between different viewpoints. This choice of different viewpoints can effectively allow users to "teleport" within a scene. If the viewpoints are spatially close enough, and / or if there is a rendering transition between different viewpoints, this teleportation can convey to the user a sense of near-continuous movement within the scene.

[0007] Such panoramic videos can be streamed to streaming clients as corresponding video streams. A key challenge is achieving seamless, or at least rapid, switching between video streams presenting different viewpoints so that the user experience is not significantly interrupted during transitions.

[0008] Corbillion et al. [1] described a multi-view (MVP) 360-degree video streaming system in which a scene is captured simultaneously by multiple omnidirectional cameras. Users can only switch their position to a predefined viewpoint (VP). The video stream can be encoded and streamed using MPEG-based HTTP Dynamic Adaptive Streaming (DASH), which provides multiple representations of the same content at different bit rates and resolutions. The video stream can be further encoded and streamed as a Motion Constrained Tile Set (MCTS) [2], in which rectangular video can be spatially divided into rectangular, independently decodeable, non-overlapping regions, known as “tiles”.

[0009] Corbillion et al. further described the ability to predict the user's next translation movement, allowing the client to decide whether to request a representation in the current viewpoint if it anticipates no movement in the near future, or to request a representation in another viewpoint if it anticipates movement soon. They further described the ability to predict the user's next head orientation rotation, enabling the client to request tiles in various predicted locations of a frame within a given viewpoint.

[0010] Therefore, Corbillion et al. addressed the problem of seamless switching between different viewpoints by having the client predict the user's movement within the scene and the predicted viewpoints already streamed. Since this could require very high bandwidth, the client could combine MPEG-DASH with streaming only specific tiles, allowing only those tiles the user is currently viewing or expects to view in the current or adjacent viewpoints to be streamed.

[0011] The drawback described by Corbillion et al. is that the bandwidth requirements of streaming clients, and conversely, the bandwidth requirements of streaming servers, may still be too high. This can be detrimental to streaming, and in some cases, to the encoding, decoding, and / or storage of video data associated with the viewpoint of the scene.

[0012] References

[0013] [1] Xavier Corbillon, Francesca De Simone, Gwendal Simon and Pascal Frossard, 2018. Dynamic adaptive streaming for multi-viewpoint omnidirectional videos, Proceedings of the 9th ACM International Multimedia Conference (MMSys'18), New York Computer Association, 237-249.

[0014] [2] K. Misra, A. Segall, M. Horowitz, S. Xu, A. Fuldseth and M. Zhou, An Overview of Tiles in HEVC, IEEE Signal Processing Special Issues, 7(6):969-977, December 2013. Summary of the Invention

[0015] When streaming multi-view video, it may be desirable to further reduce bandwidth requirements or have additional ways to reduce bandwidth requirements while continuing to achieve seamless or at least fast switching between different viewpoints.

[0016] In a first aspect of the invention, a computer-implemented method for streaming video data to a streaming client can be provided. The video data may represent a panoramic video of a scene. The panoramic video can show the scene from a viewpoint within the scene. The viewpoint may be one of multiple viewpoints within the scene. The method may include the following operations performed by a streaming server:

[0017] The first video data, which can represent a first panoramic video from a first viewpoint within the scene, is streamed to the streaming client.

[0018] In response to the decision to stream second video data representing a second panoramic video from a second viewpoint within the scene.

[0019] The second video data is streamed to the streaming client by performing at least temporarily and simultaneously the following operations in at least one operating mode of the streaming server:

[0020] - Continue streaming at least a portion of the first video data by continuing to stream shared video data representing shared video content of the scene, wherein the shared video content may include video content of the scene that is visible in the first panoramic video and also visible in the second panoramic video; and

[0021] - Streaming can represent second viewpoint-specific video data of second viewpoint-specific video content of the scene, wherein the second viewpoint-specific video content may include video content of the scene that can be part of the second panoramic video rather than part of the shared video content of the scene.

[0022] In another aspect of the invention, a streaming server for streaming video data to a streaming client can be provided. The streaming server may include:

[0023] The network interface to the network, where data communication can reach the streaming client via the network;

[0024] A processor subsystem that can be configured to perform the following operations via a network interface:

[0025] The first video data, which can represent a first panoramic video from a first viewpoint within the scene, is streamed to the streaming client.

[0026] In response to the decision to stream second video data representing a second panoramic video from a second viewpoint within the scene, the second video data is streamed to the streaming client by at least temporarily and simultaneously performing the following operations:

[0027] - Continue streaming at least a portion of the first video data by continuing to stream shared video data representing shared video content of the scene, wherein the shared video content may include video content of the scene that is visible in the first panoramic video and also visible in the second panoramic video; and

[0028] - Streaming can represent second viewpoint-specific video data of second viewpoint-specific video content of the scene, wherein the second viewpoint-specific video content may include video content of the scene that can be part of the second panoramic video rather than part of the shared video content of the scene.

[0029] In another aspect of the invention, a computer-implemented method for receiving video data can be provided. This method may include the following operations performed by a streaming client:

[0030] Receive first video data that can represent a first panoramic video from a first viewpoint within the scene via streaming;

[0031] When switching to the second viewpoint, second video data is received via streaming from the second panoramic video of the second viewpoint within the scene;

[0032] Receiving second video data via streaming may include performing the following operations at least temporarily and simultaneously in at least one operating mode of the streaming client:

[0033] - Continue receiving at least a portion of the first video data via streaming by continuing to receive shared video data that can represent shared video content of the scene, wherein the shared video content may include video content of the scene that is visible in the first panoramic video and visible in the second panoramic video field.

[0034] - Receive second viewpoint-specific video data via streaming, which can represent second viewpoint-specific video content of the scene, wherein the second viewpoint-specific video content can include video content of the scene that can be part of the second panoramic video rather than part of the shared video content of the scene.

[0035] In another aspect of the invention, a streaming client for receiving video data via streaming transmission can be provided. This streaming client may include:

[0036] The network interface to the network;

[0037] A processor subsystem, configured to perform the following operation via a network interface: receiving first video data representing a first panoramic video from a first viewpoint within the scene via streaming;

[0038] When switching to the second viewpoint, second video data from the second panoramic video within the scene is received via streaming. Receiving the second video data via streaming may include at least temporarily and simultaneously performing the following operations:

[0039] - Continue receiving at least a portion of the first video data via streaming by continuing to receive shared video data that can represent shared video content of the scene, wherein the shared video content may include video content of the scene that is visible in the first panoramic video and visible in the second panoramic video field.

[0040] - Receive second viewpoint-specific video data via streaming, which can represent second viewpoint-specific video content of the scene, wherein the second viewpoint-specific video content can include video content of the scene that can be part of the second panoramic video rather than part of the shared video content of the scene.

[0041] The aforementioned measures may involve making video data from multiple viewpoints within a scene available for streaming from a streaming server to a streaming client. The streaming server may be, for example, one or a combination of content delivery nodes (CDNs), and the streaming client may be, for example, a network node, such as an edge node of a mobile or fixed-line network, or an end-user device, such as a mobile device, computer, VR / AR device, etc. Through these measures, video data representing a first panoramic video can be streamed from the streaming server to the streaming client. The first panoramic video can depict the scene from a first viewpoint within the scene. The video data may also be referred to elsewhere as "first" video data to indicate that the video data represents a "first" panoramic video from a "first" viewpoint. Elsewhere, numerical adjectives such as "second," "third," etc., may identify other viewpoints and corresponding panoramic videos and video data, but do not imply any other limitations not described.

[0042] The first video data can be streamed from the streaming server to the streaming client at a given moment, for example, by the streaming client having previously requested the streaming of the first video data. At a later moment, it may be desirable to switch to a second viewpoint within the scene. For example, a user may have already selected a second viewpoint, such as by operating a user input device on an end-user device, or the second viewpoint may have been automatically selected, for example, by the streaming client using predictive techniques to predict that the user may soon select a second viewpoint. Accordingly, the streaming client can decide to stream second video data representing a second panoramic video from the second viewpoint within the scene, and the streaming server can receive the corresponding streaming request from the streaming client. Alternatively, another entity can decide that the streaming server streams the second panoramic video to the streaming client; this other entity may be the streaming server itself or the entity orchestrating the streaming session. Typically, in response to such a decision, the second video data can then be streamed to the streaming client. This may involve at least temporarily continuing to stream at least a portion of the first video data while additionally streaming video data specifically associated with the second viewpoint. The portion of the first video data that can continue to be streamed can be referred to as “shared” video data, while the additional streamed video data can be referred to as “second viewpoint-specific video data,” as will be explained below.

[0043] The inventors have recognized that redundancy can exist between panoramic videos of the same scene. That is, at least a portion of the video content shown in one panoramic video can also be shown in panoramic videos from adjacent viewpoints. This is because the video content can represent objects or portions of the scene visible from both viewpoints. A concrete example is a portion of the sky in an outdoor scene; most of the sky remains visible in viewpoints along the way as one moves between viewpoints. The inventors have recognized that, typically, some portions of the scene may be occluded and de-occluded when transitioning between different viewpoints, but that portion of the scene can remain visible and can only move or change size between viewpoints. In particular, if the spatial distance between viewpoints within the scene is relatively small compared to the distance from the camera acquiring the scene to these portions of the scene, such movement and size change can be (very) limited, meaning that a potentially large portion of the scene may only move or change size slightly when transitioning between viewpoints.

[0044] Therefore, redundancy may exist between panoramic videos. Conversely, this could mean that in the method of Corbillion et al. [1], where adjacent viewpoints may have already been streamed to the streaming client, several video streams could be streamed simultaneously to the streaming client if it is expected that the user can move to adjacent viewpoints. These video streams may include redundant video content. This could mean that bandwidth requirements may be unnecessarily high, even when combined with tile streaming, in which only a subset of the tiles of the panoramic video can be streamed.

[0045] In principle, utilizing redundancy between videos is known. For example, Su et al. [3] proposed using HEVC Multi-View Extension (MV-HEVC) to redundantly encode texture and depth information across different viewpoints. However, Su et al. proposed grouping all viewpoints into a single encoded bitstream containing different layers, where each layer is a viewpoint and these viewpoints have encoding dependencies. However, this may not be suitable for streaming scenarios but rather for local file playback, as the bandwidth required for streaming a single encoded bitstream can be very high, especially with many different viewpoints. This may not only be detrimental to streaming but also to the encoding, decoding, and / or storage of video data associated with the viewpoints of the scene.

[0046] The inventors have further recognized that at least a portion of the first video data can be effectively “reused” to reconstruct the second panoramic video on the streaming client, at least on a temporary basis. That is, when switching to the second viewpoint, the video content of the scene that is visible in both the first and second panoramic videos can still be obtained from the first video data. However, due to the aforementioned occlusion or de-occlusion between viewpoints, the second panoramic video cannot be fully reconstructed from the first video data. Accordingly, the above measures can make viewpoint-specific video data available for streaming, which can contain viewpoint-specific video content of the scene, in this case, for example, video content visible in the second panoramic video, rather than a portion of shared video content between the two viewpoints, for example, due to it being invisible or insufficiently visible in the first panoramic video. This may mean that when switching to the second viewpoint, additional streaming of the second viewpoint-specific video data representing the viewpoint-specific video content of that viewpoint may be sufficient, or at least when requesting the second panoramic video or when it is otherwise decided to stream said video, the streaming server may not need to stream the second panoramic video immediately and completely.

[0047] In practice, the second video data may at least temporarily consist of a portion of the first video data and additional streamed viewpoint-specific video data, and therefore may not represent the overall encoding of the second panoramic video.

[0048] This can have various advantages. For example, if the request for the second viewpoint is a prefetch request, such as when the user is expected to select a second viewpoint, it may not be necessary to stream the entire first and second panoramic videos simultaneously; instead, streaming specific portions of the first and second panoramic videos from the viewpoints may suffice. Similar to a prefetch request, the selection of the second viewpoint can be a preselection by the streaming client when the user is expected to select a second viewpoint. This can be advantageous if several panoramic videos or tiles of such videos have already been streamed simultaneously when the user is expected to select one of the corresponding viewpoints. Here, the bandwidth reduction can be considerable compared to streaming each panoramic video as a complete and independent video stream.

[0049] Even if the request is sent only after the user selects a second viewpoint, there may be other advantages. For example, switching to a second viewpoint can be faster. That is, the second viewpoint-specific video content may be limited in data size because it may only represent a portion of the second panoramic video, not the entirety. This means that shorter segments can be encoded while still achieving the given bandwidth of the specific panoramic video. Such shorter segments can correspond to random access points with more dense intervals and can correspond to shorter GOP sizes. This could mean that after selecting a second viewpoint, the streaming client can switch to the second viewpoint faster by being able to start decoding the second viewpoint-specific video data from any upcoming random access point. Conversely, because shared video content can be reused across several viewpoints, and starts and stops less frequently on average during a streaming session, longer segments (corresponding to less densely spaced random access points and longer GOP sizes) can be encoded to improve encoding efficiency and compensate for the shorter segments of the viewpoint-specific video content.

[0050] Various other advantages are envisioned. For example, shared video data can be encoded and / or stored only once for two or more viewpoints. This can reduce the computational complexity of encoding and / or storage requirements.

[0051] It is understood that the first video data may have already been streamed in the same or similar manner as the second video data, as it may include at least two distinct and separately identifiable portions: a shared portion containing video content shared with other viewpoints and a viewpoint-specific portion containing viewpoint-specific video content. Thus, the advantages described above for at least temporarily streaming the second video data in the aforementioned manner can also be applied to the streaming of the first video data, or at least temporarily if the first video data is only temporarily streamed in this manner, the same applies to the streaming of video data from any viewpoint and the switching between viewpoints.

[0052] It is understood that switching to a second panoramic video by reusing video data from the first panoramic video can be a feature provided in a specific operating mode of the streaming client and / or streaming server. In other words, the viewpoint does not necessarily need to be switched in the described manner; rather, this switching can be an optional operating mode for either entity.

[0053] Typically, the measures described in this specification can be applied in the context of rendering scenes in virtual reality (VR), augmented reality (AR), and mixed reality (MR), all three of which are also known as extended reality (XR).

[0054] The following embodiments may relate to methods for streaming video data and streaming servers. However, it is understood that any embodiment defining how data is sent to a streaming client or what type of data is sent to a streaming client is reflected in an embodiment in which the streaming client is configured to receive, decode, and / or render such data and provides corresponding methods for receiving, decoding, and / or rendering such data.

[0055] In an embodiment of the method for streaming video data,

[0056] - Streaming the first video data may include: streaming first viewpoint-specific video data representing first viewpoint-specific video content of a scene, wherein the first viewpoint-specific video content includes video content of the scene that may be part of the first panoramic video rather than part of shared video content of the scene; and streaming a shared video stream including shared video data;

[0057] - Streaming of the second video data may include at least temporarily continuing the streaming of the shared video stream.

[0058] In a corresponding embodiment, the processor subsystem of the streaming server can be configured to perform streaming in the manner described above.

[0059] According to these embodiments, the first video data can be streamed as at least two separate video streams (i.e., a shared video stream including shared video data and a viewpoint-specific video stream including viewpoint-specific video data of the corresponding panoramic video), wherein at least the shared video stream is temporarily continued to be streamed while the second video data is being streamed. In some embodiments, where the first and second video data are streamed simultaneously, for example when a user is expected to select a second viewpoint, the first and second video data can be streamed as a shared video stream, a first viewpoint-specific video stream, and a second viewpoint-specific video stream. Compared to streaming two video streams, each fully representing the corresponding panoramic video, the above-described streaming can require less bandwidth. Furthermore, for example, in response to a user actually selecting a second viewpoint for display, the streaming of the first viewpoint-specific video stream can be simply stopped independently of the shared video stream.

[0060] Typically, multiple panoramic videos can be represented by one or more shared video streams containing shared video data shared between different panoramic videos, and by viewpoint-specific video streams for each of the multiple panoramic videos. Different encoding properties can often be used for shared video streams instead of viewpoint-specific video streams. For example, viewpoint-specific video streams can be encoded with more random access points, different encoding quality, etc., compared to shared video streams.

[0061] In an embodiment of the method for streaming video data, streaming the first video data may include streaming the shared video data as a video stream in which independently decodeable portions are included.

[0062] In a corresponding embodiment, the processor subsystem of the streaming server can be configured to perform streaming in the manner described above.

[0063] According to these embodiments, the shared video data can be an independently decodeable portion of the first video stream. For example, the shared video data can be represented by one or more tiles of a tile-based video stream. This allows the streaming client to stop decoding portions of the first video stream other than the shared video data when switching to a second video stream. Thus, the computational complexity of decoding can be reduced compared to the case where the shared video data is part of the video stream, which is not allowed to be decoded independently; otherwise, it must be fully decoded by the streaming client to obtain the shared video content needed to reconstruct the second panoramic video. In some embodiments, the streaming client can also stop receiving portions of the video data that are not needed to reconstruct the second panoramic video.

[0064] In embodiments of the method for streaming video data, the method may include providing metadata to a streaming client, wherein the metadata may indicate shared video data representing a portion of first video data and a portion of second video data. In a corresponding embodiment, a processor subsystem of a streaming server may be configured to provide the aforementioned metadata.

[0065] In an embodiment of the method for receiving video data, the method may further include: receiving metadata, wherein the metadata may indicate shared video data representing a portion of first video data and a portion of second video data; and using the shared video data for rendering a second panoramic video based on the metadata. In a corresponding embodiment, the processor subsystem of the streaming client may be configured to receive and use the metadata in the manner described above.

[0066] According to these embodiments, a signal can be used to notify a streaming client that shared video data, which can be received at a given time as part of the video data of a first panoramic video, can also represent part of the video data of a second panoramic video. This signal notification can be implemented by the streaming server providing the streaming client with metadata, for example, as part of the video data of the corresponding panoramic video or as part of the general metadata of the streaming session. For example, shared video data can be identified as being "shared" among different panoramic videos in a so-called inventory associated with the streaming session.

[0067] In embodiments of the method for streaming video data, the method may further include streaming shared video data or any viewpoint-specific video data that is spatially segmented and encoded as at least a portion of a panoramic video.

[0068] In a corresponding embodiment, the processor subsystem of the streaming server can be configured to perform streaming in the manner described above.

[0069] It is known that spatial segmentation coding, as described in [2], enables the streaming of specific portions of a panoramic video to a streaming client without the need to concurrently stream the entire panoramic video. That is, such spatial segments can be streamed and decoded individually by the streaming client. This allows the streaming client to retrieve only those segments that are currently displayed or are expected to be displayed in the near future, for example, as described in [1]. This can reduce the concurrent bandwidth to the streaming client or other receiving entity and reduce the computational complexity of decoding for the receiving entity. The techniques described in this specification for making shared video data available can be combined with spatial segmentation coding techniques. For example, shared video content can be encoded as one or more independently decodeable spatial segments of a video stream, such as a panoramic video, or as a separate “shared” video stream. In some embodiments, this avoids the need to stream all shared video data if some shared video content is not currently displayed or is not expected to be displayed. In some embodiments, this allows the streaming client to continue receiving and decoding shared video data by continuing to receive and decode one or more spatial segments of a first panoramic video while stopping receiving and decoding one or more other spatial segments of the first panoramic video. In some embodiments, metadata associated with spatial segmentation coding (including, but not limited to, so-called Spatial Representation Description (SRD) metadata of MPEG-DASH) may be used to identify the spatial location and / or orientation of shared video content relative to panoramic video.

[0070] In embodiments of the method for streaming video data, the streaming of shared video data may include periodically transmitting at least one of the following:

[0071] -Images; and

[0072] - Metadata, which defines the color used to fill the space area.

[0073] The image or color represents at least a portion of the shared video content of at least a plurality of video frames.

[0074] In a corresponding embodiment, the processor subsystem of the streaming server can be configured to perform streaming in the manner described above.

[0075] Shared video data can represent the "background" of a scene in many cases because it can represent a portion of the scene that is farther away than other objects that might be the "foreground." Such a background can be relatively static because it doesn't change in appearance, or only changes to a limited extent. For example, the sky or the exterior of a building can remain relatively static across multiple video frames in a panoramic video. The same applies to video data that isn't necessarily the "background" of a scene, as this type of non-background video data can also remain relatively static across multiple video frames in a panoramic video. To improve encoding efficiency, thereby reducing bandwidth to the streaming client and lowering the computational complexity of decoding, shared video data can be streamed by periodically transmitting images representing the shared video content. This effectively corresponds to temporal subsampling of the shared video data, since only the shared video content in every 2nd, 10th, 50th, etc., video frames can be transmitted to the streaming client. This periodic streaming can also be adaptive, as images can be transmitted if the changes in the shared video content exceed a static or adaptive threshold. Therefore, the term "periodic transmission" implies that the time between image transmissions is of variable length. The image can directly represent the shared video content, or it can contain a texture that represents the shared video content when tiled within a spatial region. In some embodiments, where the shared video content is homogeneous, such as a cloudless sky, the shared video content can take the form of metadata defining colors used to fill spatial regions corresponding to the shared video content. It is also envisioned that the transmission of the image and the metadata defining the fill colors be combined.

[0076] The following embodiments may relate to methods for receiving video data and streaming clients. However, it is understood that any embodiment defining how a streaming client receives or what type of data it receives is reflected in embodiments in which a streaming server is configured to stream and / or generate such data, and corresponding methods are provided for streaming and / or generating such data.

[0077] In this embodiment, the metadata may further identify a second viewpoint-specific video stream containing second viewpoint-specific video data, wherein the second viewpoint-specific video stream can be accessed from a streaming server. In this embodiment, the method of receiving video data may further include: requesting a second viewpoint-specific video stream from the streaming server based on the metadata and when switching to a second viewpoint.

[0078] In a corresponding embodiment, the processor subsystem of the streaming client can be configured to receive and use metadata in the manner described above.

[0079] According to these embodiments, second-viewpoint-specific video data can be provided to the streaming client as a separate video stream, which can be identified to the streaming client via metadata. This allows the streaming client to request a second-viewpoint-specific video stream, for example, when the second viewpoint is selected by the user or automatically by the streaming client, for caching purposes. This allows the streaming client to additionally retrieve only those portions of the second panoramic video that have not yet been provided to the streaming client, such as the second-viewpoint-specific portions, thereby avoiding unnecessary bandwidth and decoding complexity that might be required when the second panoramic video is only available for streaming in its entirety.

[0080] In an embodiment, metadata can be a list of video data streams for a scene. For example, a streaming server can provide a Media Presentation Description (MPD) to a streaming client, which is an example of a so-called manifest or manifest file. The MPD can list and thereby identify different video streams, and in some embodiments, different versions (“representations”) of the video streams are listed and identified, for example, each with a different spatial resolution and / or bitrate. The client can then select a version of the video stream by choosing from the MPD. Thus, any shared video data and / or any viewpoint-specific video data can be identified in the manifest, for example, as a separate video stream having one or more representations.

[0081] In one embodiment, metadata may indicate the transformation to be applied to the shared video content. In this embodiment, the method of receiving video data may further include applying the transformation to the shared video content before or as part of the rendering of the second panoramic video.

[0082] In a corresponding embodiment, the processor subsystem of the streaming client can be configured to use metadata in the manner described above.

[0083] Although the objects representing shared video content may be visible from two different viewpoints, the objects may have different relative positions and / or orientations with respect to the real or virtual camera at each viewpoint. For example, if the shared video content represents the exterior of a building, one viewpoint may be closer to the building or at the same distance but to the side of another viewpoint. Therefore, the appearance of the objects representing the shared video content can vary between viewpoints and, consequently, between panoramic videos. To enable the shared video content to be used for the reconstruction of different panoramic videos, metadata can indicate the transformations to be applied to the shared video content. In some embodiments, the transformations defined in the metadata can be specific to the reconstruction of a particular panoramic video. Typically, this transformation can compensate for the appearance variations of the objects between viewpoints and, therefore, between panoramic videos. For example, the transformation can parameterize the appearance variations as an affine transformation or a higher-order image transformation. This allows the shared video content to be (re)used for the reconstruction of different panoramic videos, despite these appearance variations.

[0084] In this embodiment, the location of the viewpoint can be represented as corresponding coordinates in a coordinate system associated with the scene, wherein metadata defines the coordinate range of the shared video data. In this embodiment, the method of receiving video data may further include using the shared video data to render a corresponding panoramic video of a viewpoint whose coordinates lie within the coordinate range defined by the metadata.

[0085] In a corresponding embodiment, the processor subsystem of the streaming client can be configured to use metadata in the manner described above.

[0086] Shared video content can be used to reconstruct multiple panoramic videos. This can be indicated to the streaming client in various ways, such as using the list mentioned above. Alternatively or alternatively, if the viewpoint can be represented as corresponding coordinates in a coordinate system associated with the scene, such as geographic coordinates or XY coordinates without specific geographic significance, the coordinate range can be communicated to the streaming client. For example, by comparing the coordinates of the viewpoints of the panoramic videos to the coordinate range, the streaming client can then determine whether the shared video data will be used for the reconstruction of a particular panoramic video (and, in some embodiments, which shared video data will be used for the reconstruction of a particular panoramic video). This allows shared video content to be associated with multiple panoramic videos in a way that has a physical analogy. This approach is readily understood because objects are typically visible from multiple viewpoints within a specific coordinate range. In particular, this allows shared video data to be associated with viewpoints unknown to the streaming server, for example, in cases where the streaming client is able to reconstruct an intermediate viewpoint by interpolating between panoramic videos of nearby viewpoints.

[0087] In this embodiment, at least temporarily received shared video data is a first version of shared video content derived from a first panoramic video. In this embodiment, the method of receiving video data may further include receiving second shared video data via streaming, the second shared video data representing a second version of the shared video content derived from a second panoramic video.

[0088] In a corresponding embodiment, the processor subsystem of the streaming client can be configured to receive video data in the manner described above.

[0089] Previously received shared video data from a first viewpoint can be temporarily used for rendering the second panoramic video, but after a period of time it can be replaced by a second version of the shared video content derived from the second panoramic video. This can have various advantages. For example, while the shared video content from the first viewpoint may be sufficient to render the second panoramic video, there may be subtle differences in appearance between the objects(s) representing the shared video content between the first and second viewpoints. For example, the objects(s)(s) may be closer to the real or virtual camera of the second viewpoint, which can allow for the resolution of more details. As another example, in some embodiments, the first version of the shared video content may need to be transformed to be rendered as part of the second panoramic video, while the second version of the shared video content can be rendered as part of the second panoramic video without transformation. By switching to the second version of the shared video content, the shared video content can be shown in a manner visible from the second viewpoint. However, for example, in cases where the frequency of random access points in the shared video data is low per time unit while the frequency of random access points in the viewpoint-specific video data is high, it may be advantageous to temporarily continue streaming the first version of the shared video content compared to immediately requesting, receiving, and switching to the second version of the shared video content. That is, before starting to render the second panoramic video, it may not be necessary to wait for the random access point in the second version of the shared video content; instead, the first version can be used at least temporarily (re) for the rendering.

[0090] In this embodiment, a second version of the shared video content can be received as a video stream, and the reception of the video stream can begin from a streaming access point within the video stream. A streaming access point, which may also be referred to elsewhere as a random access point, allows the streaming client to switch to the second version of the shared video content. For example, the frequency of the random access point in the shared video data within each time unit may be lower than the frequency of the random access point in the viewpoint-specific video data. This means that streaming of the first version of the shared video content can continue until a streaming access point in the second version of the shared video content becomes available, at which point the second version can be used.

[0091] In an embodiment, the second version of the video stream containing shared video content may be a second shared video stream, wherein second viewpoint-specific video data may be received as a second viewpoint-specific video stream, and wherein the second shared video stream includes fewer streaming access points per time unit than the second viewpoint-specific video stream.

[0092] In a corresponding embodiment, the processor subsystem of the streaming client can be configured to receive video data in the manner described above.

[0093] In embodiments of the method for receiving video data, the method may further include combining shared video content with second viewpoint-specific video content by performing at least one of the following operations:

[0094] - Spatially adjoin the shared video content with the video content specific to the second viewpoint;

[0095] -Merge the shared video content with the video content specific to the second viewpoint; and

[0096] - Overlay the second viewpoint-specific video content onto the shared video content.

[0097] In a corresponding embodiment, the processor subsystem of the streaming client can be configured to perform the combination as described above.

[0098] Shared video content can be formatted so that streaming clients or other receiving entities can combine it with viewpoint-specific video content in a specific manner. This combination can be predetermined, for example, standardized, but in some embodiments, metadata may be included together with the video data of the panoramic video via a streaming server and signaled to the receiving entity. By combining these two types of video content, the receiving entity can effectively reconstruct the corresponding panoramic video or at least a portion thereof. For example, if the shared video content and the viewpoint-specific video content are spatially separated, the shared video content can be adjacent to the viewpoint-specific video content. This example relates to the fact that shared video content and viewpoint-specific video content can represent spatially separated portions of the panoramic video. These two types of video content can also be combined by overlaying one type of video content onto the other or by merging the two types of video content in any other way. Different types of combinations can also be used together. For example, if the shared video content is partially spatially separated but overlaps with the viewpoint-specific video content at its boundaries, a blending technique can be used to combine the overlapping portions of the two types of video content.

[0099] In an embodiment of the method for receiving video data, the method may further include rendering second video data to obtain rendered video data for display by a streaming client or another entity, wherein rendering may include combining shared video data and second viewpoint-specific video data.

[0100] In a corresponding embodiment, the processor subsystem of the streaming client can be configured to render video data in the manner described above.

[0101] Rendering can represent any known rendering of panoramic video and may include steps such as applying an inverse equirectangular projection to the received video data if the received video data contains spherical video content, which is then converted into a rectangular image format using the equirectangular projection. Typically, rendering may include rendering the panoramic video only within, for example, the viewport currently displayed to the user. Rendering can be performed, for example, by an end-user device or by an edge node acting as a streaming client. In the former case, the rendered video data can then be displayed by the streaming client, while in the latter case, the rendered video data can be transmitted again to another streaming client, i.e., to the end-user device where it can be displayed.

[0102] In another aspect of the invention, a streaming client for receiving video data via streaming transmission can be provided, wherein the video data may represent a panoramic video of a scene, wherein the panoramic video can show the scene from a viewpoint within the scene, and wherein the viewpoint may be one of a plurality of viewpoints within the scene. The streaming client may include:

[0103] The network interface to the network;

[0104] A processor subsystem that can be configured to perform the following operations via a network interface:

[0105] First video data, representing at least a portion of a first panoramic video from a first viewpoint within the scene, is received via streaming.

[0106] Receive another video data of the scene via streaming;

[0107] Receive metadata indicating that the other video data is shared video data representing shared video content of the scene, wherein the shared video content may include video content of the scene that is visible in the first panoramic video and in the second panoramic video from a second viewpoint within the scene;

[0108] Based on this metadata, the first video data is combined with the shared video data to obtain combined video data for rendering the first panoramic video.

[0109] For example, the first video data may include viewpoint-specific video data of the first panoramic video, which may be combined with shared video data in a manner described elsewhere in this specification. For example, the shared video data may include a version of shared video content derived from the second panoramic video. In a particular, but not limiting, example, a streaming client may switch from rendering the second panoramic video to rendering the first panoramic video and may continue streaming the shared video data derived from the second panoramic video while simultaneously starting to stream the viewpoint-specific video data of the first panoramic video. This may have the advantages set forth above following the introductory sentence “This can have various advantages,” while noting that the advantages described there are for switching from the first panoramic video to the second panoramic video, but also apply to switching from the second panoramic video to the first panoramic video. It is understood that the streaming client may be an embodiment of a streaming client as defined elsewhere in this specification, but may not necessarily be required.

[0110] In another aspect of the invention, a system is provided that includes a streaming server and a streaming client as described in this specification.

[0111] In embodiments of shared video data, the shared video includes at least one of the following:

[0112] The first panoramic video must be spatially adjacent to the first viewpoint-specific video content and / or the second viewpoint-specific video content in order to render a portion of the corresponding panoramic video.

[0113] The first panoramic video should be overlaid on or under the first viewpoint-specific video content and / or the second viewpoint-specific video content for rendering a portion of the corresponding panoramic video; and

[0114] The first panoramic video is a portion of the video content that is spatially separated from the first viewpoint-specific video content and / or the second viewpoint-specific video content.

[0115] In another aspect of the invention, a computer-implemented method for creating one or more video streams representing a panoramic video scene can be provided. This method may include:

[0116] - Access the first panoramic video, which shows the scene from a first viewpoint within the scene;

[0117] - Access a second panoramic video, which shows the scene from a second viewpoint within the scene;

[0118] - The identifier can represent shared video data of shared video content of the scene, wherein the shared video content may include video content of the scene that is visible in the first panoramic video and visible in the second panoramic video;

[0119] - An identifier may represent first-viewpoint-specific video data of first-viewpoint-specific video content for the scene, wherein the first-viewpoint-specific video content may include video content of the scene that can be part of the first panoramic video rather than part of the shared video content of the scene; and

[0120] - Encode the first panoramic video in the following way:

[0121] - Encode the first-viewpoint-specific video data into a first-viewpoint-specific video stream, and encode the shared video data into a shared video stream; or

[0122] - Encode the first viewpoint-specific video data and the shared video data into a video stream, wherein the shared video data can be included in the video stream as an independently decodable part of the video stream.

[0123] In another aspect of the invention, a system for creating one or more video streams representing a panoramic video scene can be provided. This system may include:

[0124] - Data storage interface, which is used for the following operations:

[0125] Access the first panoramic video, which shows the scene from a first viewpoint within the scene;

[0126] Access a second panoramic video, wherein the second panoramic video shows the scene from a second viewpoint within the scene;

[0127] - Processor subsystem, which can be configured to:

[0128] The identifier can represent shared video data of shared video content in the scene, wherein the shared video content can include video content of the scene that is visible in the first panoramic video and also visible in the second panoramic video;

[0129] The identifier can represent first-viewpoint-specific video data of first-viewpoint-specific video content for a scene, wherein the first-viewpoint-specific video content may include video content of the scene that can be part of the first panoramic video rather than part of shared video content of the scene; and

[0130] The first panoramic video is encoded in the following manner:

[0131] - Encode the first-viewpoint-specific video data into a first-viewpoint-specific video stream, and encode the shared video data into a shared video stream; or

[0132] - Encode the first viewpoint-specific video data and the shared video data into a video stream, wherein the shared video data can be included in the video stream as an independently decodable part of the video stream.

[0133] The video data for panoramic video can be created by an authoring system. This creation typically involves identifying viewpoint-specific video data and shared video data, and encoding both types of data, for example, into separate video streams or into a single video stream where the shared video data is a separately decodeable portion. An example of the latter is a tile video stream, where one or more tiles can be independently decoded and streamed as substreams of the tile video stream. In some examples, the authoring system may also generate metadata or portions of metadata as described in this specification.

[0134] In another aspect of the invention, a computer-readable medium may be provided, which may include transient or non-transitory data representing a computer program. The computer program may include instructions for causing a processor system to perform any of the methods described herein.

[0135] In another aspect of the invention, a computer-readable medium may be provided, which may include transient or non-transitory data defining a data structure representing metadata that can identify video data of a video stream as shared video content representing a scene, wherein the metadata may indicate that the shared video content is visible at least in a first panoramic video and a second panoramic video of the scene. This data structure may be associated with the first and second panoramic videos because it may be provided to a streaming client to enable the streaming client to render either panoramic video from video data received from a streaming server, the video data including viewpoint-specific video data and shared video data. In other words, the metadata may identify the shared video data as suitable for reconstruction of the first and second panoramic videos, i.e., by including video content shared between the two panoramic videos. For example, the metadata may include an identifier for the shared video data and an identifier for the corresponding panoramic video, and / or viewpoint-specific video data of the panoramic video, to enable the streaming client to identify the shared video data as shared video content representing the corresponding panoramic video. In some examples, the metadata may define additional properties, such as the scene, the corresponding panoramic video, the shared video data, and / or the corresponding viewpoint-specific video data, as also defined elsewhere in this specification.

[0136] According to the abstract of this specification, a streaming server and a streaming client are described. The streaming server is configured to stream video data representing a panoramic video of a scene to the streaming client. The panoramic video shows the scene from one of a plurality of viewpoints within the scene. When streaming first video data of a first panoramic video and when streaming second video data of a second panoramic video begins, at least temporarily and simultaneously, the following operations can be performed: at least a portion of the first video data can continue to be streamed by continuing to stream a portion of shared video content representing the scene that is visible in both the first and second panoramic videos, while additionally streaming second viewpoint-specific video data representing viewpoint-specific video content of the scene.

[0137] Those skilled in the art will understand that two or more of the embodiments, implementations and / or aspects of the invention mentioned above can be combined in any manner that they deem useful.

[0138] Modifications and changes to any of the systems or devices (e.g., streaming servers, streaming clients), methods, metadata, and / or computer programs (corresponding to the modifications and changes described in another of these systems or devices, methods, metadata, and / or computer programs, and vice versa) may be performed by those skilled in the art based on this specification.

[0139] Other references:

[0140] [3] T. Su, A. Sobhani, A. Yassine, S. Shirmohammadi and A. Javadtalab, A DASH-based HEVC multi-view video streaming system, Journal of Real-Time Image Processing, 12(2):329-342, August 2016 Attached Figure Description

[0141] These and other aspects of the invention will be understood and illustrated with reference to the embodiments described below. In the accompanying drawings:

[0142] Figure 1A and Figure 1B A schematic top view of a scene where omnidirectional video is captured from different viewpoints A and B is shown, wherein the scene includes objects visible from both viewpoints A and B;

[0143] Figure 2A and Figure 2BA visual representation of the video content visible from both viewpoints A and B is shown, as a dome screen covering viewpoints A and B;

[0144] Figure 3 The message exchange between the streaming client and the streaming server is shown for downloading and rendering video data representing the dome screen;

[0145] Figures 4A to 4C This demonstrates different types of overlapping dome screens for different types of scenes, namely... Figure 4A The middle is a corridor and Figure 4B and Figure 4C The center is an open space;

[0146] Figure 5 The message exchange between the streaming client and the streaming server is shown for switching between video data on different dome screens;

[0147] Figure 6A It showcases omnidirectional videos from different viewpoints A, B, and C, reconstructed from viewpoint-specific video data and from shared video data;

[0148] Figure 6B The demonstration shows the switching of viewpoints A, B, and C over time, where the first version of the shared video data continues to be streamed after switching to the new viewpoint. This first version of the shared video data is derived from the omnidirectional video of the previous viewpoint and is replaced by the second version of the shared video data from the omnidirectional video of the new viewpoint.

[0149] Figure 7 It showcases objects with different angles and sizes displayed on different dome screens;

[0150] Figure 8 It demonstrates how a smooth transition can be achieved when moving from viewpoint A to D while switching between two different domes;

[0151] Figure 9 It showed that the dome screen was nested;

[0152] Figure 10 It shows the dome boundaries defined in the coordinate system associated with the scene, indicating which viewpoints the dome is applicable to;

[0153] Figure 11 A streaming server is shown for streaming video data to a streaming client, wherein the video data represents a panoramic video of a scene;

[0154] Figure 12 A streaming client for receiving video data via streaming is shown, wherein the video data represents a panoramic video of a scene;

[0155] Figure 13An authoring system for creating panoramic videos representing a scene is shown;

[0156] Figure 14 A computer-readable medium including non-transient data is shown;

[0157] Figure 15 An exemplary data processing system is shown.

[0158] It should be noted that items with the same reference numerals in different figures have the same structural features and the same function, or the same signal. Since the function and / or structure of such items have already been explained, it is unnecessary to repeat the explanation in the specific implementation.

[0159] List of reference numerals

[0160] The following list of reference numerals and abbreviations is provided to facilitate the interpretation of the drawings and should not be construed as limiting the claims.

[0161] AD Viewpoint

[0162] CL Streaming Client

[0163] DM(X) Dome(X)

[0164] N x Angle size

[0165] OBJ(X) Object(X)

[0166] SRV Streaming Server

[0167] Schematic top view of scenes 100 and 102

[0168] 110 Ground

[0169] 120 Currently selected viewpoint

[0170] Camera at viewpoint A (130°)

[0171] 132 Camera at viewpoint B

[0172] 140 Omnidirectional image acquired from viewpoint A

[0173] 142 Omnidirectional image acquired from viewpoint B

[0174] 150 represents the dome screen displaying the shared video content between viewpoints A and B.

[0175] Transition between 160 viewpoints

[0176] 170 viewpoint

[0177] 172 Time

[0178] 200 streaming servers

[0179] 220 Network Interface

[0180] 222 Data Communication

[0181] 240 Processor Subsystem

[0182] 260 Data Storage

[0183] 300 streaming clients

[0184] 320 network interface

[0185] 322 Data Communication

[0186] 340 Processor Subsystem

[0187] 360 Display Output Terminal

[0188] 362 Display Data

[0189] 380 monitor

[0190] 400 Creation System

[0191] 420 network interface

[0192] 422 Data Communication

[0193] 440 Processor Subsystem

[0194] 460 Data Storage

[0195] 500 computer-readable media

[0196] 510 Non-transient data

[0197] 1000 Exemplary Data Processing Systems

[0198] 1002 processor

[0199] 1004 Memory Element

[0200] 1006 System Bus

[0201] 1008 Local Memory

[0202] 1010 High-capacity storage device

[0203] 1012 Input Devices

[0204] 1014 Output Device

[0205] 1016 Network Adapter

[0206] 1018 Application Detailed Implementation

[0207] The following embodiments relate to streaming video data to a streaming client. Specifically, the video data may represent panoramic video of a scene, showing the scene from a viewpoint within the scene. This viewpoint may be one of multiple viewpoints within the scene. Therefore, multiple panoramic videos of the scene can be used for streaming. For example, these panoramic videos may have been previously acquired by multiple cameras, or may have been previously synthesized, or in some examples may have been generated in real time. An example of the latter is a football match in a football stadium recorded by multiple panoramic cameras corresponding to different viewpoints within the stadium, each viewpoint being selected for streaming.

[0208] If the panoramic video was originally acquired or generated for a specific viewpoint (e.g., via an omnidirectional camera or via offline rendering), such panoramic video can also be called a pre-rendered view area (PRVA). This means that the video content is provided to the streaming client in a pre-rendered manner, rather than the streaming client having to synthesize and generate the video content.

[0209] For example, the following examples assume that the panoramic video is omnidirectional, such as 360° video. However, this is not a limitation, as the measures described with respect to these embodiments are equally applicable to other types of panoramic video, such as 180° video, etc. In this regard, note that panoramic video can be single-field-of-view video, or it can be stereoscopic video or volumetric video, for example, represented by point clouds or grids or sampled light fields.

[0210] A streaming server can be configured to stream video data for a given viewpoint, for example, in response to a request received from a streaming client. As will be explained in more detail below, when streaming video data for another viewpoint (which may also be referred to elsewhere as the "second" viewpoint) begins, that video data can be streamed by continuing to stream at least a portion of the video data for the currently streaming viewpoint (which may also be referred to as the "first" viewpoint). This video data can also be referred to as "shared" video data because it can represent shared video content of the scene visible from both the first and second viewpoints. In addition to shared video data, viewpoint-specific video data, which can represent viewpoint-specific video content of the scene, can be streamed to the streaming client. Viewpoint-specific video content can include video content not included in the shared video data.

[0211] The streaming client can then use the received shared video data and the received viewpoint-specific video data to reconstruct the panoramic video corresponding to the second viewpoint. This reconstruction may include, for example, combining the two types of video data by concatenating or overlaying the shared video data with the viewpoint-specific video data. The simultaneous but separate streaming of the shared video data and the viewpoint-specific video data described above can continue at least temporarily. In some embodiments, the streaming client may later switch to the full version of the panoramic video, or switch to another form of streaming video data.

[0212] The above and the following refer to the similarity between video content. It is understood that this generally refers to the similarity between the image content of corresponding frames of a video, where "corresponding" means that these frames were captured at substantially the same time or generally relate to points of similarity on the content's timeline. General references to video content will be understood by those skilled in the art to include references to the image content of video frames, wherein the references are to the spatial information of the video rather than its temporal information.

[0213] Figure 1A and Figure 1B A top view of scene 100 is shown, showing omnidirectional video captured at different viewpoints A and B. In this and other figures, viewpoints are shown as those identified by their respective identifiers (…). Figure 1A and Figure 1B The circled letters representing viewpoints A and B. Here and elsewhere, triangle 120 indicates the currently selected viewpoint; in this example, the triangle indicates that video data for viewpoint A can be streamed to a streaming client.

[0214] exist Figure 1A and Figure 1B In this context, scene 100 is schematically shown as a grid that can be addressed using a coordinate system. This grid-like representation can reflect that, in some examples, cameras recording the scene can be arranged in a grid-like manner in physical space, with the corresponding locations represented as coordinates within the grid. For example, these grid coordinates could correspond to actual geographical locations, but could alternatively represent coordinates without a direct physical analogy. Figure 1A and Figure 1B In the example, viewpoint A can therefore have coordinates (1,1) in the scene, while viewpoint B can have coordinates (7,1) in the scene.

[0215] from Figure 1B As can be seen from this, there can be objects in the scene, and these objects are... Figure 1BThis object is marked "OBJ" elsewhere. It is visible from both viewpoints, the main difference being that, because viewpoint B is closer to the object than viewpoint A, the object appears larger from viewpoint B than from viewpoint A. In other words, the object may appear larger in the video data at viewpoint B than in the video data at viewpoint A, but it is visible from both viewpoints. This example illustrates that a portion of a scene can be visible in several omnidirectional videos of that scene. It can be understood that in some examples, most of the scene may be visible in several omnidirectional videos, such as the sky.

[0216] It may be desirable to avoid streaming redundant video data in omnidirectional video, and in some cases, to avoid encoding, storing, and / or decoding it. In some examples, a portion of a scene may appear identical in several omnidirectional videos, not only in size but also in location. This means the video content in several omnidirectional videos can be substantially the same and thus can be considered “shared” video content. This is especially likely if an object is so far away from the camera that a movement of the viewpoint does not result in a noticeable movement of the object. Similarly, this may be the case for objects in the sky or distant locations, such as mountains or the skyline. In other examples, a portion of a scene may be visible in several omnidirectional videos, but its appearance may change between viewpoints (e.g., in terms of size, position, and / or orientation). Such video content can also represent “shared” video content, as will be explained elsewhere. In the examples above, redundancy between viewpoints can be taken advantage of by avoiding streaming redundant versions of shared video content.

[0217] Figure 2A and Figure 2B Shown Figure 1A and Figure 1B A cross-sectional view of the scene, showing ground 110, camera A 130, and camera B 132, each camera acquiring video data omnidirectionally from its respective viewpoint, such as... Figure 2A and Figure 2B The dome screens 140 and 142 indicate the omnidirectional acquisition of each camera. Figure 2A and Figure 2B The visual representation of the video content visible from both viewpoints A and B is further illustrated, as shown by the overall dome screen 150 covering viewpoints A and B (in Figures 2A to 2B (marked as "DM" in Chinese).

[0218] The analogy of a "spherical screen" can be understood as follows: Consider how, in many cases, omnidirectional video can be displayed by projecting acquired video data onto the interior of a spherical screen or similar object and placing a virtual camera within the spherical screen. Thus, omnidirectional video, or a portion thereof, can be represented as a virtual spherical screen for this reason. Furthermore, as will be explained elsewhere, shared video data in many cases can contain video data of objects far from the camera, such as the sky, a city skyline, distant trees, etc. For these distant objects, it is obvious that movement between viewpoints may result in these objects having little or no parallax. In effect, to the user, such video data may appear to be projected onto a large, overall virtual spherical screen. Again for this reason, shared video data can be referred to as a spherical screen below and is visually represented as a spherical screen in the accompanying drawings; therefore, video data defining a spherical screen is an example of shared video data.

[0219] It is understandable that, although the following text may refer to the encoding, streaming, decoding, rendering, etc. of the dome screen, it can be understood as referring to the encoding, streaming, decoding, rendering, etc. of the video data of the dome screen.

[0220] Continuing with the "dome" analogy, shared video content can also be considered as representing a (partial) dome. For example, if the shared video data involves the sky, then that video data can be represented as a hemispherical dome. Understandably, the adjective "partial" may be omitted below, but it's important to understand that a dome representing shared video data may only contain a portion of the omnidirectional video data, such as one or more spatial regions.

[0221] Although the sky is frequently referred to as an example of shared video content below, it is understandable that shared video data can also involve objects that are relatively close to the camera but visible from at least two different viewpoints. Figure 2B An example of such an object is shown, which demonstrates Figure 1B The object OBJ in the image is visible from viewpoints A and B but has a different size in each acquired omnidirectional video, as indicated by an object with an angular size N2 in the field of view of camera B, which is larger than the angular size N1 in the field of view of camera A. It can be further understood that, for example, in the case of the sky in an outdoor environment, the shared video data may not cover or need to cover most of the corresponding omnidirectional video, but may relate to one or more smaller spatial regions (e.g., representing smaller individual objects). Thus, the shared video data may consist of several separate spatial regions, or, as described in the embodiments below, may represent one of the spatial regions. Therefore, referring to the shared video data as a “dome” should not be construed as limiting, but rather as an intuitively understandable representation of common types of shared video data.

[0222] Basic dome screen

[0223] Once the scene content has been captured or synthesized and the locations of the omnidirectional cameras (real or virtual) in space are known, the dome can be generated as shared video data or as metadata identifying the shared video data. Dome creation can utilize image recognition techniques, such as simply comparing image data from video frames acquired by camera A with image data from video frames acquired by camera B, to identify pixels (or voxels or other image elements) representing shared video content frame by frame. A very simple algorithm can subtract video data from camera A pixel by pixel, and a pixel can be considered to belong to shared video data when the difference is below a certain threshold. More complex algorithms can find correspondences between video frames that take into account variations in appearance such as size, position, and / or orientation. Finding correspondences between video frames is well-known in the general field of video analysis, such as the motion estimation subfield for estimating correspondences between temporarily adjacent video frames, the disparity estimation subfield for estimating correspondences between left and right video frames from a stereo pair, or the (elastic) image registration subfield, etc. Such techniques can also be used to find image data believed to be shared between video frames acquired from different viewpoints of a scene. As described elsewhere, not only can correspondences be determined, but they can also be encoded as, for example, metadata, enabling streaming clients to reconstruct the appearance of shared video data within a specific viewpoint. In a particular example, correspondences can be estimated by estimating one or more affine transformations between image portions of different videos. These affine transformations (multiple transformations) can then be signaled to the streaming client as metadata.

[0224] Any algorithm used to identify shared video content can preferably be robust to camera-captured noise, for example, by using techniques to increase noise robustness, as is known per se. Furthermore, different cameras can be synchronously locked to acquire video frames from different cameras at the same time. If the cameras are not synchronously locked, the algorithm might assume the video frame rate is high enough that noticeable movement in the scene can be considered negligible between the closest video frames captured by different cameras at any given time.

[0225] Dome creation typically generates a video stream for each viewpoint, containing viewpoint-specific video content, and one or more video streams represent corresponding domes, each containing shared video content between two or more viewpoints. In a streaming client, the viewpoint-specific video content and one or more domes can then be used to reconstruct the panoramic video, thereby reconstructing the PRVA for that specific viewpoint.

[0226] Information about the location of the dome can be included in the inventory file, for example, as described elsewhere in this specification, such as in a Media Presentation Descriptor (MPD). Here, the term "location" can refer to information that allows the dome to be associated with one or more PRVAs and corresponding viewpoints. Furthermore, the inventory file may include or indicate network locations through which the corresponding dome can be retrieved.

[0227] Another example of creating dome-shaped and viewpoint-specific video streams is using depth information. That is, such depth information can be used to calculate the distance from the optical center of each viewpoint to a point in the scene represented by pixels (or voxels or other image elements) in the video data. Pixel regions can then be clustered across several viewpoints based on this depth information. For example, a dome can be generated by clustering pixels at equidistant distances from each viewpoint. In a specific example, pixels at infinity could be considered to belong to the shared video content "sky."

[0228] Continue to refer to Figure 2A and Figure 2B A viewpoint-specific video stream can be generated for viewpoint A, and another viewpoint-specific video stream can be generated for viewpoint B. Each video stream omits the pixels of the sky, but the omitted pixels are transmitted once for both video streams using a shared video stream representing the DM 150 dome. In this example, individual viewpoint-specific video streams can be generated by horizontally dividing the video into equal rectangular sections, meaning dividing the equal rectangular video into a top and bottom section, thus creating the top section of the dome and the viewpoint-specific bottom section. The video data of the dome itself can also be projected using equal rectangular projection, or typically any other type of projection, through which omnidirectional video data can be transmitted in rectangular video frames.

[0229] In another example, the sky, as an example of a background object visible from different viewpoints, might be partially occluded by objects whose positions shift between viewpoints. This could be the case, for example, when the scene is an interior scene in a building with a glass ceiling, where the sky is visible through the glass ceiling and beams of light partially obscure it. In such an example, a spherical screen can be generated by selecting sky pixels visible from both viewpoints and projecting these pixels onto a new spherical screen, for example, using an equal rectangular projection. Pixels specific to each viewpoint can be omitted from this spherical screen. Then, video streams for viewpoints A and B can be generated, for example, by masking the pixels added to the spherical screen with an opaque color (e.g., green). In effect, this masking "removes" shared video content from the original panoramic video, thus generating viewpoint-specific video.

[0230] Shared video content can also be encoded as independently decodeable spatial segments, such as "tiles." Metadata describing the tile locations can be signaled to the streaming client in various ways (e.g., in filenames or as properties within the MPD). The streaming client can download these tiles at specific points in time, determine their location on the dome, and render them accordingly. This tile-based encoding and rendering can be beneficial if the spatial region representing the shared video content is relatively small but has high resolution.

[0231] In some examples, the dome can also be generated as a still image, or may include a still image that can, for example, represent at least a portion of shared video content of at least a plurality of video frames by indicating duration. In some examples, the image can contain textures that can be tiled to fill spatial regions. In some examples, the dome can also be generated as a spatial region filled with color, or may include a spatial region filled with color. Metadata can describe the color and the spatial region in the dome to be filled with that color. This can be defined in the manifest file, for example, as follows:

[0232]

[0233] • Region: A region can be defined that includes at least three points on the dome, each defined by x and y values. At least three points may be required to define the shape. The x and y values ​​can be separated by the following:

[0234] • PTS, short for "Presentation Timestamp": This displays the presentation time of this dome screen.

[0235] • Duration: The displayed duration, in seconds (decimal).

[0236] • Color: Colors can be defined using RGB, HSL, hexadecimal, etc.

[0237] Note that in this and other examples, a dome can be defined as a PRVA that can be used to reconstruct one or more PRVAs identified by corresponding identifiers, in this example the PRVAs are identified by identifiers “a” and “b”.

[0238] Download and render dome screen

[0239] Figure 3 The diagram illustrates message exchange between streaming client 300 (labeled "CL") and streaming server 200 (labeled "SRV") for downloading and rendering video data representing a dome. This message exchange may involve the following: The dome may need to be downloaded before shared video data in the form of a dome can be used to render omnidirectional video. Figure 3An example of message exchange between streaming client 300 and streaming server 200 for downloading dome screens is shown. Specifically, streaming client 300 may first retrieve a list from streaming server 200, and then, during the running session, may determine its position within the scene by referring to the position of the currently selected viewpoint. If the position changes, for example, due to user selection or selection by streaming client 300 itself, streaming client 300 may determine, based on the list, which(s) of dome screens(s) are needed to render omnidirectional video at the new location, and subsequently retrieve that(s) of dome screen(s) from streaming server 200. After the dome screen(s)(s) have been retrieved from streaming server 200, in some embodiments, streaming client 300 may then render the dome screen(s)(s) as part of the omnidirectional video. For this purpose, although... Figure 3 Not shown, but streaming client 300 can additionally retrieve viewpoint-specific video content for the new location from streaming server 200. Viewpoint-specific video content and (multiple) domes can be combined by streaming client 300 in various ways, for example, by projecting the video content of the viewpoint-specific video content onto the top of the domes, such as by overlaying video content. To enable simple types of overlaying on the streaming client, for example, arrangements can be made at the streaming server or at the authoring system when creating the video data such that the coordinate system and projection type used for the viewpoint-specific video content and the domes are the same. In some examples, the viewpoint-specific video content can be formatted as a dome, for example, using an equal rectangular projection, but this video content may include portions that do not contain video data. These portions can be made transparent so that the corresponding video data on the domes remains visible even when the viewpoint-specific video content is overlaid on the domes.

[0240] Create a manifest file using dome information.

[0241] The following describes the creation of a manifest file containing dome information to enable streaming clients to correctly use the dome when rendering omnidirectional video based on the retrieved dome. For example, the following references the MPD as defined in common pending application EP 19 219597.2, but it will be understood that this dome information may also be provided as part of a different type of manifest, or generally as different types of metadata. In the following, it is assumed that the dome is formatted, for example, using an equal rectangular projection, into omnidirectional video that can be rendered in the same manner as PRVA. To enable streaming clients to render the dome, the following attributes may be included in the MPD, or generally in the manifest:

[0242] • ID: Name of the dome

[0243] • Offset: The dome will be rendered at this point in time relative to the timeline of the base media (also known as the base timeline).

[0244] • URI: Web location, which can be used to retrieve the dome screen.

[0245] The location of the dome can depend on which PRVAs the dome covers, or in other words, on which viewpoint-specific video data the dome's shared video data can be combined with. The dome can contain references to PRVA identifiers. This allows for the reconstruction of PRVAs from viewpoint-specific video content and from multiple domes (see also the "Overlapping Domes" section).

[0246] Here is an example of dome screen information:

[0247]

[0248] Tile streaming

[0249] A dome can be omnidirectional. In many examples, a streaming client may only be able to render a portion of the omnidirectional video at a time, such as the portion the user is currently viewing using a head-mounted display (HMD). This could mean that not all the omnidirectional video can be displayed at any given time. Therefore, it may not be necessary to stream all the omnidirectional video. To reduce the amount of data transmitted, tile streaming can be used. For this purpose, the dome and / or viewpoint-specific video content can each be divided across one or more tiles. The streaming client can then retrieve only those tiles that are currently needed, such as those used for rendering the omnidirectional video or for pre-caching purposes. Any retrieved tiles can be stitched together locally on the streaming client and rendered in their respective locations.

[0250] Overlapping Dome

[0251] On the one hand, a one-to-many relationship can exist between viewpoint and viewpoint-specific video data and shared video data on the other hand, because an omnidirectional video as a PRVA can be reconstructed using several instances of viewpoint-specific video data and shared video data (e.g., several domes). In fact, a PRVA can therefore be part of one or more domes, and a dome can contain one or more PRVAs. This creates the possibility of overlapping domes.

[0252] Figures 4A to 4C This demonstrates different types of overlapping dome screens for different types of scenes, namely... Figure 4A The corridor and Figure 4B and Figure 4C The open space within. Here, each viewpoint A through D can have its own viewpoint-specific content, and each dashed circle / ellipse can represent a dome screen. For example... Figure 4AAs can be seen in the image, in the corridor, the user can transition from viewpoint C to A to B to D. A dome can be defined as applicable to two viewpoints, with each viewpoint overlapping to create a series of domes. In an open space, all cameras are within each other's line of sight and collectively observe the scene. Figure 4B As shown, if applied with Figure 4A Using the same principles to create domes could result in generating a large number of domes for a relatively small number of viewpoints. Each dome could be a separately decodeable stream. This could create significant redundancy between the domes. Figure 4B The dome screens are combined to form a single overall dome screen, such as Figure 4C As shown. Whether this is beneficial can depend on the entropy of the encoded dome and the size difference between the overall dome and the individual domes. For example, for a clear sky, the size of the overall dome can be almost the same as the size of all the individual domes combined. However, if the sky contains many irregularities, such as hot air balloons and clouds, the sum of the individual domes can easily exceed the size of the overall dome. The overall dome can be created "offline," for example, during creation within the creation system. However, the overall dome can also be created in real time.

[0253] Figure 5 The diagram illustrates message exchange between streaming client 300 and streaming server 200 for switching between video data from different dome screens. In this example, streaming client 300 can first determine, for example, the resolution of the HMD display, then retrieve a list from streaming server 200, and then, during the running session, determine its position within the scene by referencing the position of the currently selected viewpoint. Streaming client 300 can then, for example, use a function called "calculate_distance_to_next_dome()" to determine the distance to the next dome screen. If this distance is less than a threshold, this prompts streaming client 300 to retrieve the next dome screen from streaming server 200. In this example, the user can... Figure 4A The scenario shown depicts a movement from viewpoint C to viewpoint A. First, a first dome covering viewpoints A and C can be downloaded and rendered. When the function result is less than 0.25, 3 / 4 of the distance to viewpoint A is covered. At this point, streaming client 300 can download and render a second dome covering viewpoints A and B. This way, at any given time, only one dome needs to be downloaded and decoded.

[0254] Shared video content

[0255] Different versions of the video content may exist. For example, if viewpoint A shows the sky and viewpoint B shows the same sky, a dome can be generated based on the sky at viewpoint A and / or the sky at viewpoint B. Conversely, panoramic videos of viewpoints A and B can be reconstructed based on viewpoint-specific video content at viewpoint A, viewpoint-specific video content at viewpoint B, and a dome representing shared video content, such as a version derived from panoramic video at viewpoints A or B, or in some examples, a version derived from both panoramic videos at viewpoints A or B. In the latter example, some portions of the shared video content can be derived from panoramic video at viewpoint A, while other portions can be derived from panoramic video at viewpoint B, for example, by selecting portions based on image quality standards or spatial resolution of objects in the video data (e.g., deriving image data of objects from panoramic video where the objects are closer to the camera).

[0256] In some examples, formatting the shared video content into separate files or video streams can be omitted. Instead, metadata can be generated that defines which portion of the panoramic video content represents the shared video content. In some examples, the panoramic video can be formatted to allow for independent retrieval and decoding of spatial segments, for example, in tile-encoded form. In such examples, the metadata can indicate the spatial segments or tiles representing the shared video content, allowing streaming clients to independently decode and, in some examples, retrieve the spatial segments or tiles of the shared video content.

[0257] In some examples, multiple versions of shared video content can be provided, and the streaming server and client can switch between streaming one version and streaming another. Various such examples exist, which may involve the use of one or two decoders. When the streaming client has multiple hardware decoders, using two decoders can be beneficial, such as an H.265 decoder and an H.264 decoder.

[0258] The first example could include:

[0259] 1. A streaming server can use signals to indicate which part of the video stream from viewpoint A is visible from viewpoint B. This part can represent the shared video content between viewpoints A and B, essentially a "dome".

[0260] 2. When a user moves from viewpoint A to viewpoint B within a scene, the streaming client can continue to receive stream A and additionally retrieve viewpoint-specific video data for viewpoint B.

[0261] 3. A streaming client can render view-specific video data for viewpoint B while still rendering shared video data from the stream from viewpoint A.

[0262] 4. When the streaming client reaches the streaming access point in the full video stream at viewpoint B, it can switch to that full video stream.

[0263] The second example can utilize tile streaming. This example assumes that the entire omnidirectional video includes viewpoint-specific video content below the horizon and a blue sky, which can be seen from several viewpoints (e.g., A, B, and C) above the horizon. The omnidirectional video of viewpoint B can then be reconstructed using a version of the sky from viewpoint A, and then the omnidirectional video of viewpoint C can be reconstructed using versions of the sky from viewpoint B, etc.

[0264] Figure 6A A second example is shown, in which omnidirectional videos of different viewpoints A, B, and C are reconstructed from viewpoint-specific video data and from shared video data. Specifically, it can be assumed that viewpoint A can be reconstructed at a given time from video data representing the sky at viewpoint A and from video data representing the scene below the horizon (“ground”) in viewpoint A. Therefore, both types of video content can be derived from video data acquired omnidirectionally by a camera at viewpoint A, such as… Figure 6A The images show A1 (sky) and A2 (ground) with the same shadow. When transitioning to viewpoint B... Figure 6A (Arrow 160 in the image) The ground B2 of viewpoint B can be streamed to the streaming client and used to reconstruct the omnidirectional video of viewpoint B, while the sky A1 of viewpoint A can continue to be streamed to the streaming client and used to reconstruct the omnidirectional video of viewpoint B. Then, the streaming client can switch to the sky B1 of viewpoint B at a certain point in time (…). Figure 6A (Not explicitly shown in the text), for example, by switching to the overall panoramic video of viewpoint B when a streaming access point is reached in the panoramic video of viewpoint B. When transitioning later to viewpoint C, the ground C2 of viewpoint C can be streamed to the streaming client and used to reconstruct the omnidirectional video of viewpoint C, while the sky B1 of viewpoint B can continue to be streamed and used to reconstruct the omnidirectional video of viewpoint C.

[0265] Figure 6A Therefore, it can be effectively shown that sky A1 can be used as a substitute for B1 to reconstruct the omnidirectional video of B. This may mean that sky B1 does not need to be immediately transmitted to the streaming client.

[0266] Figure 6B It demonstrates the switching of viewpoints A, B, and C over time. Specifically, Figure 6B Time 172 is shown along the horizontal axis and viewpoint 170 is shown along the vertical axis, which in this example are viewpoints A, B, and C, while also being compared with... Figure 6A A similar approach is used to show the omnidirectional video as portions above and below the horizon. The various portions derived from an initially acquired omnidirectional video are shown with the same shade. Specifically, Figure 6BThis could involve the following: Shared video content, such as different versions derived from different omnidirectional videos, may be similar to each other or significantly different. For example, the sky or other video content between viewpoint A and viewpoint B may appear very similar. This allows the streaming client to continue receiving the sky from the omnidirectional video of viewpoint A when switching from viewpoint A to viewpoint B. However, the sky in viewpoint C may be significantly different from the sky in viewpoint A. Therefore, in temporal instance 3, the streaming client can retrieve the version of the sky initially acquired at viewpoint B, thereby rendering the omnidirectional video of viewpoint B that was initially acquired. This allows the streaming client to temporarily continue receiving the sky from the omnidirectional video of viewpoint B before switching to a different version of the sky from the omnidirectional video of viewpoint C when switching to viewpoint C.

[0267] Dome projections, or shared video content in general, can be generated in various ways, such as automatically or semi-automatically by an authoring system, or manually by the user. Two examples of dome projection creation could be:

[0268] 1. Non-tile method

[0269] a. It is possible to compare the video content of two or more panoramic videos, for example, in the form of equal rectangular projections, to identify shared video content.

[0270] b. You can use dome elements that identify shared video content as spatial areas, such as creating a manifest in MPD format.

[0271] 2. Tile Streaming Method

[0272] a. Tiles from two or more panoramic videos can be compared to identify shared video content.

[0273] b. An inventory can be created, for example, in the form of an MPD, which can contain references to tiles representing panoramic video content of this kind.

[0274] c. For example, if the streaming server has limited storage, duplicate tiles representing different versions of the shared video content can be removed. Conversely, if sufficient storage is available, duplicate tiles can be retained, which allows switching between different versions of the shared video content, such as selecting the version that displays objects at the highest spatial resolution. Continuing with the generation of the reference manifest (which is the MPD in the example below), the MPD can be generated or modified to indicate which part of the panoramic video can represent the shared video content, for example, as follows:

[0275]

[0276] The URI for the dome here indicates that the dome uses the resources of panoramic video A. The parameters x, y, width, and height can define the portion of the area within the shared (equal-size rectangle) content between panoramic videos A and B, thus representing the shared video content or the dome.

[0277] For an example using tile streaming, an MPD can be generated to define a dome by referencing the appropriate tiles. In this case, tiles that do not need to be spatially connected can be referenced. Below is an example of how a dome can be defined in tile streaming:

[0278]

[0279] Each dome can include multiple tiles, and a single tile can be part of multiple domes. This many-to-many relationship can be created as follows:

[0280]

[0281] Dome tiles can reference multiple reference tiles, referring to the tiles initially acquired or generated from the panoramic video. Specifically, such multiple reference tiles can represent different versions of shared video content. This can be useful when sharing video content between panoramic videos, but objects are shown in one panoramic video at a higher spatial resolution due to the camera's spatial location within the scene. Streaming clients can decide to switch between reference tiles based on their location within the scene. In cases where two tiles contain the same content, the streaming server can decide to remove these duplicate tiles.

[0282] Transformation

[0283] The video content between two panoramic videos can be different or substantially different, but they will still look similar. A dome can be created using video content from one panoramic video. To reconstruct another panoramic video, a transformation function can be applied to the dome. An MPD example illustrating this is shown below.

[0284]

[0285]

[0286] For example, the transformation could be a warp transformation, a perspective transformation, or any other type of spatial transformation. Another specific example is an affine transformation. A manifest file allows the streaming client to determine which dome to apply the transformation to, what the transformation is, and any other parameters the transformation might require. For example, the transformation could be defined as a function that takes video content and parameters as input and provides the transformed video content as output. The manifest file could indicate that only a portion of the dome needs to be processed by the transformation, for example, not the entire "dome1" in the example above, but only a portion of "dome1". Alternatively, the transformation could also be provided to the streaming client via a URL "transform-url," where the streaming client can download an executable function based on, for example, JavaScript, WebGL, C, etc., to transform the video content of one dome to obtain another.

[0287] For each shared-dome-tile, a transformation can be applied in a similar manner to that described in the preceding paragraphs. This allows even tiles that don't perfectly match each other to be used. Furthermore, a single tile plus the transformation function may be sufficient to create most of the dome screen. Examples are given in the table below.

[0288]

[0289] Nested dome

[0290] Figure 7 The diagram illustrates objects OBJ with different angular sizes displayed on different domes DM1 and DM2, such as an object with angular size N3 in dome DM1, which is smaller than angular size N4 in dome DM2. When moving from viewpoint B to C, a switch from dome DM1 to dome DM2 can be made midway between the two domes. However, this may result in a visible jump in the perceived quality of the object OBJ. To avoid or reduce this visible jump, the video content of dome DM1 can be mixed with the video content of dome DM2, for example, using a mixing function. However, this may require both domes to be streamed and decoded simultaneously, at least temporarily.

[0291] Figure 8 It shows the relationship with Figure 7 A similar example is shown, but dome screen DM1 is shown for viewpoints A and B, and dome screen DM2 is shown for viewpoints C and D. Furthermore, Figure 8Two objects, OBJ1 and OBJ2, are shown. OBJ1 is closest to dome DM1, and OBJ2 is closest to dome DM2. When generating the dome, it can be decided to encode only the video data of OBJ1 as a part of dome DM1 and only the video data of OBJ2 as a part of dome DM2. This may result in dome nesting, such as an onion-shaped arrangement of domes, or... Figure 9 As shown. This likely means that in order to render omnidirectional video from viewpoint A (where the first object OBJ1 is shown nearby and the second object OBJ2 is shown in the distance), video content from both dome elements DM1 and DM2 can be retrieved. To indicate this to streaming clients, the "parent" property can be added to the dome element, which can contain multiple values ​​using the following content separator:

[0292]

[0293]

[0294] When retrieving dome screens recursively, this could cause the streaming client to retrieve all the dome screens in the scene. This could consume too much bandwidth. To reduce the necessary bandwidth, the streaming client's viewing direction can be used to retrieve only a subset of the dome screens. Here, the term "viewing direction" can be understood as referring to the client rendering only a subset of the omnidirectional video along a specific viewing direction in the scene. For this purpose, the function "determine_domes_to_get(pts,position)" mentioned above can be extended with a parameter named "viewing_direction". For example, this parameter could contain the viewing direction in degrees between 0° and 360°. This allows the streaming client to determine which subset of the dome screens to retrieve from the streaming server.

[0295] Dome Boundary

[0296] A dome can be defined as an application to a specific set of viewpoints; the term "application" refers to the fact that the video content of the dome can be used to reconstruct the corresponding viewpoints. If a viewpoint can be represented as coordinates in a coordinate system associated with the scene, then the dome can be defined as the boundary of that coordinate system, where the dome applies to all viewpoints within the boundary. Figure 10The boundaries of spherical screens DM1 and DM2 in a coordinate system associated with scene 100 are shown. It can be seen that spherical screen DM1 is defined to apply to viewpoints A, B, and C, while spherical screen DM2 is defined to apply to viewpoints D and E. It is understood that each spherical screen can also apply to intermediate viewpoints, for example, by generating intermediate viewpoints using viewpoint compositing techniques mentioned elsewhere in this specification. However, if viewpoint compositing is used to generate viewpoints outside the spherical screen boundaries, such as when transitioning between viewpoints C and D, it may be unclear which spherical screen to retrieve. As a possible solution, the streaming client can determine which spherical screen to retrieve based on the spherical screen boundaries, for example, determining which spherical screen is closest. This could mean that in Figure 10 In the example, the streaming client can switch from dome DM1 to dome DM2 at approximately the same location as object OBJ.

[0297] It is understandable that, in addition to viewpoint and dome, objects can also be defined based on their location within the scene's coordinate system. This allows streaming clients to adjust their streaming based on the relative location of the currently rendered or to-be-rendered viewpoint and the objects in the scene. For example, based on the object's location, a streaming client can retrieve different versions of shared video content depicting an object, for instance, by retrieving the version depicting the object at the highest spatial resolution. Accordingly, an MPD can be efficiently generated to provide a spatial map of the scene, defining the viewpoint, dome boundaries, and / or objects. This spatial map can be in the form of a grid. For example, objects can be defined in a manifest file as follows:

[0298]

[0299] Here, x and y can be decimals, where y = 1.0 can be the top-left corner of an object in the grid, and y = 0 can be the bottom-left corner. width, depth, and height can define the size of the object, for example, in meters, such as 200 × 200 × 200 meters, or in any other suitable manner. On the authoring end, the location of the object can be determined manually or automatically, for example, using a depth camera or image recognition technology for 3D depth reconstruction to reconstruct the size and / or location of objects in the scene from the acquired omnidirectional video content. By providing object coordinates, the streaming client can, for example, determine which spheres are best suited for acquiring video data for the sphere by comparing the object coordinates to the sphere boundaries. Thus, if such high-resolution rendering is desired, the streaming client can, for example, retrieve the sphere displaying the object at the highest resolution or retrieve a portion of the sphere displaying the object.

[0300] Figure 11A streaming server 200 is shown for streaming video data representing a panoramic video scene to a streaming client. The streaming server 200 may include a network interface 220 connected to a network. Data communication 222 can reach the streaming client via the network. The network interface 220 may be, for example, a wired communication interface, such as an Ethernet interface or a fiber-optic-based interface. The network may be, for example, the Internet or a mobile network, wherein the streaming server 200 is connected to a fixed portion of the mobile network. Alternatively, the network interface 220 may be, for example, designed for... Figure 12 The streaming client 300 describes the type of wireless communication interface.

[0301] The streaming server 200 may further include a processor subsystem 240, which may be configured, for example by hardware design or software, to perform the operations described in this specification, operations relating to the streaming server or generally relating to streaming video data of panoramic video of a scene to a streaming client. Typically, the processor subsystem 240 may be implemented by a single central processing unit (CPU), such as an x86 or ARM-based CPU, but may also be implemented by a combination or system of such CPUs and / or other types of processing units. In embodiments where the streaming server 200 is distributed across different entities (e.g., across different servers), the processor subsystem 240 may also be distributed across, for example, the CPUs of these different servers. Similarly, as... Figure 11 As shown, the streaming server 200 may include a data storage 260 that can be used to store data, such as a hard disk drive or hard disk drive array, a solid-state drive or solid-state drive array, etc.

[0302] Typically, the streaming server 200 can be a content delivery node, or multiple content delivery nodes can be implemented in a distributed manner. The streaming server 200 can also be implemented by another type of server or a system of such servers. For example, the streaming server 200 can be implemented by one or more cloud servers or by one or more edge nodes of a mobile network.

[0303] Figure 12A streaming client 300 is illustrated for receiving video data representing a panoramic video scene via streaming. The streaming client 300 may include a network interface 320 to a network to enable communication with a streaming server via data communication 322. The network interface 320 may be, for example, a wireless communication interface, which may also be referred to as a radio interface and can be configured to connect to mobile network infrastructure. In some examples, the network interface 320 may include a radio and an antenna, or include a radio and antenna connection. In specific examples, the network interface 320 may be a 4G or 5G radio interface for connecting to a 4G or 5G mobile network conforming to one or more 3GPP standards, or it may be a Wi-Fi communication interface for connecting to Wi-Fi network infrastructure, etc. In other examples, the network interface 320 may be, for example, for... Figure 11 The streaming server 200 describes a wired communication interface of the type described. Note that data communication between the streaming client 300 and the streaming server 200 may involve multiple networks. For example, the streaming client may connect to the infrastructure of a mobile network via a radio access network and to the Internet via the mobile network infrastructure, wherein the streaming server is a server that is also connected to the Internet.

[0304] The streaming client 300 may further include a processor subsystem 340, which may be configured, for example by hardware design or software, to perform the operations described in this specification, operations relating to the streaming client or generally relating to receiving panoramic video data of a scene via streaming. Typically, the processor subsystem 340 may be implemented by a single central processing unit (CPU), such as an x86 or ARM-based CPU, but may also be implemented by a combination or system of such a CPU and / or other types of processing units, such as a graphics processing unit (GPU). The streaming client 300 may further include a display output terminal 360 for outputting display data 362 to a display 380. The display 380 may be an external display or an internal display of the streaming client 300. Figure 12 An external display is shown, and it can typically be head-mounted or non-head-mounted. The streaming client 300 can display the received panoramic video using the display output 360. In some embodiments, this may involve the processor subsystem 340 rendering the panoramic video; the term "rendering" refers to one or more processing steps that transform the video data of the panoramic video into a displayable form. For example, rendering may involve mapping the video data of the panoramic video as a texture onto objects in a virtual environment, such as the interior of a sphere. The rendered video data can then be provided to the display output 360.

[0305] Typically, a streaming client 300 can be implemented by a (single) device or apparatus, such as a smartphone, personal computer, laptop computer, tablet device, game console, set-top box, television, monitor, projector, smartwatch, smart glasses, media player, media recorder, etc. In some examples, the streaming client 300 can be a so-called user equipment (UE) of a mobile telecommunications network such as 5G or next-generation mobile networks. In other examples, the streaming client can be an edge node of the network, such as the aforementioned mobile telecommunications edge node. In these examples, the streaming client may lack a display output, or may at least not use a display output to display the received video data. Instead, the streaming client can receive video data from a streaming server and reconstruct a panoramic video from it, which can then be used for streaming (e.g., via tile streaming) to further downstream streaming clients, such as end-user equipment.

[0306] Figure 13 An authoring system 400 is shown for creating one or more video streams representing a panoramic video scene. The authoring system 400 may include a data storage 460 capable of storing multiple panoramic videos from various viewpoints of the scene, such as a hard disk drive or hard disk drive array, a solid-state drive or solid-state drive array, etc. The authoring system 400 is further shown including a network interface 420 for data communication via a network 422. The network interface 420 may be for... Figure 11 Streaming server 200 or Figure 12 The streaming client 300 describes the type. In some examples, data storage 460 may be external storage that can be accessed over a network via network interface 420. The authoring system 400 may further include a processor subsystem 440, which may be configured, for example by hardware design or software, to perform the operations described herein, operations relating to the authoring system or generally relating to the creation of video streams, which may include the generation of shared video streams (“dome”), metadata, etc.

[0307] For example, processor subsystem 440 can be configured to identify shared video data representing shared video content of a scene, wherein the shared video content includes video content of the scene that is visible in a first panoramic video and also visible in a second panoramic video. As explained elsewhere, such shared video content can be identified, for example, by finding correspondences between video frames. Processor subsystem 440 can be further configured to identify first viewpoint-specific video data representing first viewpoint-specific video content of a scene, wherein the first viewpoint-specific video content includes video content of the scene that is part of the first panoramic video but not part of the shared video content of the scene. Processor subsystem 440 can be further configured to encode the first panoramic video by encoding the first viewpoint-specific video data into a first viewpoint-specific video stream and by encoding the shared video data into a shared video stream, or by encoding the first viewpoint-specific video data and the shared video data into a video stream, wherein the shared video data is included in the video stream as an independently decodable portion of the video stream.

[0308] Processor subsystem 440 can typically be targeted at Figure 11 Streaming server 200 or Figure 12 The streaming client 300 describes the type. Typically, the authoring system 400 can be implemented by a (single) device or apparatus, such as a personal computer, laptop computer, workstation, etc. In some examples, the authoring system 400 can be distributed across various entities, such as local servers or remote servers.

[0309] Typically, each entity described in this specification can be implemented as a device or apparatus, or implemented within a device or apparatus. The device or apparatus may include one or more (micro)processors executing appropriate software. The processor of the corresponding entity may be implemented by one or more of these (micro)processors. The software implementing the functionality of the corresponding entity may have been downloaded and / or stored in the corresponding one or more memories, for example, in volatile memory such as RAM or non-volatile memory such as flash memory. Alternatively, the processor(s) of the corresponding entity may be implemented in the device or apparatus as programmable logic, for example, as a field-programmable gate array (FPGA). Any input and / or output interfaces may be implemented by the corresponding interface of the device or apparatus. Typically, each functional unit of the corresponding entity may be implemented as a circuit or circuit system. The corresponding entities may also be implemented in a distributed manner, for example, involving different devices or apparatuses.

[0310] Note that any method described in this specification, such as any method described in any claim, can be implemented on a computer as a computer-implemented method, dedicated hardware, or a combination of both. Instructions for a computer (e.g., executable code) can be stored, for example... Figure 14 On the computer-readable medium 500 shown, executable code may be stored, for example, in the form of a series of machine-readable physical marks 510 and / or as a series of elements with different electrical (e.g., magnetic) or optical properties or values. Executable code may be stored in a transient or non-transient manner. Examples of computer-readable media include memory devices, optical storage devices, integrated circuits, servers, online software, etc. Figure 14 The optical storage device 500 is illustrated by way of example.

[0311] In alternative embodiments of the computer-readable medium 500, the computer-readable medium 500 may include transient or non-transient data 510 in the form of a data structure representing the metadata described in this specification.

[0312] Figure 15 This is a block diagram illustrating an exemplary data processing system 1000 that can be used in the embodiments described in this specification. Such a data processing system includes the data processing entities described in this specification, including but not limited to streaming servers, streaming clients, and authoring systems.

[0313] The data processing system 1000 may include at least one processor 1002 coupled to a memory element 1004 via a system bus 1006. Thus, the data processing system can store program code within the memory element 1004. Furthermore, the processor 1002 can execute program code accessed from the memory element 1004 via the system bus 1006. In one aspect, the data processing system may be implemented as a computer suitable for storing and / or executing program code. However, it is understood that the data processing system 1000 may be implemented in the form of any system including a processor and memory capable of performing the functions described herein.

[0314] Memory element 1004 may include one or more physical memory devices, such as local memory 1008 and one or more mass storage devices 1010. Local memory may refer to random access memory or other non-persistent memory(s) typically used during the actual execution of the program code. Mass storage devices may be implemented as hard disk drives, solid-state drives, or other persistent data storage devices. Data processing system 1000 may also include one or more cache memories (not shown) that provide temporary storage for at least some program code to reduce the number of times program code is otherwise retrieved from mass storage device 1010 during execution.

[0315] Input / output (I / O) devices, depicted as input device 1012 and output device 1014, may optionally be coupled to the data processing system. Examples of input devices include, but are not limited to, microphones, keyboards, pointing devices such as mice, game controllers, Bluetooth controllers, VR controllers, and gesture-based input devices. Examples of output devices include, but are not limited to, monitors or displays, speakers, etc. Input and / or output devices may be coupled to the data processing system directly or via an intermediate I / O controller. Network adapter 1016 may also be coupled to the data processing system to enable it to be coupled to other systems, computer systems, remote network devices, and / or remote storage devices via an intermediate private or public network. The network adapter may include a data receiver for receiving data transmitted to the data processing system by the system, device, and / or network, and a data transmitter for transmitting data to the system, device, and / or network. Modems, cable modems, and Ethernet cards are examples of different types of network adapters that can be used with the data processing system 1000.

[0316] like Figure 15 As shown, memory element 1004 can store application program 1018. It should be understood that data processing system 1000 can further execute an operating system (not shown) capable of facilitating the execution of the application program. The application program, implemented in the form of executable program code, can be executed by data processing system 1000 (e.g., by processor 1002). In response to executing the application program, the data processing system can be configured to perform one or more operations, which will be described in further detail herein.

[0317] For example, data processing system 1000 can be represented as shown in the reference. Figure 11 And the streaming server described elsewhere in this specification. In this case, application 1018 may represent an application that, when executed, configures data processing system 1000 to perform the functions described with reference to the entity. In another example, data processing system 1000 may represent as described with reference to Figure 12 And the streaming client described elsewhere in this specification. In this case, application 1018 may represent an application that, when executed, configures data processing system 1000 to perform the functions described with reference to the entity. In another example, data processing system 1000 may represent as described with reference to Figure 13 The authoring system described elsewhere in this specification. In this case, application 1018 may represent an application that, when executed, configures the data processing system 1000 to perform the functions described with reference to the entity.

[0318] It should be noted that the embodiments mentioned above are illustrative and not limiting of the invention, and those skilled in the art will be able to devise many alternative embodiments without departing from the scope of the appended claims.

[0319] In the claims, any reference numerals placed between parentheses should not be construed as limiting the claims. The use of the verb "comprising" and its variations does not exclude the presence of elements or levels other than those stated in the claims. The article "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. Expressions such as "at least one" preceding a list or group of elements indicate the selection of all elements or any subset of elements from the list or group. For example, the expression "at least one of A, B, and C" should be understood to include only A, only B, only C, both A and B, both A and C, both B and C, or all of A, B, and C. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In device claims enumerating several means, several of these means can be implemented by the same hardware. The fact that certain measures are stated in mutually different dependent claims does not imply that combinations of these measures cannot be advantageously used.

Claims

1. A method implemented by a computer for streaming video data to a streaming client, wherein, The video data represents a panoramic video of a scene, where the panoramic video shows the scene from one of multiple viewpoints within the scene. This method includes the following operations performed by the streaming server: - Stream the first video data, representing the first panoramic video from the first viewpoint within the scene, to the streaming client as the first video stream; - Provide metadata to the streaming client; - In response to the decision to stream second video data representing a second panoramic video from a second viewpoint within the scene, the second video data is streamed to the streaming client by performing at least temporarily and simultaneously the following operations in at least one operating mode of the streaming server: - Continue streaming at least a portion of the first video stream by continuing to stream shared video data representing shared video content of the scene, wherein the shared video content includes video content of the scene visible in the first panoramic video and visible in the second panoramic video, and wherein the metadata indicates shared video data representing a portion of the first video data and a portion of the second video data, and further identifies a second viewpoint-specific video stream containing second viewpoint-specific video data representing second viewpoint-specific video content of the scene; and - Stream the second viewpoint-specific video data as the second viewpoint-specific video stream, wherein the second viewpoint-specific video content includes video content of the scene that is part of the second panoramic video rather than part of the shared video content of the scene.

2. The method according to claim 1, wherein: - The streaming of the first video data includes: streaming first viewpoint-specific video data representing first viewpoint-specific video content of the scene, wherein the first viewpoint-specific video content includes video content of the scene that is part of the first panoramic video but not part of shared video content of the scene; and streaming a shared video stream including the shared video data; - The streaming of the second video data includes at least a temporary continuation of the streaming of the shared video stream.

3. The method according to claim 1, wherein, The streaming of the first video data includes streaming the shared video data as a video stream in which independently decodeable parts are included.

4. The method according to any one of claims 1 to 3, further comprising streaming the shared video data or any viewpoint-specific video data that is spatially segmented and encoded as at least a part of a panoramic video.

5. The method according to any one of claims 1 to 4, wherein, The streaming of the shared video data includes the periodic transmission of at least one of the following: -Images; and - Metadata, which defines the color used to fill the space area. The image or color represents at least a portion of the shared video content of at least a plurality of video frames.

6. A computer-implemented method for receiving video data, wherein, The video data represents a panoramic video of a scene, wherein the panoramic video shows the scene from a viewpoint within the scene, and the viewpoint is one of multiple viewpoints within the scene. The method includes the following operations performed by a streaming client: - Receive first video data representing a first panoramic video from a first viewpoint within the scene as a first video stream via streaming; - Receive metadata; -When switching to the second viewpoint Second video data is received via streaming from a second panoramic video from a second viewpoint within the scene. Receiving the second video data via streaming includes, at least temporarily and simultaneously, performing the following operations in at least one operating mode of the streaming client: - Continue receiving at least a portion of the first video stream via streaming by continuing to receive shared video data representing shared video content of the scene, wherein the shared video content includes video content of the scene that is visible in the first panoramic video and also visible in the second panoramic video. - Receive second viewpoint-specific video data representing second viewpoint-specific video content of the scene via streaming as a second viewpoint-specific video stream, wherein the second viewpoint-specific video content includes video content of the scene that is part of the second panoramic video rather than part of the shared video content of the scene; - Based on the metadata, the shared video data is used for rendering the second panoramic video, wherein the metadata indicates shared video data representing a portion of the first video data and a portion of the second video data and further identifies the second viewpoint-specific video stream containing second viewpoint-specific video data.

7. The method according to claim 6, wherein, The second viewpoint-specific video stream is accessible from the streaming server, and the method further includes requesting the second viewpoint-specific video stream from the streaming server based on the metadata and when switching to the second viewpoint.

8. The method according to any one of claims 6 or 7, wherein, This metadata is a list of the video data streams for this scene.

9. The method according to any one of claims 6 to 8, wherein, The metadata indicates a transformation to be applied to the shared video content for the reconstruction of the second panoramic video, and the method further includes applying the transformation to the shared video content before or as part of the rendering of the second panoramic video.

10. The method according to any one of claims 6 to 9, wherein, The location of the viewpoint can be represented as corresponding coordinates in a coordinate system associated with the scene, wherein the metadata defines the coordinate range of the shared video data, and wherein the method further includes using the shared video data to render a corresponding panoramic video of a viewpoint whose corresponding coordinates are located within the coordinate range defined by the metadata.

11. The method according to any one of claims 6 to 10, wherein, At least temporarily received shared video data is a first version of the shared video content derived from the first panoramic video, wherein the method further includes receiving second shared video data via streaming, the second shared video data representing a second version of the shared video content derived from the second panoramic video.

12. The method according to claim 11, wherein, The second version of the shared video content is received as a video stream, and the reception of the video stream begins from the stream access point in the video stream.

13. The method according to claim 12, wherein, The second version of the video stream containing the shared video content is a second shared video stream, wherein the second viewpoint-specific video data is received as the second viewpoint-specific video stream, and wherein the second shared video stream includes fewer streaming access points per time unit than the second viewpoint-specific video stream.

14. The method according to any one of claims 6 to 13, further comprising combining the shared video content with the second viewpoint-specific video content by at least one of the following operations: - Spatially adjoin the shared video content with the video content specific to the second viewpoint; -Merge the shared video content with the video content specific to the second viewpoint; and - Overlay the second viewpoint-specific video content onto the shared video content.

15. The method according to any one of claims 6 to 14, further comprising: - Render the second video data to obtain rendered video data for display by the streaming client or another entity, wherein the rendering includes combining the shared video data and the second viewpoint-specific video data.

16. A computer-implemented method for creating one or more video streams representing a panoramic video scene, the method comprising: - Access the first panoramic video, which shows the scene from a first viewpoint within the scene; - Access a second panoramic video, which shows the scene from a second viewpoint within the scene; -Identifiers represent shared video data of shared video content for the scene, wherein the shared video content includes video content of the scene that is visible in the first panoramic video and also visible in the second panoramic video; - Identify first viewpoint-specific video data representing first viewpoint-specific video content of the scene, wherein the first viewpoint-specific video content includes video content of the scene that is part of the first panoramic video rather than part of the shared video content of the scene; - Generate metadata, wherein the metadata indicates shared video data representing a portion of the first viewpoint-specific video data and a portion of the second viewpoint-specific video data representing the second viewpoint-specific video content of the scene, and further identifies the second viewpoint-specific video stream containing the second viewpoint-specific video data; and - Encode the first panoramic video in the following way: - Encode the first-viewpoint-specific video data into a first-viewpoint-specific video stream, and encode the shared video data into a shared video stream; or - The first viewpoint-specific video data and the shared video data are encoded into a video stream, wherein the shared video data is included in the video stream as an independently decodable portion of the video stream.

17. A computer-readable medium comprising transient or non-transient data representing a computer program, the computer program including instructions for causing a processor system to perform the method according to any one of claims 1 to 16.

18. The computer-readable medium of claim 17, wherein, The transient or non-transient data also defines a data structure that represents metadata identifying the video data of the video stream as shared video content representing a scene, wherein the metadata indicates that the shared video content is visible at least in the first and second panoramic videos of the scene.

19. A streaming server for streaming video data to a streaming client, wherein, The video data represents a panoramic video of a scene, where the panoramic video shows the scene from one of multiple viewpoints within the scene. The streaming server includes: - A network interface to the network, through which data communication can reach the streaming client; - A processor subsystem configured to perform the following operations via the network interface: The first video data, representing the first panoramic video from the first viewpoint within the scene, is streamed as the first video stream to the streaming client. Provide metadata to the streaming client; In response to the decision to stream second video data representing a second panoramic video from a second viewpoint within the scene, the second video data is streamed to the streaming client by at least temporarily and simultaneously performing the following operations: - Continue streaming at least a portion of the first video stream by continuing to stream shared video data representing shared video content of the scene, wherein the shared video content includes video content of the scene visible in the first panoramic video and visible in the second panoramic video, and wherein the metadata indicates shared video data representing a portion of the first video data and a portion of the second video data, and further identifies a second viewpoint-specific video stream containing second viewpoint-specific video data representing second viewpoint-specific video content of the scene; and - Streaming means that the second viewpoint-specific video data is a second viewpoint-specific video stream, wherein the second viewpoint-specific video content includes video content of the scene that is part of the second panoramic video rather than part of the shared video content of the scene.

20. A streaming client for receiving video data via streaming transmission, wherein, The video data represents a panoramic video of a scene, wherein the panoramic video shows the scene from a viewpoint within the scene, wherein the viewpoint is one of multiple viewpoints within the scene, and wherein the streaming client includes: -The network interface to the network; - A processor subsystem configured to perform the following operations via the network interface: First video data representing a first panoramic video from a first viewpoint within the scene is received via streaming as a first video stream. - Receive metadata; When switching to the second viewpoint, second video data from the second panoramic video within the scene is received via streaming, wherein receiving the second video data via streaming includes at least temporarily and simultaneously performing the following operations: - Continue receiving at least a portion of the first video stream via streaming by continuing to receive shared video data representing shared video content of the scene, wherein the shared video content includes video content of the scene that is visible in the first panoramic video and also visible in the second panoramic video. - Receive second viewpoint-specific video data representing second viewpoint-specific video content of the scene via streaming as a second viewpoint-specific video stream, wherein the second viewpoint-specific video content includes video content of the scene that is part of the second panoramic video rather than part of the shared video content of the scene; - Based on the metadata, the shared video data is used for rendering the second panoramic video, wherein the metadata indicates shared video data representing a portion of the first video data and a portion of the second video data and further identifies the second viewpoint-specific video stream containing second viewpoint-specific video data.

21. The streaming client according to claim 20, wherein, The streaming client is one of the following: - Network nodes, such as edge nodes; - End user equipment, which is configured to connect to a network.

22. A system for creating panoramic videos representing a scene, the system comprising: - Data storage interface, which is used for the following operations: Access the first panoramic video, which shows the scene from a first viewpoint within the scene; Access a second panoramic video, which shows the scene from a second viewpoint within the scene; - Processor subsystem, which is configured to: The identifier represents shared video data of shared video content for the scene, wherein the shared video content includes video content of the scene that is visible in the first panoramic video and also visible in the second panoramic video; The identifier represents first-viewpoint-specific video data of first-viewpoint-specific video content of the scene, wherein the first-viewpoint-specific video content includes video content of the scene that is part of the first panoramic video rather than part of the shared video content of the scene; - Generate metadata, wherein the metadata indicates shared video data representing a portion of the first viewpoint-specific video data and a portion of the second viewpoint-specific video data representing the second viewpoint-specific video content of the scene, and further identifies the second viewpoint-specific video stream containing the second viewpoint-specific video data, and The first panoramic video is encoded in the following manner: - Encode the first-viewpoint-specific video data into a first-viewpoint-specific video stream, and encode the shared video data into a shared video stream; or - The first viewpoint-specific video data and the shared video data are encoded into a video stream, wherein the shared video data is included in the video stream as an independently decodable portion of the video stream.

Citation Information

Patent Citations

  • An apparatus for transmitting a video, a method for transmitting a video, an apparatus for receiving a video, and a method for receiving a video

    WO2020036384A1