Peer-to-peer ultra-low latency streaming of real-time media

CN122556064APending Publication Date: 2026-08-11MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-11
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]交互式会议由于它们的低延迟而支持在会议参与者之间的交互,但是它们难以缩放至大量用户,并且甚至缩放至中等数量的用户也可能增加显著的复杂性(例如,就计算机和网络硬件、带宽等而言)

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122556064A_ABST
    Figure CN122556064A_ABST
Patent Text Reader

Abstract

This disclosure provides techniques and solutions for facilitating low-latency media streaming. Streaming techniques include sending or receiving blocks, wherein each block contains a sequence identifier for a media type and a single discrete sample of that specific media type. The block does not contain a sample of another media type. The sequence identifier can be used for purposes such as reducing the length of a block for a specific media type, reordering blocks, or copying duplicate blocks. The sequence identifier also facilitates peer-to-peer streaming techniques because it helps process blocks received by a streaming client from multiple peers.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Based on the need for interaction and participation, online meetings can be broadly categorized into interactive meetings (e.g., fast-lane meetings) and passive meetings (e.g., slow-lane meetings). Interactive meetings are typical online meetings where participants are free to contribute media to the session (e.g., chat, screen sharing, etc.). Passive meetings effectively deliver content to meeting participants without offering them the option to interact with or contribute media content to the online meeting. Passive meetings typically utilize streaming technologies and conventional Content Delivery Networks (CDNs) to deliver streaming media.

[0002] Passive conferencing can scale to handle media delivery on a global scale. However, passive conferencing experiences significant latency. For example, latency in passive conferencing (e.g., introduced by CDN and / or caching) can be approximately 30 seconds. The inherent latency associated with passive conferencing excludes any kind of meaningful interaction or participation with the presenter or other meeting participants.

[0003] Interactive meetings support interaction between participants due to their low latency; however, they are difficult to scale to large numbers of users, and even scaling to a moderate number of users can introduce significant complexity (e.g., in terms of computer and network hardware, bandwidth, etc.). For example, an interactive meeting can accommodate up to approximately 1,000 users. Therefore, there is room for improvement. Summary of the Invention

[0004] The choice of concepts to introduce a simplified form is provided in this summary of the invention, which will be further described in the following detailed implementation. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.

[0005] In one aspect, this disclosure provides a process at a streaming client to perform processing operations relative to a block comprising media samples to facilitate low-latency media streaming by reducing continuous execution of specific media sample types. A first data block is received. The first data block includes a first sequence identifier for a first media type and a first single discrete media sample of the first media type. The first data block does not include media samples of media types different from the first media type.

[0006] Receive multiple data blocks comprising corresponding single discrete media samples of the first media type. The single discrete media samples of the first media type are ordered in the stream. Receive a second data block. The second data block comprises a second sequence identifier for the second media type and a first single discrete sample of the second media type. The second data block does not include media samples of a media type different from the second media type. The second media type is a media type different from the first media type.

[0007] A threshold is determined to be met by identifying the number of consecutive discrete media samples of the first media type. Based on this determination, a first discrete sample of the second media type is inserted between consecutive media samples of the first media type in the stream. The first single discrete media sample of the first media type and the second discrete media sample of the second media type are provided to the media player for presentation at the streaming client. The streaming client receives data blocks for a stream comprising multiple data blocks from multiple sources.

[0008] On the other hand, this disclosure provides a process for performing streaming operations at a streaming client relative to a block including media samples to facilitate low-latency media streaming by reordering media samples of a certain type. A first data block is received at a first time. The first data block includes a first sequence identifier for a first media type and a first single discrete media sample of the first media type. The first data block does not include media samples of a media type different from the first media type.

[0009] A second data block is received at a second time. The second time is after the first time. The second data block includes a second sequence identifier for the first media type and a second single discrete sample of the first media type. The second data block does not include media samples of media types different from the first media type.

[0010] The second sequence identifier is determined to be lower than the first sequence identifier. The first single discrete media sample and the second single discrete media sample of the first media type are reordered such that the second single discrete media sample of the first media type is provided to the media player in the stream before the first single discrete media sample of the first media type. The first single discrete media sample and the second discrete media sample of the first media type are provided to the media player for presentation at the streaming client. The streaming client receives data blocks for a stream comprising multiple data blocks from multiple sources.

[0011] In another aspect, this disclosure provides a process at a streaming client to perform processing operations relative to a block including media samples to facilitate low-latency media streaming by discarding duplicate blocks. A first data block is received at a first time. The first data block includes a first sequence identifier for a first media type and a first single discrete media sample of the first media type. The first data block does not include media samples of media types different from the first media type.

[0012] A second data block is received at a second time. The second time is after the first time. The second data block includes a second sequence identifier for the first media type and a second single discrete media sample of the second media type. The second data block does not include media samples of a media type different from the first media type. It is determined that the second sequence identifier is the same as the first sequence identifier. The second data block is discarded, and the second single discrete media sample of the first media type is not provided for presentation at the client device. The first single discrete media sample of the first media type is provided to the media player for presentation at the streaming client. The streaming client receives data blocks for a stream comprising multiple data blocks from multiple sources.

[0013] On the other hand, this disclosure provides a process for performing operations as a streaming client in peer-to-peer delivery of streaming media content with low latency. A first data block is sent to a first receiving computing system. The first data block includes a first sequence identifier for a first media type and a first single discrete media sample of the first media type. The first data block does not include media samples of media types different from the first media type.

[0014] A second data block is sent to a second receiving computing system. The second data block includes a second sequence identifier for a second media type and a first single discrete sample of the second media type. The second data block does not include media samples of media types different from the second media type. The second media type is either the first media type or a media type different from the first media type. The second receiving computing system is either the first receiving computing system or a computing system different from the first receiving computing system.

[0015] A first single discrete media sample of the first media type and a second single discrete media sample of the second media type are presented. The first sequence identifier and the second sequence identifier are used to order the media samples for presentation at a first streaming client. The streaming client receives data blocks for a stream comprising multiple data blocks from multiple sources.

[0016] This disclosure also provides a procedure for operating in the low-latency streaming of media content. A first data block is received. The first data block includes a first sequence identifier for a first media type and a first single discrete media sample of the first media type. The first data block does not include media samples of media types different from the first media type.

[0017] A second data block is received. The second data block includes a second sequence identifier for a second media type and a first single discrete sample of the second media type. The second data block does not include media samples of a media type different from the second media type. The second media type is the first media type or a media type different from the first media type. The first single discrete media sample of the first media type and the second discrete media sample of the second media type are provided to the media player for presentation at the client device. The client device receives data blocks for a stream including the first data block and the second data block from multiple sources.

[0018] On the other hand, this disclosure provides operation in the delivery of media content via low-latency streaming. A first data block is sent to a first receiving computing system. The first data block includes a first sequence identifier for a first media type and a first single discrete media sample of the first media type. The first data block does not include media samples of media types different from the first media type.

[0019] A second data block is sent to a second receiving computing system. The second data block includes a second sequence identifier for a second media type and a first single discrete sample of the second media type. The second data block does not include media samples of a media type different from the second media type, and the second media type is either the first media type or a media type different from the first media type. The second receiving computing system is either the first receiving computing system or a computing system different from the first receiving computing system. The first sequence identifier and the second sequence identifier are used to order the media samples for presentation at a first streaming client. The first streaming client receives data blocks from a stream comprising the first data block and the second data block from multiple sources.

[0020] This disclosure also includes computing systems and tangible, non-transitory computer-readable storage media configured to perform or include instructions for performing the methods described above. As described herein, various other features and advantages can be incorporated into the art as needed. Attached Figure Description

[0021] Figure 1 It is a block diagram depicting an example environment for low-latency real-time streaming of media content.

[0022] Figure 2 This is a diagram depicting an example environment for low-latency real-time streaming of media content, including streaming of multiple predefined audio / video streams.

[0023] Figure 3 This is a flowchart of an example method for low-latency real-time streaming of media content.

[0024] Figure 4 This is a flowchart of an example method for low-latency real-time streaming of media content, including sending streaming metadata.

[0025] Figure 5 This is a flowchart of an example method for low-latency real-time streaming of media content, including using stream metadata to select audio / video streams.

[0026] Figure 6 It is a block diagram depicting an example environment for low-latency streaming of media content using lossless protocols.

[0027] Figure 7 This is a diagram depicting an example environment for selectively discarding media content when performing low-latency streaming using a lossless protocol.

[0028] Figure 8 This is a diagram depicting an example layer configuration for streaming video data that supports selectively discarding layers when performing low-latency streaming.

[0029] Figure 9 This is a flowchart of an example method for low-latency streaming of media content using lossless protocols.

[0030] Figure 10 This is a flowchart of an example method for low-latency streaming of media content using a semi-lossy protocol.

[0031] Figure 11 A diagram is provided illustrating example block formats used in the disclosed technology, as well as different block types that can be represented in the example block formats.

[0032] Figure 12 Provides the ability to provide for those with Figure 11 The attribute table provided by the media sample block in the block format.

[0033] Figure 13 Provides the ability to include targets with Figure 11 Another table showing the properties of media sample blocks in the example block format.

[0034] Figure 14 It is used for distributing with Figure 1Example block format of a peer-to-peer network graph.

[0035] Figure 15 It is a diagram illustrating the operations that can be performed by the interleaver component of a streaming client, including sample rearrangement, rejection of duplicate samples, and compensation for missing samples.

[0036] Figure 16 It is a diagram illustrating operations that can be performed by the interleaver component of a streaming client, including rearranging media samples to avoid long-term operation of media samples of the same media sample type.

[0037] Figure 17 This is a flowchart of the process at the streaming client to perform processing operations relative to blocks containing media samples in order to facilitate low-latency media streaming by reducing the continuous execution of specific media sample types.

[0038] Figure 18 This is a flowchart of the process at the streaming client to perform processing operations relative to a block containing media samples to facilitate low-latency media streaming by reordering a type of media samples.

[0039] Figure 19 This is a flowchart of the process at the streaming client to perform processing operations on blocks containing media samples to facilitate low-latency media streaming by discarding duplicate blocks.

[0040] Figure 20 It is a flowchart of the process by which a streaming client performs operations in the peer-to-peer delivery of streaming media content with low latency.

[0041] Figure 21 It is a flowchart of the process of performing processing operations on media content in low-latency streaming using block sequence identifiers.

[0042] Figure 22 It is a flowchart of the process of delivering streaming media content with low latency using block sequence identifiers.

[0043] Figure 23 This is a diagram of an example computing system in which some of the described embodiments can be implemented.

[0044] Figure 24 This is an example cloud-supported environment that can be used in conjunction with the technologies described in this article. Detailed Implementation

[0045] Example 1 - Overview The following describes a technique for low-latency real-time streaming of media content. For example, it is possible to receive streaming media content from a media source, wherein the streaming media content includes audio and / or video content. The audio / video stream can be streamed to one or more streaming clients. The audio / video stream is streamed as a sequence of encoded audio and / or video frames, which are independently encoded audio and / or video frames that are not grouped into chhunks for streaming. Furthermore, the sequence of encoded audio and / or video frames is streamed as a unidirectional stream to one or more streaming clients, and no requests for subsequent frames or chhunks are received from the one or more streaming clients.

[0046] The techniques described herein enable efficient (e.g., low-overhead) low-latency streaming of audio / video content to streaming clients. These techniques offer various advantages over existing streaming solutions (e.g., existing Content Delivery Network (CDN) solutions), including those that organize streaming content into chunks for streaming (also known as file-based streaming or segmented streaming). Using existing solutions that organize streaming content into chunks, the client periodically receives a list of requests for the next chunk of streaming content (e.g., the client periodically requests the next chunk, such as the next 2-second chunk of video and / or audio data). Furthermore, such existing solutions cache (e.g., cache) these chunks at various locations within the network (e.g., at various content delivery nodes). Therefore, such existing solutions suffer from high latency and high overhead.

[0047] The techniques described herein reduce the overhead of receiving video and / or audio samples because the streaming client does not need to request audio and / or video samples from the computing device (e.g., a server) that sends them. In other words, during streaming, the client does not send any requests to the server (i.e., the client only receives streaming audio and / or video frames without any requests or polling). By not sending any requests or performing any polling operations during streaming, the techniques described herein provide reduced overhead (e.g., reduced utilization of computing resources such as processors, network bandwidth, and memory) and reduced latency (e.g., the client does not spend time sending requests for the next set of chunks and waiting for a response). This contrasts with existing solutions, such as existing content delivery network solutions, where the client requests (e.g., as a polling operation) each next set of chunks of streaming media content.

[0048] The techniques described in this paper enable efficient forking of audio / video streams. Forking is performed using techniques that result in lower overhead, lower latency, and reduced utilization of computational resources. For example, forking can be achieved by sending audio and / or video frames to one or more additional streaming clients without modifying the header information of each frame (e.g., frames can be sent to additional streaming clients while being received without any modification or other processing). Forking can be performed with minimal setup (e.g., sending a list from which clients or delivery nodes select one or more streams to receive). Once minimal setup is performed, streaming can begin and continue without any additional requests from clients or delivery nodes.

[0049] This disclosure also describes the ability to monitor multiple streaming clients to determine if any of them has fallen behind in the streaming of the media stream. When a streaming client falls behind, a portion of the video data to be streamed to the streaming client can be selectively discarded based on scalability information and / or Long-Term Reference (LTR) frame information. Low-latency streaming can be performed without using per-client quality feedback from multiple streaming clients. When streaming using a semi-lossy protocol, multiple delivery modes can be used, each for different types of encoded video data and providing different levels of reliability.

[0050] This disclosure also provides specific block formats that facilitate low-latency streaming. In one aspect, the block format includes a sequence number that identifies the position of a media sample within a block within a stream for a specific media type. The sequence number can be used to order blocks within a stream of that media type, or to determine whether a duplicate block has been received or whether a block has been lost. Compared to, for example, using timestamps associated with media samples included in a block, the sequence number provides a faster way to manage packets. The sequence number can be useful when varying network latency may be encountered, allowing blocks to be received out of order.

[0051] The sequence number also facilitates a distribution topology where streaming clients can receive blocks from multiple sources, including blocks with the same media type. As a specific example, the disclosed technique facilitates the use of peer-to-peer networks for block delivery. If the same block is received from multiple clients, this can be quickly determined using the sequence number. Similarly, if blocks are received out of order, the sequence number can be used to place the blocks in the correct sequence.

[0052] Chunks can be of different types, including chunks of a specific type for a media sample or chunks containing content different from that of a media sample. A given chunk can have an identifier for its chunk type, which therefore allows media samples to be placed in the correct media type stream, and where the sequence number allows for the ordering of samples within a given media stream.

[0053] Block types that do not include media samples can include block types that initialize "tracks" for a specific media stream, such as information about codecs that can be used by the media player, wherein, as mentioned above, after the track is initialized at the streaming client, subsequent media blocks do not need to include such information.

[0054] Another block type can be used for key management, including key rotation. For example, a block type can include a mapping from a GUID for an encryption key to a local key ID. Another block type can be used for arbitrary system messages, such as indicating when a stream can be restarted, for example due to a resolution change, or indicating when a stream schedule ends.

[0055] The disclosed block format can be used to send data for use in other streaming protocols such as HLS or DASH. Therefore, block types can include: block types that include file data; and block types that can be used to delete files, such as those used for cache management.

[0056] As discussed, various aspects of the block format can be used to implement peer-to-peer block distribution. Sequence numbers allow for rapid analysis of samples received from multiple sources, ensuring samples are placed in the correct order within the correct media sample stream and rejecting duplicate samples. At the streaming client, these and other operations can be performed by the interleaver component. Other actions that can be performed by the interleaver component include: rearranging sample types in the stream to avoid large numbers of consecutive samples of the same type, and taking action to identify and potentially repair gaps in samples of a specific media type stream.

[0057] The disclosed streaming technology and block format can have additional features that enable improved streaming. For example, blocks can "pack" media samples within a block, such as helping to provide blocks with media samples of a fixed size or a range of desired sizes. Assuming that the overhead associated with sending blocks is relatively unaffected by the size of the media samples, it is possible to use media of a fixed duration (such as 100 milliseconds) or a fixed buffer size to help "amortize" the cost of the network "send" operation.

[0058] Example 2 - Terminology The term "media source" refers to the source of streaming audio and / or video content. In some implementations, the streaming audio and / or video content is real-time streaming audio and / or video content. For example, real-time streaming audio and / or video content can originate from a real-time video conference or meeting (e.g., generated by combining audio and / or video content from multiple participants into a composite audio and / or video stream). The media source can provide streaming audio and / or video content in various formats. For example, streaming audio and / or video content can be provided as unencoded (e.g., raw) audio and / or video samples (e.g., generated by locally connected or remote audio and / or video capture devices). Streaming audio and / or video content can also be provided as encoded audio and / or video data (e.g., encoded according to a corresponding audio and / or video coding standard).

[0059] The term "audio / video stream" refers to a stream containing a sequence of encoded audio and / or video frames (e.g., including corresponding audio and / or video samples). The disclosed techniques enable media block types that include single media samples of a single media type. These media samples of the disclosed block format can be referred to as "discrete media samples." The encoded audio and / or video frames are encoded via corresponding audio and / or video codecs. A given audio / video stream is encoded at a specific predefined quality (e.g., a specific resolution, bitrate, etc.). The term "predefined quality" indicates that the quality is client-independent and not specific to any given streaming client. In other words, the streaming techniques described herein enable streaming clients to select from multiple audio / video streams of predefined quality.

[0060] The term "stream metadata" refers to information describing audio / video streams available from a given delivery node for a specific streaming media content. This information includes indications of predefined quality for each available audio / video stream (e.g., resolution, bitrate, etc.). For example, there may be three available audio / video streams with different predefined qualities for the streaming media content identified for streaming. The stream metadata can identify the three available audio / video streams with labels such as "high," "medium," and "low" quality. The stream metadata can also provide more specific information describing the three available audio / video streams (e.g., indicating that the first audio / video stream has 720p video quality, the second audio / video stream has 1080p video quality, etc.). The stream metadata can also identify the specific audio and / or video codecs used for a given audio / video stream (e.g., indicating that the first audio / video stream contains AAC-encoded audio data and H.264-encoded video data).

[0061] The term "delivery node" refers to software and / or hardware configured (e.g., via software instructions) to perform low-latency, real-time streaming of audio / video streams. A delivery node can be a primary delivery node or a client delivery node. A primary delivery node typically operates in a cloud environment (e.g., via cloud computing services) and distributes the audio / video stream to end-user clients and client delivery nodes. Client delivery nodes typically operate within a network (such as an organization's network) and distribute the audio / video stream to end-user clients and other client delivery nodes (e.g., within the organization).

[0062] Delivery nodes acting as clients can be used in peer-to-peer block distribution. Peer-to-peer block distribution can include enabling a given client to send and receive blocks. Furthermore, a client can receive blocks from or supply blocks to multiple other peer clients. Peer-to-peer implementations can include other features of peer-to-peer networks, including at least some aspects discovered by the client's peers, the distribution of distributed hash tables, and the maintenance of routing tables. Peer-to-peer network implementations can include at least certain functionalities performed by centralized components or assisted by centralized components.

[0063] The term "streaming client" refers to a client that receives audio / video streams. A streaming client can be an end-user client that is the destination of the audio / video stream. For example, an end-user client can be a software application running on a computing device (e.g., a laptop or desktop computer, tablet, smartphone, or other type of computing device) that decodes and renders (e.g., via audio and / or video playback) the received audio / video stream. A streaming client can also be a client delivery node that further distributes the audio / video stream (e.g., distributes it to other end-user clients and / or other client delivery nodes).

[0064] The term "media content" refers to encoded audio and / or video content. Encoded audio and / or video content comprises a sequence of encoded audio and / or video frames (e.g., including corresponding audio and / or video samples). The encoded audio and / or video frames are encoded via corresponding audio and / or video codecs. By sending individual media samples, such as video frames, additional encoding can be avoided, at least in some implementations, and therefore, the samples can be directly provided to the media player and rendered using an appropriate codec.

[0065] The term "streaming" refers to the transmission of media content as a media stream from a first computing device to a second computing device via a computer network.

[0066] The term "lossless protocol" refers to one or more network protocols that provide reliable transmission of data over a computer network from a first computing device (e.g., a server) to a second computing device (e.g., a client). When data fails to reach its destination computing device, a lossless protocol retransmits the data until it is successfully received (thus providing reliable transmission). Examples of lossless protocols include Transmission Control Protocol (TCP), Hypertext Transfer Protocol (HTTP), HTTP Live Streaming (HLS), and Dynamic Adaptive Streaming over HTTP (DASH). In some implementations, lossless transmission protocols such as TCP are used to ensure reliable transmission of data to the client.

[0067] The term "semi-lossless protocol" refers to one or more network protocols that provide semi-reliable transmission of data over a computer network from a first computing device (e.g., a server) to a second computing device (e.g., a client). Semi-lossless protocols can provide different levels of reliability (e.g., different delivery modes).

[0068] Example 3 - Example Environment for Low-Latency Real-Time Streaming of Media Content Figure 1 This is a diagram depicting an example environment 100 for low-latency real-time streaming of media content. Example environment 100 includes a media source 110. Media source 110 is a source for streaming audio and / or video content (e.g., unencoded or encoded audio and / or video content).

[0069] Example environment 100 includes a primary delivery node (PDN) 120. The primary delivery node 120 receives streaming audio and / or video content from a media source 110. For example, the primary delivery node 120 can receive streaming audio and / or video content as unencoded audio and / or video samples or as encoded audio and / or video frames (e.g., at one or more quality levels). In some implementations, the primary delivery node 120 receives unencoded audio and / or video content from the media source 110 and performs audio and / or video encoding operations (e.g., using audio and / or video codecs) to generate encoded audio and / or video streams at one or more quality levels. In some implementations, the primary delivery node 120 receives streaming audio and / or video content in an already encoded format (e.g., encoded at one or more quality levels) and is capable of performing relay and / or transcoding operations. The primary delivery node 120 provides audio / video streams to the client delivery nodes (including client delivery node 140) and the end-user clients (including end-user clients 130).

[0070] Example environment 100 includes client delivery nodes 140, 142, and 144. For example, client delivery nodes 140, 142, and 144 may be located within an organization's network to serve clients locally within the organization. Client delivery nodes 140, 142, and 144 deliver audio / video streams to other client delivery nodes and / or end-user clients. As depicted, client delivery node 140 delivers audio / video streams to client delivery nodes 142 and 144. Client delivery node 142 delivers audio / video streams to end-user client 150. Client delivery node 144 delivers audio-video streams to end-user client 152.

[0071] The audio / video streaming within example environment 100 (by master delivery node 120 and client delivery nodes 140, 142, and 144) is performed by streaming the sequence of encoded audio and / or video frames as independently encoded audio and / or video frames without grouping the frames into chunks for delivery, as depicted at 160. In other words, end-user clients (including end-user clients 130, 150, and 152) receive the audio / video stream without sending requests for the next chunk of the streaming content (as is typically done using a content delivery network). As a result, audio / video streaming is performed within example environment 100 with low overhead and very low latency (e.g., no requests for subsequent frames from end-user clients).

[0072] Additionally, audio / video streaming within the streaming architecture depicted in Example Environment 100 is performed by directly streaming the encoded audio and / or video frames (e.g., directly from the primary delivery node 120 to the client delivery node 140, to the client delivery node 144, and ultimately to the end-user client 152). In at least some implementations, direct streaming is performed without any caching or buffering at the delivery nodes, as depicted at 160. For example, when client delivery node 140 receives encoded frames of audio and / or video content from primary delivery node 140, client delivery node 140 sends the encoded frames to client delivery nodes 142 and 144 (e.g., by making copies of the received encoded frames as needed) without any caching or buffering and without waiting for requests for additional frames from client delivery nodes 142 and 144. In this way, once audio / video streaming has begun, it is a unidirectional stream. As a result, audio-video streaming was performed in the example environment with low overhead and very low latency. However, in other implementations, the delivery node is able to buffer samples (e.g., in packets containing media samples such as frames). For example, if the client delivery node 140 sends samples to multiple receivers, it is possible to maintain the samples in a single buffer or multiple buffers until the media samples have been sent to the relevant receivers.

[0073] Additionally, example environment 100 supports efficient forking of audio / video streams (e.g., with low overhead). For example, audio / video streams can be forked so that they can be delivered to additional end-user clients and / or additional delivery nodes. Forking is performed by sending stream metadata to the new end-user client and / or delivery node. Once the end-user client and / or delivery node has selected the stream (or multiple streams) they want, streaming begins and continues without any additional requests from the new end-user client and / or delivery node. In this way, forking can be performed efficiently to deliver audio / video streams to many additional streaming clients. For example, if thousands of new streaming clients join, they can be added efficiently by sending stream metadata, receiving requests for the selected audio / video streams, and forking the audio / video streams (which have already been received by the delivery node) to send them to the new streaming clients.

[0074] In a typical streaming scenario, the primary delivery node 120 provides various audio / video streams of different qualities, all of which encode the same audio and / or video content received from the media source 110. For example, the primary delivery node 120 may provide high-quality, medium-quality, and low-quality audio / video streams, where the low-quality stream contains the same audio and / or video content received from the media source 110, encoded only at different qualities. Each end-user client and delivery node can select which audio / video streams to receive. For example, a given end-user client can receive stream metadata describing the available audio / video streams and select one to receive. A given client delivery node can receive all available audio / video streams (e.g., enabling it to offer all options to downstream client delivery nodes and / or end-user clients) or receive only a subset of the available audio / video streams (e.g., if the client delivery node is only currently serving end-user clients who have already selected those quality streams, then the client delivery node may only receive low-quality and medium-quality streams).

[0075] The different quality audio / video streams provided within example environment 100 are client-independent audio / video streams not specific to any given streaming client. For example, if a new end-user client wants to begin receiving streams, it first receives stream metadata describing the available audio / video streams, each with its associated predefined quality. The end-user client selects the audio / video stream of the desired quality to receive and begins receiving the selected audio / video stream. The selected audio / video stream may not be suitable for a particular streaming client, but the streaming client is able to select the audio / video stream of the desired quality (from those available) based on various criteria such as available computing resources at the streaming client, current network bandwidth, and conditions.

[0076] Example environment 100 is capable of delivering audio / video streams using various transport layer network protocols. For example, it is possible to perform the delivery of audio / video streams using lossless transport protocols (e.g., Transmission Control Protocol (TCP)) or lossy transport protocols (e.g., User Datagram Protocol (UDP)).

[0077] Using the techniques described above with respect to example environment 100, streaming of audio and / or video content to numerous end-user clients (e.g., 100,000 or more end-user clients) can be performed efficiently and with low latency (e.g., less than one second). For example, real-time media sessions (e.g., real-time audio and / or video conferencing) can be provided by media source 110, traveling through various master delivery nodes (e.g., master delivery node 120) and / or client delivery nodes (e.g., client delivery nodes 140, 142, and 144), and arriving at end-user clients (e.g., end-user clients 130, 150, and 152) with low latency and low overhead. Therefore, using the techniques described herein, the streaming architecture can practically support any number of end-user clients while still allowing end-user clients to have meaningful interaction or participation with the streaming media content when needed (e.g., if the streaming media content is part of a real-time conference, the end-user client can interact with conference participants in real time).

[0078] The streaming client can switch between audio / video streams of various qualities as needed. For example, a given end-user client can send a message to a delivery node to switch to an audio / video stream of a different quality. In response, the delivery node stops sending the old quality audio / video stream and begins sending the new quality audio / video stream. The streaming client can perform the switching based on previously received stream metadata, or request (or receive) updated stream metadata (e.g., the available predefined quality audio-video streams may have changed). New track setup information can be sent to the streaming client before sending packets for the newly selected stream.

[0079] Example environment 100 depicts a single master delivery node 120. However, other implementations may include any number of master delivery nodes, each receiving streaming audio and / or video content from media source 110 and / or from another master delivery node. Furthermore, each master delivery node is capable of streaming to any number of client and / or delivery nodes (e.g., master delivery nodes or client delivery nodes).

[0080] Figure 2 This is a diagram depicting an example environment 200 for low-latency real-time streaming of media content, including streaming of multiple predefined audio / video streams. Example environment 200 is similar to... Figure 1 The example environment 100 described in the document further illustrates how several predefined audio / video streams operate.

[0081] As depicted in example environment 200, the primary delivery node 120 has a set of predefined available audio / video streams. In this example, the set of predefined available audio / video streams includes high-quality (H) audio / video streams, medium-quality (M) audio / video streams, and low-quality (L) audio / video streams. This set of predefined available audio / video streams can be generated by the primary delivery node 120. For example, the primary delivery node 120 can receive streaming audio and / or video content from media source 110 and generate high, medium, and low-quality representations of the streaming audio and / or video content by encoding it using one or more audio and / or video codecs. The primary delivery node 120 can also receive high, medium, and low-quality streams as already encoded streams representing the streaming audio and / or video content (e.g., from media source 110 or from another source, such as an intermediate media encoding and / or compositing service).

[0082] The primary delivery node 120 streams one or more of a predefined set of available audio / video streams to the end-user client and / or the client delivery node. In some implementations, the primary delivery node 120 provides stream metadata to the requesting end-user client and the client delivery node, which then request to receive one or more of the predefined set of available audio / video streams. For example, end-user client 210 has selected a high-quality audio / video stream to receive, while end-user client 212 has selected a medium-quality audio / video stream to receive. The client delivery node 140 has selected the entire set of predefined available audio / video streams (high, medium, and low quality) to receive and will therefore be able to stream any audio / video stream from that set.

[0083] In this example, client delivery node 142 has requested both high-quality and medium-quality predefined audio / video streams and streams the high-quality audio / video stream to end-user client 214, and streams the medium-quality audio / video stream to end-user client 216. Client delivery node 144 has requested both medium-quality and low-quality predefined audio / video streams and streams the medium-quality audio / video stream to end-user client 218, and streams the low-quality audio / video stream to end-user client 220.

[0084] In some implementations, an end-user client can transform into a client delivery node and stream audio / video streams to other end-user clients and / or client delivery nodes. In this example, end-user client 210 has been transformed into a client delivery node, as depicted at 230. After becoming a client delivery node, end-user client 210 is able to provide streaming metadata to other end-user clients and / or client delivery nodes, receive requests for available predefined audio / video streams, and stream selected audio / video streams. In this example, end-user client 210 (which now also operates as a client delivery node) streams a high-quality audio / video stream to end-user client 235.

[0085] In some implementations, media source 110 represents multiple components (e.g., multiple cloud services, which may run in a localized or distributed manner). For example, media source 110 may include: a media processor component that receives real-time audio and / or video content (e.g., from a live conference); a media compositing runtime component that receives real-time audio and / or video content from the media processor and composites the content into one or more audio / video streams (e.g., including multiple streams of predefined quality) that are provided to one or more master delivery nodes; and a lookup service that facilitates discovery and communication between the master delivery nodes and other components.

[0086] Example 4 - Example Streaming Protocol The techniques described herein provide a novel streaming protocol for low-latency, real-time streaming of media content. In some implementations, this novel streaming protocol is a unidirectional streaming protocol that does not allow requests to be received.

[0087] The new streaming protocol includes a setup process. During this setup process, the streaming client sends a request for available audio / video streams (e.g., to a delivery node). In some implementations, this request is sent as a Hypertext Transfer Protocol (HTTP) request. In response, the streaming client receives stream metadata (also known as a track list) describing the available predefined audio / video streams. For example, multiple predefined quality audio / video streams may be available, such as high-quality streams, medium-quality streams, and low-quality streams.

[0088] The streaming client then selects one of the predefined audio / video streams to receive. In some implementations, a request for one of the predefined audio / video streams is sent as an HTTP request (e.g., to a delivery node). In response, the streaming client receives one or more track setup objects. Each track setup object describes the attributes of the corresponding track for the selected audio / video stream. An example track setup object describing an H.264 video track may include information such as frame rate, sequence parameter set (SPS) and picture parameter set (PPS) information, profile information, and / or other information that allows the streaming client to configure its decoder for receiving and decoding H.264 encoded video frames. An exemplary track setup object describing an Opus audio track may include information such as sample rate, sample duration, number of channels, and / or other information that allows the streaming client to configure its decoder for receiving and decoding Opus encoded audio data.

[0089] The encoded audio and / or video frames are then streamed to the streaming client using the new streaming protocol. The new streaming protocol streams the encoded audio and / or video frames as independently encoded audio and / or video frames, without grouping them into chunks for streaming. Furthermore, the encoded audio and / or video frames streamed by the new streaming protocol are unsearchable (i.e., the client cannot send requests for specific frames or chunks, such as using timestamps). In other words, the new protocol is a streaming-based protocol that streams encoded audio and / or video frames in real time without buffering or cached frames. The client receives the audio / video stream and begins real-time decoding (e.g., starting from the next keyframe).

[0090] In some implementations, the new streaming protocol is a one-way streaming protocol that does not allow receiving requests. In these implementations, different network protocols, such as HTTP, are used to receive requests from downstream streaming clients. For example, the primary delivery node is able to receive HTTP requests from streaming clients for available audio / video streams and HTTP requests from streaming clients for selected audio / video streams to initiate streaming. The primary delivery node can use the new streaming protocol when sending data to streaming clients (e.g., when sending streaming metadata, when sending track setting objects, and when streaming encoded audio and / or video frames).

[0091] The new streaming protocol is defined at a layer higher than the transport layer. Therefore, the new streaming protocol can use various transport layer network protocols to deliver encoded audio and / or video frames. For example, the new streaming protocol can use lossless transport protocols (e.g., TCP) or lossy transport protocols (e.g., UDP).

[0092] Example 5 - Example methods for low-latency real-time streaming of media content In any of the examples herein, a method for low-latency real-time streaming of media content can be provided. This method can be performed by a delivery node (e.g., a primary delivery node and / or a client delivery node) or by a streaming client.

[0093] Figure 3 This is a flowchart of an example method 300 for low-latency real-time streaming of media content. For example, example method 300 can be executed by a delivery node, such as a primary delivery node 120 or a client delivery node 140, 142 or 144.

[0094] At point 310, streaming media content is received from a media source. The streaming media content includes audio and / or video content.

[0095] At position 320, the audio / video stream is streamed to one or more streaming clients. The audio / video stream is streamed as a sequence of encoded audio and / or video frames generated from media content. The sequence of encoded audio and / or video frames is streamed as independently encoded audio and / or video frames without grouping the frames into chunks. In some implementations, the audio / video stream is streamed to one or more streaming clients without caching or buffering the encoded audio and / or video frames.

[0096] At point 330, the sequence of encoded audio and / or video frames is streamed as a unidirectional stream to one or more streaming clients, and no requests for subsequent frames or chunks are received from the one or more streaming clients. In some implementations, the sequence of encoded audio and / or video frames is streamed using a streaming protocol that does not support sending requests for subsequent frames or chunks to the delivery node.

[0097] Figure 4 This is a flowchart of an example method 400 for low-latency real-time streaming of media content, including sending streaming metadata. For example, example method 400 can be executed by a delivery node, such as a primary delivery node 120 or a client delivery node 140, 142 or 144.

[0098] At point 410, streaming media content is received from a media source. The streaming media content includes audio and / or video content.

[0099] At 420, a request for an available audio / video stream is received from the streaming client. For example, the request may include an indication (e.g., a unique identifier) ​​of the streaming media content the streaming client wants to receive. In some implementations, the request is received from the streaming client via an HTTP request.

[0100] At 430, in response to the request at 420, streaming metadata is sent to the streaming client. This streaming metadata describes a set of predefined available audio / video streams (with corresponding predefined qualities) for streaming the streaming media content. This set of predefined available audio / video streams are client-independent audio / video streams that are not specific to any given streaming client.

[0101] At 440, a selection of audio / video streams (from the set of predefined available audio / video streams) is received from the streaming client. In some implementations, the request is received from the streaming client via an HTTP request.

[0102] At 450, in response to the selection at 440, the selected audio / video stream is streamed to the streaming client. The selected audio / video stream is streamed as a sequence of encoded audio and / or video frames, which are independently encoded audio and / or video frames without grouping frames into chunks. In some implementations, the audio / video stream is streamed to the streaming client without buffering or caching the encoded audio and / or video frames.

[0103] At 460, the sequence of encoded audio and / or video frames is streamed as a unidirectional stream to the streaming client, and no requests for subsequent frames or chunks are received from the streaming client. In some implementations, the sequence of encoded audio and / or video frames is streamed using a streaming protocol that does not support sending requests for subsequent frames or chunks to the delivery node.

[0104] Figure 5 This is a flowchart of an example method 500 for low-latency real-time streaming of media content, including using stream metadata to select audio / video streams. For example, example method 500 can be executed by a streaming client, such as an end-user client or a client delivery node.

[0105] At point 510, a request for an available audio / video stream is sent (e.g., to the delivery node) for streaming media content identified by the streaming service. For example, the request may identify (e.g., using a unique identifier) ​​the streaming media content.

[0106] At 520, in response to the request at 510, streaming metadata is received. This streaming metadata describes a set of predefined, available audio / video streams for streaming the streaming media content. This set of predefined audio / video streams are client-independent audio / video streams that are not specific to any given streaming client.

[0107] At point 530, (e.g., to the delivery node) a selection of an audio / video stream from the set of predefined available audio / video streams is sent. For example, the stream metadata can include indications of predefined available audio / video streams (e.g., unique identifiers) that can be used when selecting an audio / video stream.

[0108] At 540, the audio / video stream selected at 530 is received. The audio / video stream is received as a sequence of encoded audio and / or video frames, which are independently encoded audio and / or video frames without grouping the frames into chunks.

[0109] At 550, the sequence of encoded audio and / or video frames is received as a unidirectional stream, and no requests for subsequent frames or chunks are sent (e.g., to the delivery node). In some implementations, the sequence of encoded audio and / or video frames is received using a streaming protocol that does not support the delivery node sending requests for subsequent frames or chunks.

[0110] Example 6 - Example Media Stream and Logical Buffer As already described, in the disclosed technology, media content is delivered to the client as a media stream (e.g., as a live media stream). For example, the media content of a media stream can be used for live delivery to the client.

[0111] Due to the various network conditions between the server and the client, each client may be at a different position in the media stream. For example, the first client may be up-to-date (e.g., there is no media content waiting to be delivered to the first client), while the second client may be behind (e.g., there is five seconds of media content waiting to be delivered to the second client).

[0112] In some implementations, each client (also referred to as a streaming client) has an associated logical buffer (also referred to as a buffer) that indicates the client's position in the media stream. In other words, the logical buffer contains the amount of media content waiting to be delivered to the client. In this way, the logical buffer also indicates whether the client is behind the media stream. For example, if a given client has 10 seconds of media content in its logical buffer that has not yet been delivered to the given client, then the given client is 10 seconds behind the media stream. The logical buffer can be implemented in various ways. For example, the media data can be kept in a single memory buffer along with the position of each client's location in a single memory buffer. In some implementations, the collection of logical buffers for the streaming client is called a buffer queue.

[0113] Example 7 - Example environment for low-latency streaming of media content using lossless protocols Figure 6 This is a diagram depicting an example environment 600 for low-latency streaming of media content using lossless protocols. Example environment 600 includes a streaming service 610. Streaming service 610 provides media content for streaming to streaming clients (including streaming clients 630, 632, and 634). Streaming service 610 can be implemented using various types of software and / or hardware computing resources, such as server computers, cloud computing resources, audio and / or video encoding software, streaming software, etc.

[0114] As depicted at 612, the streaming service 610 performs various operations for streaming media streams to a streaming client. These operations may include, for example, receiving audio and / or video data (e.g., samples), encoding the audio and / or video data using various audio and / or video codecs, sending network packets containing the encoded audio and / or video data to the streaming client using a lossless protocol, managing a buffer of the encoded audio and / or video data, and / or selectively discarding audio and / or video data based on various criteria (e.g., buffer size).

[0115] In this example, streaming service 610 is serving media streams to three streaming clients: streaming client 630, streaming client 632, and streaming client 634. Streaming service 610 maintains a buffer (also referred to as a logical buffer) for each streaming client. Specifically, streaming client 630 is associated with buffer 620, streaming client 632 with buffer 622, and streaming client 634 with buffer 624. Each buffer indicates that encoded audio and / or video data is awaiting delivery to its corresponding streaming client.

[0116] The media stream has a current (latest) position, which is located at the head of the stream (e.g., the current position in a live media stream). The head of the stream is indicated by the top of buffers 620, 622, and 624. As new audio and / or video data is generated in the media stream, it is encoded and added to the top of the buffers, as depicted by the arrows leading into buffers 620, 622, and 624. The position of a given streaming client in its corresponding buffer is indicated by the bottom of its corresponding buffer (also referred to as the tail). The amount of encoded audio and / or video data in a given buffer indicates how far behind the streaming client is in the media stream. In other words, the position in a given streaming client's buffer indicates how far behind the given streaming client is in the media stream. For example, if buffer 620 indicates that seven seconds of encoded audio and / or video data is waiting to be delivered to streaming client 630, then streaming client 630 is seven seconds behind in the media stream.

[0117] When it is determined that a given streaming client is behind in the streaming of a media stream, the streaming service 610 selectively discards the media content. For example, streaming client 630 may be behind in the media stream due to the size of its associated buffer 620, as depicted at 626 (e.g., buffer 620 may indicate that seven seconds of encoded audio and / or video data has been queued and is waiting to be delivered to streaming client 630, exceeding a five-second threshold). In another implementation, when a streaming client is behind, instead of discarding the media content, the streaming service may switch to a lower-quality stream in addition to discarding the media content. A given streaming client 630 can be considered to be behind based on various criteria (such as the amount of encoded audio and / or video data in its buffer (e.g., based on the number of frames or other metrics) and / or the amount of time represented by the encoded audio and / or video data in its buffer (e.g., the number of playback seconds represented by the encoded audio and / or video data in the buffer)). In another implementation, the streaming client 630 can be considered to be behind if the latency exceeds a threshold (such as a configurable threshold, including latency due to network loss or bandwidth limitations).

[0118] Streaming service 610 can determine whether a given streaming client is behind by monitoring buffers (e.g., by monitoring buffers 620, 622, and 624). Streaming service 610 can monitor the buffers continuously or periodically. Based on the monitoring, streaming service 610 can determine when a streaming client has fallen behind. In some implementations, streaming service 610 determines that the streaming client associated with the buffer is behind when the buffer indicates more than a threshold amount of encoded audio and / or video data (or the corresponding playback time).

[0119] When streaming service 610 determines that a given streaming client is behind the streaming media stream (e.g., when there is more than a threshold amount of encoded audio and / or video data in the buffer of the given streaming client), streaming service 610 selectively discards audio and / or video data to be streamed to the streaming client. When streaming service 610 selectively discards audio and / or video data, it discards a portion of the audio and / or video data that would otherwise be streamed to the streaming client. For video data, streaming service 610 selects the portion to discard based on scalability information (e.g., discarding one or more layers) and / or long-term reference (LTR) frame information (e.g., discarding one or more frames). For audio data, streaming service 610 is able to adjust the forward error correction (FEC) reliability level (e.g., this results in the discarding of redundant audio data that would otherwise be transmitted to the streaming client). For example, if there are multiple network connections between the streaming service 610 and a given streaming client, audio data (e.g., redundant audio data) from one or more network connections can be discarded.

[0120] Streaming service 610 operates without using per-client quality feedback from streaming clients (e.g., from streaming client 630, streaming client 632, and streaming client 634). For example, when streaming service 610 determines whether a given streaming client is lagging in the streaming of the media stream, streaming service 610 does not use quality feedback from the streaming client (e.g., the streaming client does not send any information to the streaming service indicating that the streaming client is experiencing network loss or latency or decoding slowdown). In some implementations, when determining whether any of the multiple streaming clients is lagging, streaming service 610 uses only server-side streaming information. For example, server-side streaming information may include buffer conditions (e.g., how much encoded media data is in a given buffer waiting to be streamed), network statistics (e.g., packet delivery or loss statistics observed by streaming service 610), and / or other information observed by streaming service 610 but not received from streaming client.

[0121] In some implementations, the streaming service 610 selectively discards video data when the streaming client has fallen behind at least partially based on scalability information. For example, the streaming service 610 can discard one or more video data layers defined by temporal scalability information, spatial scalability information, and / or quality scalability information (e.g., by discarding one or more layers other than the base layer). In some implementations, the scalability information is defined according to the Scalable Video Coding (SVC) standard.

[0122] In some implementations, the streaming service 610 selectively discards video data when the streaming client has fallen behind at least in part based on the LTR frame information. For example, the LTR frame information can indicate the dependency structure of frames (e.g., including Instant Decoder Refresh (IDR) frames, LTR frames, prediction (P) frames, and / or other types of frames). Based on the LTR frame information, specific frames (e.g., predicted frames) can be discarded.

[0123] In some implementations, the streaming service 610 selectively discards video data when the streaming client has fallen behind at least in part based on both scalability information and LTR frame information. For example, one or more layers can be discarded, wherein the layers are scalability layers defined based on LTR frame information (e.g., having a specific LTR dependency structure for the frames).

[0124] Figure 7 This diagram depicts an example environment 700 for selectively discarding media content when performing low-latency streaming using a lossless protocol. Example environment 700 includes a streaming service 710. Streaming service 710 provides media content to streaming clients, including streaming client 720. Streaming service 710 can be implemented using various types of software and / or hardware computing resources, such as server computers, cloud computing resources, audio and / or video encoding software, streaming software, etc.

[0125] As depicted in example environment 700, packetizer 712 receives the media stream as a sequence of encoded audio samples and / or video frames and generates network packets. These network packets are stored in buffer 714 for transmission to streaming client 720. As depicted at 730, streaming service 710 delivers the network packets from buffer 714 to streaming client 720 using a lossless protocol (e.g., as quickly as possible). Depending on various factors, such as network conditions, the number of network packets in buffer 714 (and the corresponding amount of encoded audio and / or video data) may increase or decrease.

[0126] As depicted in example environment 700, streaming service 710 includes network congestion analyzer 716. Network congestion analyzer 716 analyzes buffer 714 to determine whether streaming client 720 is behind (e.g., whether there is more than a threshold amount of encoded audio and / or video data waiting to be sent to streaming client 720 in buffer 714).

[0127] When the streaming client 720 has fallen behind, the network congestion analyzer 716 signals the packetizer 712 to selectively discard media content (also known as throttling), as depicted at 718. For example, throttling can be performed by selectively discarding non-critical scalability (e.g., SVC) layers. In some implementations, multiple layers exist, one or more of which can be discarded based on various criteria, such as the amount of data in buffer 714 and / or other criteria, such as the rate at which buffer 714 is growing.

[0128] When the streaming client 720 has caught up (e.g., when the amount of data in buffer 714 falls below a threshold), the media stream can be restored to its initial state (e.g., media data will no longer be discarded). However, if long-term network conditions (e.g., bandwidth limitations) are detected, long-term changes to the media stream are possible. For example, the streaming client 720 can be moved to a different predefined media stream encoded at a lower video resolution and / or a lower frame rate.

[0129] Example environment 700 depicts one way of implementing streaming service 710. Other implementations may use different components (e.g., components for encoding audio and / or video data, components for generating network packets, components for monitoring buffer conditions, components for transmitting network packets, components for selectively discarding media data, and / or other components). In some implementations, streaming service 610 uses the components depicted for streaming service 710.

[0130] For ease of illustration, example environment 700 depicts a streaming client 720 and its associated buffer 714. However, example environment 700 supports any number of streaming clients and associated buffers.

[0131] In some implementations, streaming is performed by streaming a sequence of encoded audio and / or video frames without grouping the frames into chunks for delivery (e.g., within example environments 600 or 700). In other words, the streaming client receives the encoded audio and / or video data without sending requests for the next chunk of the streaming content (as would typically be done using a content delivery network). As a result, streaming is performed with low overhead and very low latency (e.g., no requests for subsequent chunks of the streaming content from the streaming client).

[0132] In the streaming technology described herein, the media stream streamed to the streaming client is not specific to any given streaming client. In other words, the streaming client does not negotiate streaming parameters specific to the streaming client. However, multiple predefined media streams (e.g., with varying quality and configuration) can be provided by a streaming service (by streaming service 610 or 710), and the streaming service can assign a streaming client to one of the available predefined media streams.

[0133] Using the techniques described above for example environments 600 and / or 700, streaming of audio and / or video content to numerous streaming clients (e.g., 100,000 or more streaming clients) can be performed efficiently and with low latency (e.g., less than one second). Low latency can be maintained by selectively discarding audio and / or video data when a streaming client lags behind. Therefore, using the techniques described herein, a streaming environment can practically support any number of streaming clients while still allowing streaming clients to have meaningful interaction or engagement with the streaming media content when needed (e.g., if the streaming media content is part of a live conference, the streaming client can interact with conference participants in real time).

[0134] In some implementations, multiple network connections (e.g., multiple TCP connections) are used to stream media content to a given streaming client. These multiple network connections can include one or more reliable, semi-reliable, and / or unreliable network connections (also referred to as network channels). For example, a first network connection providing a higher level of reliability (e.g., with more retries) can be used to deliver higher-priority media content (e.g., base layer content), while a second network connection providing a lower level of reliability (e.g., with fewer retries) can be used to deliver lower-priority media content (e.g., layers other than the base layer). In some implementations, multiple network connections (e.g., two or more) are used to stream media content, each with its own reliability level (e.g., number of retries) and configuration (e.g., forward error correction (FEC) configuration).

[0135] In some implementations, quality can be improved by using multiple network connections. For example, if the same content (e.g., encoded audio and / or video data) is sent via multiple connections, the first arriving content can be utilized, thereby reducing latency. In some implementations, redundant media content is distributed across multiple connections (e.g., in a circular manner or using another distribution scheme). Such solutions can reduce latency (e.g., by using the first arriving copy of the media content) and / or improve reliability (e.g., if a given network connection degrades or fails).

[0136] In some implementations, the number of network connections is dynamically determined based on various criteria. For example, the number of network connections can be dynamically determined (and dynamically changed) based on parameters such as round-trip time (RTT), profile information (e.g., location, network type, etc.), machine learning models, etc. In some implementations, streaming begins with a single network connection, and multiple network connections are activated when poor network conditions are detected.

[0137] Example 8 - Example Layer Configuration In the techniques described herein, video data is selectively discarded when a streaming client has fallen behind, at least in part, based on scalability information (e.g., temporal scalability, spatial scalability, and / or quality scalability) and / or LTR frame information. For example, a configuration using a scalability layer (e.g., an SVC layer) is possible, where the video frames are also organized according to LTR frame information. Then, when streaming to a streaming client that has fallen behind in streaming the media stream, one or more scalability layers can be discarded.

[0138] To enable the discarding of layers during streaming media content, the concept of layers is used for scalability. Depending on the configuration, video frames of a layer can be discarded without significantly affecting video quality. For example, layers can be encoded using different frame rates, resolutions, or other encoding parameters.

[0139] When layer configurations are created by defining dependency chains (also known as dependency structures) for each layer, LTR frame information can be used. If none of the P frames are received, video playback can be resumed at the next LTR frame. Depending on the frequency of LTR frames, the resumption may be subtle or barely noticeable.

[0140] Figure 8This is a schematic diagram depicting an example layer configuration 800 for streaming video data, which supports selectively discarding layers when performing low-latency streaming using a lossless protocol. Example layer configuration 800 has three layers: Layer 0, Layer 1, and Layer 2. Layer 0 is the base layer and contains IDR frames and two LTR frames. Layer 1 contains P-frames with a specific dependency structure. Layer 2 also contains P-frames with a specific dependency structure.

[0141] Using the example layer configuration 800, video content from the media stream can be selectively discarded incrementally when the streaming client falls behind. For example, if the streaming client is already behind, layer 2 can be discarded, which still allows the streaming client to decode and play the video content at a reasonable quality.

[0142] The following streaming options are available using the example layer configuration 300.

[0143] - Layer 0 + Layer 1 + Layer 2, maximum quality playback (e.g., 30 frames per second (FPS)). - Layer 0 + Layer 1, reduce playback quality (e.g., 15 FPS). - Layer 0, further reduced playback (e.g., 7.5 FPS).

[0144] Example layer configuration 800 is merely one example layer configuration that can be used for streaming data, and it supports selectively discarding layers. Other configurations can use different numbers of layers and / or different frame dependencies (e.g., IDR frames, LTR frames, P frames, and / or different arrangements and / or frequencies of different frame types).

[0145] Example 9 - Low-latency streaming of media content using semi-reliable mode In some implementations, protocols that provide semi-reliable network packet delivery (also known as semi-lossy protocols) are used to stream media content. Examples of protocols that provide semi-reliable network packet delivery include Stream Control Transport Protocol (SCTP) and QUIC.

[0146] In unreliable mode, network packets are delivered without retries. In semi-reliable mode, there are limits on the total number of retries and / or delivery attempt timeouts. Therefore, semi-reliable mode offers a better delivery probability than unreliable mode, but at the cost of potential additional latency.

[0147] In some implementations, the streaming service (e.g., streaming service 610 or 710) uses a semi-lossy protocol instead of a lossless protocol to stream the media stream. For example, when streaming using the SCTP protocol, the streaming service controls the semi-reliable delivery properties and applies different delivery modes based on the type of video content being streamed. In some implementations, the streaming service applies a first delivery mode when streaming individual video frames (e.g., I-frames, LTR frames, and / or other types of individual frames), and applies a second delivery mode when streaming incremental video frames (e.g., P-frames). For example, the first delivery mode can perform up to five retries when delivering individual video frames to the streaming client, while the second delivery mode only attempts once when delivering incremental frames to the streaming client. Other implementations can use a different number of delivery modes (e.g., more than two delivery modes), each for one or more specific types of video frames and having its own reliability properties.

[0148] In some implementations, streaming is performed using a semi-lossy protocol based on layer configuration (e.g., based on example layer configuration 800). For example, the first layer (e.g., layer 0) is able to deliver with higher reliability (e.g., with more retries), while other layers (e.g., layers 1 and 2) are delivered with lower reliability (e.g., with fewer or no retries).

[0149] In some implementations, multiple network connections are used to stream media content to a given streaming client. These multiple network connections can include one or more reliable, semi-reliable, and / or unreliable network connections. For example, a first network connection providing a higher level of reliability (e.g., with more retries) can be used to deliver higher-priority media content (e.g., base layer content), while a second network connection providing a lower level of reliability (e.g., with fewer retries) can be used to deliver lower-priority media content (e.g., layers other than the base layer). In some implementations, multiple network connections (e.g., two or more) are used to stream media content, each with its own reliability level (e.g., number of retries) and configuration.

[0150] When using a semi-lossy protocol, streaming clients can receive packets out of order. Therefore, streaming clients can collect audio and / or video data in a jitter buffer. The size of the streaming client's jitter buffer can be adjusted based on network conditions.

[0151] Example 10 - Example methods for low-latency streaming of media content In any of the examples presented herein, methods for low-latency streaming of media content using lossless and / or semi-lossy protocols are provided. Low-latency streaming in real-time streaming environments can be achieved by selectively discarding media content when the streaming client is already behind (e.g., based on a monitoring buffer).

[0152] Figure 9 This is a flowchart of an example method 800 for low-latency streaming of media content using a lossless protocol. For example, example method 800 can be performed by a streaming service such as streaming service 610 or streaming service 710.

[0153] At point 910, a lossless protocol is used to stream media to multiple streaming clients. The media stream includes encoded video data. The media stream can also include encoded audio data.

[0154] At 920, the status of the streaming clients is checked to determine if any streaming client is lagging behind in the streaming of the media stream. For example, the streaming client buffers can be monitored to determine how much encoded video data is waiting to be delivered. When the amount of streaming data in the buffer of a given streaming client exceeds a threshold, it can be determined that the given streaming client is lagging behind. This determination is made without using per-client quality feedback from the streaming clients.

[0155] At point 930, when it is determined that the streaming client is behind, a portion of the video data to be streamed to the streaming client is selectively discarded. This selective discarding is performed based on scalability information and / or LTR frame information of the video data. In some implementations, one or more scalability layers are discarded. In some implementations, both scalability information and LTR frame information are used to determine portions of the video data to be selectively discarded (e.g., based on layer configurations also defined by the LTR frame dependency structure).

[0156] In some implementations, example method 900 selectively discards audio content (e.g., in addition to selectively discarding video content). For example, audio content that would otherwise be streamed to the streaming client (e.g., redundant audio content) can be discarded (not delivered to the streaming client) when the streaming client has fallen behind (e.g., based on monitoring indications of a buffer of encoded audio content to be delivered to the streaming client). In some implementations, selective discarding of audio content is performed by adjusting the FEC reliability level of the audio data to be streamed to the streaming client.

[0157] Figure 10This is a flowchart of an example method 1000 for low-latency streaming of media content using a semi-lossy protocol. For example, example method 1000 can be executed by a streaming service such as streaming service 610 or streaming service 710.

[0158] At point 1010, a semi-lossy protocol is used to stream the media to multiple streaming clients. The semi-lossy protocol uses multiple delivery modes for corresponding different types of encoded video data, each of which provides a different level of reliability. The media stream contains encoded video data. The media stream can also include encoded audio data.

[0159] At 1020, encoded video data of the first type is transmitted to the streaming client using a first delivery mode. The first delivery mode uses a first number of retries. In some implementations, the encoded video data of the first type is individual video frames (e.g., IDR frames, LTR frames, I-frames, and / or other types of individual video frames).

[0160] At 1030, the second type of encoded video data is sent to the streaming client using a second delivery mode. The second delivery mode uses a second number of retries different from the first number of retries. In some implementations, the second type of encoded video data is incremental video frames (e.g., P-frames, B-frames, and / or other types of incremental video frames). In some implementations, the first number of retries is greater than the second number of retries, which provides increased reliability when transmitting using the first delivery mode compared to the second delivery mode.

[0161] In some implementations, example method 1000 also determines whether any streaming client is behind in the media stream and selectively discards portions of video data when necessary (e.g., based on scalability information and / or LTR frame information). For example, example method 1000 is also capable of performing some or all of the operations described for example method 900.

[0162] In some implementations, example method 1000 uses different delivery modes for delivering encoded audio content (e.g., in addition to delivering encoded video content). For example, different delivery modes can be used to provide different levels of reliability (e.g., by adjusting the FEC reliability level of the audio data to be streamed to the streaming client). Example method 1000 can also selectively discard audio content (e.g., in addition to selectively discarding video content). For example, audio content that would otherwise be streamed to the streaming client (e.g., redundant audio content) can be discarded (not transmitted to the streaming client) when the streaming client has fallen behind (e.g., based on monitoring of the buffer of encoded audio content to be delivered to the streaming client). In some implementations, selective discarding of audio content is performed by adjusting the FEC reliability level of the audio data to be streamed to the streaming client.

[0163] Example 11 - Example media block format supporting client-side interleaving and peer-to-peer distribution This disclosure provides a packet format, referred to as a block (having a block type), that can be used with the described streaming technology. As described, the disclosed streaming technology can operate by sending track setup information to a client, which can be a format or protocol other than the protocol used to deliver streaming content to the client. For example, the track setup information can be provided via HTTP, including in response to a client's HTTP request. In other scenarios, the track setup information can be sent using the same protocol used to provide streaming media blocks, but in a different block than the block containing media samples (such as, as previously mentioned, for example, frames of audio or video). Thus, consistent with the previous discussion, the disclosed technology can provide a client with information that can be used to participate in media streaming, but using a "push" protocol, where the client sends media samples in a media block without requesting the streaming media server.

[0164] Figure 11 A general block structure 1104 that can be used in the disclosed technology is illustrated. A given block can include a block type identifier 1108, followed by one or more attribute identifiers 1112 (shown as attributes 1112a, 1112b) and their corresponding attribute values ​​(1114a, 1114b). The block can then include a payload 1118. The attributes of attributes 1112 and payload 1118 can differ between different block types.

[0165] The disclosed streaming protocol can include all block types described herein, can omit specific block types, or can include all block types except those specified herein. Figure 11Block types other than the block types shown, or those that can be included Figure 11 One or more block types are shown, but blocks can be implemented in ways different from those shown and described. Having a block type identifier 1108 and an attribute identifier 1112 would be useful, including for providing backward or cross-compatibility. For example, streaming clients could simply ignore block types and attribute types that are not programmed to be recognized.

[0166] In some implementations, media streaming provides different types of content, and the streaming client can be provided with one or more content streams. Specifically, the disclosed techniques provide separate streaming of audio samples, video samples, or text samples. For example, in Figure 11 As shown, Table 1130 lists various types of block types that can be used in the disclosed technology, wherein column 1134 lists the different block types and column 1138 contains a description of the block types.

[0167] Lines 1142a-1142c provide track setup blocks for audio tracks, video tracks, and text tracks, respectively. As described, track setup can be used by a client at the beginning of a streaming session, whereby the client subsequently automatically receives appropriate blocks providing samples of a given track type. As shown in the common description of unit 1146 for lines 1142a-1142c, track setup information can include information that can be used to initialize the media codec for a given track type, including information specifying the relevant codec to be used. Audio and video tracks can include information such as bitrate and resolution, while setup information for text tracks can include human-readable (text) comments. Text track setup information can include the encoding scheme for the text.

[0168] Lines 1142d-1142f relate to blocks containing actual media samples, which are the main block types used during streaming. As described in unit 1148, the sample blocks can contain unencrypted or encrypted media samples. When a stream contains multiple media sample types, different connections between the provider and the streaming client can be used for different media types, or multiple media types can be multiplexed on the same connection. Furthermore, the streaming client can receive the same or different media types from different providers, which can be other streaming clients or one or more main provider nodes.

[0169] Another block type, represented by line 1142g, provides decryption key information, specifically mapping the local encryption key ID to a GUID for the encryption key. This block type can be used for key rotation in sample encryption / decryption, where the same local key ID is used, but can be mapped to different encryption key GUIDs.

[0170] Line 1142h specifies the block type for streaming messages. As indicated by the corresponding entry in column 1138, streaming messages can be any string, including strings providing messages related to the state of the stream, including identifiers of alternative streaming sources. In some cases, streaming clients can use the content of streaming messages to obtain new / updated track setting information. Streaming messages can also be used to provide information about connection quality, such as information provided by provider nodes about the connection quality they perceive, for potential use by streaming clients or downstream provider nodes (which can also be streaming clients). In peer-to-peer implementations using the disclosed techniques, which will be described further, streaming messages can be used for purposes such as distributing hash keys or providing information about nodes within a cluster, including which nodes act as seeders.

[0171] Line 1142i relates to publishing a file message, while line 1142j relates to deleting a file message. As indicated in unit 1150, these block types can be used when transferring files (such as files for HLS or DASH protocols) using the described protocols. That is, the block types can be used as commands to publish or delete specific blocks or groups of blocks. Where the file-based content is provided by multiple providers, a streaming source identifier can be included in the message, which can be a monotonically increasing number, wherein the number increments for each new streaming source.

[0172] The publish file block 1142i can include files with media content such as HLS or DASH files, while the delete file block 1142j can be used to indicate that a file is no longer in use. File deletion is particularly useful for cache management.

[0173] In specific examples, file-based streaming can be used with the exposed chunk format, either by sending media samples that are not of the media sample chunk type (lines 1142d-1142f), or by sending media samples of the media sample chunk type. This is useful when the live stream has been converted to HLS or DASH chunking. Although the media content chunks are sample-based rather than chunk-based, they can still be delivered in "push" mode instead of being requested by the streaming client for additional chunks.

[0174] As an example of how the disclosed technology can be used with the disclosed block format, a media source can provide encoded samples and create media segments, playlists, and catalogs from them. Files corresponding to the media samples can be payloads in the disclosed block format, such as for transmission between intermediate server nodes in a cloud environment. In some cases, these blocks are provided to content delivery networks for delivery to streaming clients.

[0175] The delay-tracking block type is provided in line 1142k. As previously described and as will be discussed further, streaming clients can receive blocks from multiple sources, such as one or more client delivery nodes, where one or more client nodes obtain blocks from a primary delivery node. Additionally, streaming clients can receive blocks for a given stream through multiple paths capable of varying degrees of indirection. As described in column 1108, delay-tracking blocks can include “coordinates” for transmission points in the path, such as a node identifier, and a timestamp associated with the block being processed by the transmission point. This information allows for estimation of transmission latency, enabling streaming clients to switch, add, or remove streaming providers. In one implementation, whenever a given delay-tracking block is processed by a node, the node adds its coordinates (identifier) ​​along with the timestamp to a list of coordinates.

[0176] The end of the streaming block type in line 1142l indicates the end of the stream, which allows the streaming client to perform session termination operations such as stream termination, stream cleanup, and resource release. In the case of the disclosed block format used for file-based streaming, the end of the streaming block type can be used to indicate that a final playlist should be generated and published.

[0177] Figure 12 Provides the means to be included in media sample blocks (e.g., Figure 11 Table 1200 contains additional information about media sample blocks 1142d-1142f of Table 1130. For example, the information in Table 1200 can correspond to attribute 1112 of block format 1104. Table 1200 includes a column 1202 indicating the location of attributes in the block header information, a column 1204 identifying a specific attribute 1112, and a column 1206 providing a description of the corresponding attribute.

[0178] In line 1212a, the first attribute of the set of metadata attributes represented in Table 1200 provides a track identifier 1204a. The track identifier 1204a can be used to multiplex individual streams to provide a total / composite stream. For example, a specific total stream can be associated with audio and video tracks, which are separate streams multiplexed during playback. In some cases, an identifier for a specific track is provided in the request to join a streaming session, and track setup information can be provided for those streams.

[0179] Line 1212b provides a timestamp attribute 1204b for samples within a block, which can be measured in microseconds in some cases (especially when the samples correspond to individual frames of a video stream or to corresponding audio or text content for a video frame). The value of timestamp attribute 1204b can be used for track synchronization, ensuring that buffering results in the same latency across all streaming tracks.

[0180] Lines 1212c and 1212d can be used to encrypt related information. Specifically, the Key ID attribute 1204c indicates which encryption key the streaming client should use. A seed can be provided as a value for attribute 1204d, which can be used to recover the initialization vector for the encrypted sample. In some cases, the seed can be provided in the initial block, and new seeds can be provided subsequently as part of a key rotation.

[0181] Figure 13 Table 1300 provides example content for media sample block types, particularly for audio samples. For other types of media samples, similar information can be included.

[0182] Table 1300 includes column 1304, which identifies the location within the block for a specific piece of content identified in column 1308. Column 1312 provides the data type for the content, while column 1316 provides a description of the content.

[0183] Table 1300 is shown as having rows 1320a-1320f. Row 1320a provides information on the block size 1308a for the audio sample block (such as measured using the Vint data type). The block size content 1308 can include the size of the media sample itself, as well as the metadata / header information for the block and any padding data, as will be described further.

[0184] Line 1320b provides information for object type 1308b of the block, which can correspond to Figure 11The block types in Table 1100. For example, object type 1308b can have a value 1 indicating the audio sample type or a value 2 indicating the video sample type. Other values ​​can indicate text media samples or other media types that can be used in various streaming applications.

[0185] As mentioned in comment column 1316 regarding line 1320b, object type 1308b can be used by the media player for various purposes, such as synchronizing different sample streams (e.g., synchronizing audio and video blocks for a composite stream). The media player can be configured to ignore block types it does not recognize. For example, suppose the media player is programmed to combine audio and video blocks for playback, but not to recognize object type values ​​for text content. In this case, the media player can simply ignore blocks for text content. However, different or newer media players can be programmed to recognize blocks with text content types and synchronize those blocks with blocks for audio and video samples.

[0186] Lines 1320c-1330f represent specific attributes of the media sample, such as attribute 1112 of block format 1104. Each attribute can include an attribute descriptor and an attribute value. The value can be a binary or integer value, but other types of values ​​can also be used. In some cases, a hexadecimal type representation can be used, where the left byte provides the attribute descriptor / type (with values ​​0-F), and the right byte indicates the payload type (with values ​​0-F). For example, the attribute descriptor can have a value of 0x40, where 4 indicates the attribute type, and where 0 indicates the data type used to represent the value for the attribute.

[0187] Line 1320c corresponds to the payload size coverage attribute, which can have a 0x10 attribute descriptor, where 1 indicates the coverage type and 0 indicates the payload type is VInt. Attribute values ​​represented as VInt values ​​can follow the attribute descriptor. For example, having a payload size coverage value is useful when the encryption scheme uses a fixed block size and the payload is smaller than that size. That is, it is possible to use data to fill samples to meet the block size, and the payload size coverage can be used to determine the size of the actual payload (corresponding to the media sample).

[0188] As will be further described, the disclosed technology facilitates operations during media playback or, in some cases, at provider nodes (such as master provider nodes or client provider nodes). For example, components of the computing system (node) can include functions for rearranging blocks containing media samples or the samples themselves, discarding / ignoring duplicate blocks / samples, reordering blocks / samples, or compensating for missing content / samples. A media block (containing audio sample blocks) can include attribute 1308d of line 1320d, the value of which can be used to sort the block (and its associated samples). This attribute 1308d can be referred to as a recovery index. In a particular example, the attribute recovery index has an attribute descriptor 0x20, where 2 indicates that the attribute is an attribute recovery index, and 0 indicates that the attribute payload is represented as a Vint value.

[0189] The actual value can represent a sequence value for a specific media track, which, depending on the implementation, can be specific to a particular media track or correspond to a sequence for one or more media types (such as for the overall media stream). For example, audio samples can have a higher sample rate than video samples, and therefore, even if the audio is to be synchronized with a particular video sample, the sequence number of the sequence for that particular audio sample in the stream may be different from the sequence number of the sample in the video stream.

[0190] The restored index attribute 1308d can be used as a... Figure 12 The purpose of the timestamp complement in row 1212b of Table 1200 is as follows: That is, the timestamp corresponding to the presentation time can be used to synchronize samples for playback, while the recovery index attribute 1308d can be used for more “transmission” related issues, including sorting sample blocks, discarding sample blocks, or identifying lost sample blocks.

[0191] One advantage of the sequence numbers provided by the recovery index attribute 1308d is that they can be quickly determined from block metadata and directly analyzed for sorting, duplicate detection, etc. For example, sequences can be determined directly without having to subtract the timestamp for the sample in the block from the reference timestamp.

[0192] It should be noted that the recovery index attribute 1308d differs from other sequence numbers that can be used. Specifically, the sequence number of attribute 1308d is different from sequence numbers that might be associated with packets (such as TCP packets) used to transport blocks. Blocks can be split into multiple TCP packets and transmitted, wherein packets can have numbers used to manage transmission, including reassembling them into packets to provide blocks. A separate sequence number is used for said blocks, which is available after the blocks are reassembled after transmission is complete, and is used to manage blocks as blocks, rather than using blocks to manage / order samples.

[0193] Note that the sequence identifier of attribute 1308 is used even when using a reliable transport protocol, rather than for sorting packets using lossy or semi-lossy protocols. Furthermore, note that sorting is performed using sequence numbers instead of information such as timestamps or information based on media samples (such as attributes of video frames, such as whether they are I-frames, p-frames, reference frames, etc., for sorting).

[0194] Line 1312e corresponds to wall clock attribute 1308e, which can have a descriptor of 0x40, where "4" identifies the attribute as a wall clock attribute, and "0" again indicates that the wall clock attribute value is provided as a VInt data type. In this case, the wall clock value corresponds to the point at which the sample is ingested into the delivery pipeline. The wall clock value can be compared with the rendering time at the playback component (the time at which the sample is intended to be played back at the streaming client) to determine media delivery latency.

[0195] That is, latency can be determined in several ways. In one way, the wall clock value of line 1312e can be subtracted from the real-time clock value at the playback component of the streaming client, where the result indicates media delivery latency. In other words, the offset value for the streaming client can be calculated by adding the presentation time (the offset from the start of streaming at the streaming source) to the real-time clock value recorded by the streaming client at the start of playback. This value can be subtracted from the current real-time clock value to provide another measure of media delivery latency.

[0196] In some cases, one or more techniques for calculating media delivery delays can be used. For example, it is possible to periodically send the wall clock value of line 1312e, such as every few seconds, to allow for more accurate delay measurements, and to use the presentation time in between. In other examples, only one of the techniques is used, but the value is not provided in each block. For example, it is possible to provide the wall clock value once every few seconds, with this feature omitted in intermediate blocks.

[0197] In this regard, note that the format described in Table 1300 allows the use of attributes to be optional. If a streaming client does not recognize an attribute, it can simply ignore it. Similarly, a streaming client can include logic to identify whether a particular attribute is optional or only provided periodically. For example, even when a streaming client is configured to use the wall clock value of line 1312e for media delay calculation, the streaming client is also configured to process the block when a value for that attribute is available, but omit the media delay calculation when no wall clock value is provided, or perform such calculations using a different technique.

[0198] Row 1312f of Table 1300 provides information about media data attributes. Media data attributes can have a 0x30 descriptor, where "3" indicates the media data attribute and "0" indicates a value provided as a VInt data type. In a particular embodiment, the media data attribute includes an attribute description, a series of metadata values ​​represented as VInt values, and a series of bytes corresponding to a media sample, up to the end of the block. The attribute description and values ​​can correspond to... Figure 12 Those provided in Table 1200.

[0199] Specifically, different metadata values ​​can include a VInt value providing a track ID, which identifies which media stream a given sample is associated with. Another VInt value can provide a timestamp of the presented sample, such as in microseconds. An encryption key identifier is provided in another VInt value. In some cases, a single key can be used to decrypt all individual streams in a composite stream for a specific streaming source. An IV seed can be provided in another VInt value. For unencrypted samples, the key ID and IV seed can be set to zero.

[0200] In some cases, it can be useful to encrypt media metadata and media data separately, and therefore, it can be useful to have encrypted information for media tracks and encrypted information for media data. For example, accessing track metadata to determine whether a particular block should be used can be useful, where the media data is decrypted later if media samples in the block are to be processed for playback.

[0201] Combination Figures 11-13 The described streaming format offers various advantages. Unlike files in protocols such as HLS or DASH, the disclosed technique, by delivering samples, does not add additional latency beyond that required for the underlying codec. That is, the block format does not require parsing of blocks. Furthermore, a more limited amount of metadata is needed because the metadata required for streaming rendering can be provided at the track / stream setup and does not need to be included in subsequent blocks, particularly subsequent media sample blocks, which are intended to be the main block type used after the streaming is initiated. Media sample blocks can simply indicate the specific type of sample stream they belong to, such as for multiplexing or interleaving.

[0202] As previously stated, unlike the request / response model used in technologies such as HLS or DASH, the disclosed technology operates on a "push" basis. The push nature of communication reduces both latency and network traffic—saving both computational resources in terms of processing (since there is no need to generate or process requests) and network resources (since there is no need to send requests).

[0203] Furthermore, a given stream can easily fork at specific nodes, either as part of converting a client node into a provider node or by adding another streaming client to a node that is already acting as a provider node (which can include streaming clients acting as delivery nodes, as in peer-to-peer implementations). In some cases, such as text for at least a specific type of audio or text sample, forking can be initiated at arbitrary sample blocks. In the case of video samples, specific types of samples (such as keyframes) may be needed before forking can occur, but these frames occur periodically, thus introducing minimal forking latency.

[0204] The block format supports end-to-end encryption, including the use of key rotation, where encryption can be applied at multiple levels. The blocks providing the streaming message allow for or can be modified during streaming, such as changes in stream resolution, where new setup information can be provided to allow the codec to be reinitialized for updated stream parameters.

[0205] The disclosed techniques can support fixed or variable sampling rates, such as for audio or video. However, in some cases, a fixed sampling rate can be assumed, which can help optimize bandwidth usage.

[0206] Example 12 - Example Peer-to-Peer Media Streaming The disclosed technology also supports peer-to-peer delivery of content using a sample-based (rather than chunk-based) scheme, where clients receive samples via a "push"-based protocol rather than a request / response protocol such as HLS or DASH (including those used with Enterprise Content Delivery Networks (EDCN)). That is, while EDCN can provide peer-to-peer implementations of HLS or DASH, the protocol still uses a request / response implementation, which suffers from the aforementioned drawbacks.

[0207] Another advantage of the disclosed block formats is that they are "read-to-fork" for use with one or more streaming clients. That is, the block formats include media samples and, unlike file-based protocols such as HLS or DASH, do not require "truncation" before blocks can be forwarded. Furthermore, metadata is typically sent only during stream initialization, rather than (as in HLS and DASH) in each block (which needs to be updated before forwarding). This simplifies the forking process and reduces latency. However, new metadata can be sent when specific events occur, such as when streaming has been restarted, for example, when the stream resolution has changed. This new metadata can be used to reinitialize the codec at the streaming client.

[0208] When a peer-to-peer implementation is used for a specific stream, all streaming clients can participate in the peer-to-peer network, such as seeders or leaves, or some streaming clients can receive content in another way, such as directly from the master delivery node or from client nodes, such as regarding... Figure 1 As described.

[0209] Figure 14 The illustration shows an example peer-to-peer computing environment 1400 capable of implementing the disclosed technology. The computing environment 1400 may include a media source 1410 and a master delivery node 1414, which can correspond to... Figure 1 The media source 110 and the primary delivery node 120. The primary delivery node 1414 communicates with one or more groups 1418. A given group (such as group 1418a) includes multiple streaming clients 1422 (shown as streaming clients 1422a-1422j). Technically, the primary delivery node 1414 communicates with one or more streaming clients in the 1422 of group 1418.

[0210] Group 1418a illustrates components for an example streaming client 1422a, which can represent streaming client 1422. In a particular implementation, streaming client 1422a can communicate with a server (such as master delivery node 1414) and with other streaming clients 1422. At various times, a given client 1422 can act as a seeder or a leech. As used herein, a seeder refers to a streaming client 1422 that sends streaming blocks to other streaming clients. A leech refers to a streaming client 1422 that only receives streaming blocks from other streaming clients. A given streaming client 1422 can act as both a seeder and a leech, but will still be referred to as a seeder. That is, streaming client 1422 provides blocks to other streaming clients, but also receives blocks from other streaming clients without a direct connection to master delivery node 1414.

[0211] Client 1422a is shown as including server connector 1430, peer seed component 1432, and peer leech component 1434. Server connector 1430, peer seed component 1432, and peer leech component 1434 can communicate with other components (master delivery node 1414 or other streaming clients 1422) using the same network protocol or different network protocols. Furthermore, these components may be able to communicate using various protocols, and specific protocols may be used for specific connections. In a particular example, communication between streaming client 1422 using server connector 1430 and master delivery node 1414 may occur using HTTP long polling, TCP, UDP, QUIC (Fast UDP Internet Connection), SRT (Secure Reliable Transport), RTMP (Real-Time Messaging Protocol), or gRPC (gRPC Remote Procedure Call). Communication between streaming clients can use protocols such as TCP, UDP, WebRTC (Web Real-Time Communication) data channels, QUIC, SRT, RTP (Real-Time Transport Protocol), WebSocket, or BitTorrent.

[0212] Server connector 1430 is used by a streaming client to connect to master delivery node 1414, which can occur as described elsewhere in this disclosure, including as per [the description of]... Figure 1 and Figure 2 As described. Peer seed component 1432 includes functionality for forming content streams to provide content to multiple peers, wherein forking capability can be as described regarding Figure 6 And as described in Examples 5 and 6. More specifically, the peer seed component 1432 can include logical buffers to various receivers, or can have a single buffer with pointers to buffers maintained for each receiver. The peer seed component 1432 can also perform operations such as discarding blocks when a receiver has been determined to be too far behind in block delivery.

[0213] Peer-leeching component 1434 is configured to connect to and receive blocks from the same media stream from multiple streaming clients 1422. Peer-leeching component 1434 and peer-seeding component 1432 have access to directory 1440, which includes information about available peers for a particular stream. Directory 1440 can be implemented as in other peer-to-peer protocols, including at least part of a distributed hash table of streaming clients 1422 in queue 1448. Directory 1440 can also include a routing table, or streaming client 1422a can otherwise include a routing table. In a particular implementation, the routing table provides an identifier for a specific target client in directory 1440 for the “next-hop” streaming client.

[0214] The peer-to-peer component 1434 can use the connection manager 1444 to manage connections to other clients. For example, the connection manager 1444 can perform operations such as identifying new streaming clients 1422 that can be used to seed a specific stream, prioritizing specific seed streaming clients 1422, or discarding connections to streaming clients. For example, a seed streaming client 1422 used as a seeder can be discarded if the time between blocks exceeds a threshold (or if a specific block is not received from the seed streaming client), or if the aggregate latency metric for the seed streaming client (e.g., average latency over a period of time) exceeds a threshold.

[0215] and Figure 1 Compared to the computing environment 100, the advantages of the disclosed peer-to-peer network implementation are not only that a single streaming client 1422 can receive data streams from multiple streaming clients, but also that such sources / connections can change over time, yet a given streaming client 1422 can simultaneously receive streaming data from multiple streaming clients. Streaming client 1422a is shown to include an interleaver component 1450. The interleaver component 1450 is responsible for assembling a stream of chunks that can be provided to a media player 1454 (such as a web browser or other type of streaming client) for streaming playback. Operations performed by the interleaver component 1450 can include sorting chunks received from one or more clients into a sequential stream, such as using... Figure 13 The recovery index value of attribute 1308d in table 1300.

[0216] Interleaver component 1450 can include a buffer where blocks are held for a period of time, enabling the provision of an ordered sequence of blocks to the media player. For example, suppose blocks numbered 3, 4, and 6 are received. Interleaver component 1450 can hold blocks 3, 4, and 6 in the buffer for a period of time until block 5 is received. Then, interleaver component 1450 can place the blocks in the correct order and provide them to media player 1454. Note that the buffer (in...) Figure 15 The buffer shown as buffer 1510 is different from a buffer used as part of a transport protocol such as TCP because a buffer contains assembled blocks after delivery is complete, rather than specific packets of a single block.

[0217] When no block is received within a threshold time period, the interleaver component 1450 or another component of the streaming client 1422a can take various actions. These actions can include discarding a specific playback time or discarding a specific sample type from the playback time. For example, suppose that for a specific playback time, audio samples have been received, but the corresponding video samples have not yet been received. The media player 1454 can be provided with a sample stream that includes only the audio samples. The interleaver component 1450 or other components can also perform actions such as copying samples. In the example above, suppose that audio samples 3-6 have been received, but only audio samples 3, 4, and 6 have been received. In one implementation, the interleaver component 1450 can send ordered audio samples and corresponding ordered video samples to the media player 1454, but copy video sample 4 to compensate for the lack of video sample 5.

[0218] If a connection to one or more seeding streaming clients 1422 is interrupted, the leeching / receiving streaming client 1422 can resume its connection to the master delivery node 1414, such as using server connector 1430. If a suitable number of connections (which can include a single connection) to the streaming client 1422 acting as the seeder have been established, the streaming client 1422a can return to the receiving block via the peer leeching component 1436.

[0219] At least initially, the overlay network can be generated by the network generation and management component 1460 of the primary delivery node. The network generation and management component 1460 can perform actions such as initially forming a group 1418 from the streaming clients 1422, using information such as physical or network location, or based on connection speed or reliability. The overlay network can also include an initial set of routing tables used by the streaming clients 1422, or distribute keys for a distributed hash table among the streaming clients.

[0220] Depending on the desired level of centralized oversight, the network generation and management component 1460 can also handle actions such as adding or removing streaming clients from a group. The network generation and management component 1460 can also perform actions such as updating routing tables or hash keys for streaming client 422. In other scenarios, at least some of these actions can be performed directly by streaming client 1422.

[0221] In some cases, utilizing the functionality of an existing or otherwise external peer-to-peer network manager 1464 may be useful. This peer-to-peer network manager 1464 can include network generation and management components 1460. Using an existing or external peer-to-peer network manager 1464 can avoid replication functionality on the primary delivery node 1414. It can also be used to reduce the computational or networking load on the primary delivery node 1414.

[0222] Figure 14 The diagram illustrates various types of connections between streaming client 1422 and other streaming clients, or between a streaming client and master delivery node 1414. Arrows between streaming clients 1422 and between a streaming client and master delivery node 1414 indicate possible communication connections. Dashed leads indicate inactive connections, while solid leads indicate active connections. In the case of streaming client 1422, solid leads thus indicate that the streaming client is acting as a seeder, while dashed connections indicate that the streaming client is acting as a leech. The connections shown are dynamic; as explained above, streaming client 1422 can switch between seeder and leech states. Similarly, in some cases, streaming client 1422 can have a connection to master delivery node 1414 that can be a connection that introduces blocks in a group as part of “normal” operation, or that can be established when streaming client 1422 cannot connect to its peer streaming client 1422. In this "normal" operation, a streaming client 1422 with a connection to the master delivery node 1414 can act as a seed for other streaming clients, but can also act as a leech. For example, in some cases, content can be received from a peer streaming client 1422 before being received from the master delivery node. Alternatively, a streaming client 1422 can have a connection to the master delivery node 1414 for one media type, but can receive blocks for other media types from other streaming clients.

[0223] More specifically, in Figure 14 In the diagram, streaming clients 1422a, 1422b, 1422c, 1422e, 1422f, and 1422g act as seeds because they have outgoing solid arrows. Streaming clients 1422h, 1422i, and 1422j act as leeches because they have no active outgoing connections. Streaming client 1422d represents a scenario where streaming client 1422 does not communicate peer-to-peer with other streaming clients and has therefore established a fallback connection directly to the master delivery node 1414.

[0224] Note that even if the disclosed technology is not implemented in a peer-to-peer environment, streaming client 1422 may still have one or more of the components shown for streaming client 1422a. For example, the streaming client may omit peer seed component 1432 and peer leeching component 1434, or such components may be included but inactive.

[0225] Example 13 - Example Interleaving Operation Figure 15 The diagram illustrates what can be achieved by Figure 14 The operations performed by the interleaver component 1450. Specifically, Figure 15 The diagram illustrates how the interleaver can be responsible for sorting samples and handling missing or duplicate samples. It should be noted that... Figure 15 This allows for simplified operation of the interleaver component 1450. That is, it describes a single type of media content. Figure 15 The interleaver component 1450 is capable of interleaving different types of media content. For example, the interleaver can be responsible for creating sample streams that are output to a media player, where the samples have different media types, such as a mixed stream with audio and video samples. (Regarding...) Figure 16 It describes the functionality for handling different types of media samples.

[0226] Go to Figure 15 The interleaver component 1450 is capable of maintaining a buffer 1510 that can be used for temporary storage of media samples 1514 (shown as media samples 1514a-1514g). For Figure 15 For the purpose of this study, it is assumed that all media samples 1514 are of the same type, such as audio samples or media samples. The illustration of media sample 1514 can be simplified, wherein buffer 1510 stores blocks, wherein said blocks contain information about... Figure 13 The described recovery index value (sequence identifier) ​​is used to sort and process / manipulate samples / blocks. As an example, sample 1514a is also shown as having its corresponding block 1516a, and sample 1514d is shown as having its corresponding block 1516d.

[0227] exist Figure 15 In the example scenario, Figure 14The interleaver component 1450 of a specific streaming client 1422 receives media samples 1514 (in corresponding blocks, such as blocks 1516a and 1516d) from multiple providers, including a first seeder 1522a and a second seeder 1522b. The interleaver component 1450 assembles a coherent stream from these media samples 1514. For example, the interleaver component 1450 is capable of placing samples in a buffer 1510 regardless of which seed streaming client 1422 the sample was received from—the first received sample is buffered. Although shown as including two seeders 1522a, 1522b, a given streaming client can receive blocks from more than two seeders, can receive blocks from one or more seeders and one or more connections to the master delivery node, can receive blocks from one seeder and one or more connections to the master delivery node, or can receive blocks from multiple connections to the master delivery node without receiving from the streaming client. Figure 15 The scenario can represent a peer-to-peer delivery scenario, such as regarding Figure 14 As described.

[0228] In the example scenario, a first sample 1514a and a second sample 1514b are first received from a first sower 1514a, and then placed in a buffer 1510 in a sequential order (e.g., based on a wall clock or presentation timestamp included in the corresponding block containing samples 1514a and 1514b, such as having information about...). Figures 11 to 13 The block described in the format.

[0229] Now, suppose interleaver 1450 receives a second block containing second sample 1514b from second seeder 1522b. Since buffer 1510 already contains the second sample 1514b received from first seeder 1522a, interleaver 1450 ignores the additional copy of the second sample and does not place it in buffer 1510. If the second sample 1522b has already been received from second seeder 1514b, then a copy of the second sample received from the second seeder will be placed in buffer 1510, and the copy of the second sample received from first seeder 1514a will be ignored.

[0230] The third sample 1514c is received from the second sower 1522b and placed in the buffer 1510 in the correct order.

[0231] The interleaver component 1540 is capable of handling cases where no samples are received within a specific time frame. That is, there may be a sample 1514 that can reside in the buffer 1510 for a limited time. If no sample 1514 is received in that time frame, the interleaver component 1540 can perform actions to account for the lost sample, or can release the sample from the buffer (such as releasing the corresponding block containing such a sample), even if one or more samples in the sample sequence are lost.

[0232] In some cases, the interleaver component 1450 can replicate sample 1514 to account for missing samples, such as using sample generator 1544. Specifically, Figure 15 The illustration depicts the case where a fourth sample is not received. In this case, the immediately preceding sample (third sample 1522c) is copied to provide a sample for the next sample position 1530a in the sequence. The exact sample or sample 1522 used to generate the sample to replace the missing sample for sequence position 1530a can depend on the specific type of sample (such as whether the sample is an audio sample or a video sample), the specific implementation technology, and the contents of buffer 1510. Taking a video sample as an example, the missing sample can be replaced by a copy of the third sample 1514c, where the third sample can correspond to an I-frame, P-frame, or B-frame, provided that a reference I-frame is available when the third sample is a P-frame or B-frame.

[0233] As another example of how to handle missing samples, consider sequence position 1530b. If no sample is received for position 1530b, it can be left as an empty sequence position in the stream. In a particular implementation, the media player 1540, which receives sample 1514 from buffer 1510, includes functionality for compensating for missing samples. For example, for video samples, the media player 1540 can include error concealment algorithms, such as algorithms for estimating or interpolating missing frames, which can use adjacent samples 1514 in the process. For audio samples, missing samples can be compensated for by inserting silent audio samples. In other words, the sample stream contains samples that meet the expected sampling rate, but audio samples are presented as silent. Similar techniques can be used for video samples.

[0234] In various scenarios, even if the samples are in the correct order, samples can be omitted from the stream. For example, consider sample 1514f. When the sample is a P-frame of a reference I-frame, sample 1514f can represent a sample that depends on another sample (such as for a video sample). If buffer 1510 does not include the I-frame required for the P-frame of sample 1514f, then sample 1514f can be omitted from the buffer / stream, and gaps can be filled using gap or compensation techniques in the buffer, as described for sample 1514c.

[0235] The buffer 1510 can have a fixed size or a variable size. In some cases, the buffer 1510 is initialized with a default size (such as 200 microseconds) but can grow depending on input conditions. However, typically, as discussed, the sample 1514 is held in the buffer for a maximum time, after which one of the compensation techniques described above can be used, or the sample stream can be provided to the media player 1540 with missing samples. Furthermore, although compensation techniques for missing samples have been described as being performed by the interleaver component 1540, at least some of these techniques, such as compensating for missing samples, can be performed by the media player 1540.

[0236] Figure 16 The diagram illustrates the process of... Figure 14 The interleaver component 1450 provides additional functionality. That is, media players (such as...) Figure 14 A media player (1454) might expect a large number of runs of different types of media streams that do not include samples of one type of media. For example, a media player that renders synchronized audio and video might expect audio and video samples to be provided relatively close to each other in the stream for a specific rendering time. Given the sample-based properties of a disclosed chunk format (as opposed to file-based properties), samples can reach relatively long runs of the same media type. This can be particularly significant in peer-to-peer implementations or implementations that allow streaming clients to receive chunks from multiple sources (provider nodes).

[0237] Therefore, the interleaver component 1450 is able to manage the sample buffer to reorder the media sample types (also known as "modals") to distribute the media sample types more evenly across the stream provided to the media player. Figure 16 The illustration shows sample buffers 1610a and 1610b. Sample buffer 1610a shows the initial ordering of two types of media samples. Sample buffer 1610b is the same buffer as 1610a, but after the samples in the sample buffer are reordered by the interleaver component 1450. Buffers 1610a and 1610b can correspond to... Figure 15 The specific state of cache 1510. Although not shown, samples in caches 1610a and 1610b can be included in the corresponding blocks, as per the description of... Figure 15 As described.

[0238] The operations described for interleaver component 1540 and buffer 1510 are capable of relating to... Figure 16The described combination of operations. That is, the interleaver component 1450 can concurrently perform operations on a single buffer, such as reordering media samples based on a sequence for a given media sample type, removing duplicate media samples, or compensating for gaps in the stream of media samples, and reordering the stream of samples to distribute samples of different media types more evenly.

[0239] Buffers 1610a and 1610b are shown to include multiple audio samples 1614 and multiple video samples 1618. In buffer 1610a, there exists a relatively long sequence of only audio samples 1614 or a relatively long sequence of only video samples 1618. Interleaving component 1450 reorders at least a portion of the audio samples 1614 and video samples 1618 to provide a buffer state for buffer 1610b, which has a more uniformly distributed sample type.

[0240] Depending on the implementation, the interleaver component 1450 can impose a strict ordering of media sample types, or it can order the sample types more flexibly. Strict ordering can involve ensuring that each audio sample 1614 is followed by a video sample 1618, and then another audio sample. A more lenient ordering can involve the interleaver component 1450 ensuring that consecutive runs of a given media sample type do not exceed a threshold, but where such runs can be broken down into smaller runs of the media sample type, and where such smaller runs can optionally vary in length. Furthermore, although shown as reordering audio and video samples, the interleaver component 1450 can reorder other or additional sample types. For example, in an overall stream including audio, video, and text samples, the interleaver component 1450 can reorder all sample types as appropriate to avoid long runs of a single sample type.

[0241] Example 14 - Example Technical Advantages As described, in general, the disclosed techniques can provide improved media streaming by reducing latency in media transmission. For example, sending media samples without client requests reduces the latency associated with generating and processing requests. Furthermore, because individual media samples are sent in the disclosed techniques, playback can be initiated faster and performed more smoothly (because any media gaps will be smaller). Additionally, the use of individual samples allows for greater interactivity from the streaming client because the lower latency in presenting samples at the streaming client enables faster response to input from the streaming client.

[0242] These improvements are achieved, at least in part, through the use of a unique packing format. This format provides blocks for various media sample types, as well as for block types of non-media sample content (such as system messages or encrypted information). Using separate media block types improves playback compared to file-based streaming technologies where multiple sample types are included in a single file. That is, if a file is discarded or delayed, all media content is delayed. When using different media sample block types, one type of media content can still be rendered even if another type is discarded or delayed. For example, if a video sample is unavailable, an audio sample for the relevant presentation time can still be rendered at the streaming client.

[0243] Media block types are lightweight because metadata can be reduced by having separate block types for media track settings information. This information is only sent at the start of playback or at specific events, and therefore the size of the transmission unit can be reduced compared to file-based technologies, requiring less processing from streaming clients.

[0244] Media sample blocks are associated with sequence IDs for blocks of a specific media block type. The sequence IDs allow individual media blocks within a stream of a specific media type to be quickly sorted or reordered, simply by examining the block's header / metadata without extracting the sample content from the block. Streaming clients can include an interleaver component capable of taking various actions based on the sequence IDs, such as reordering blocks when they are out of order, reordering blocks of different media types to avoid long runs of samples of a single media type, or detecting duplicate blocks. Therefore, the disclosed techniques can provide smoother, lower-latency processing at the streaming client, resulting in improved media playback and facilitating greater interactivity for the streaming client.

[0245] This disclosure also enables peer-to-peer topologies. These topologies provide improved distribution of blocks, including blocks containing media samples, to streaming clients and among streaming clients. By making more sources available, blocks have more transmission paths and can be delivered faster, reducing congestion compared to scenarios with a more limited number of predefined transmission paths. Furthermore, peer-to-peer implementations can provide more reliable delivery because multiple sources are available for the same type of media blocks.

[0246] The disclosed block format supports these peer-to-peer topologies, such as by allowing block reordering or allowing the detection of duplicate blocks as described above.

[0247] Example 15 - Example Operation Figure 17This is a flowchart of process 1700, which involves performing processing operations at the streaming client relative to a block of media samples to facilitate low-latency media streaming. A first data block is received at 1705. The first data block includes a first sequence identifier for a first media type and a first single discrete media sample of the first media type. The first data block does not include media samples of media types different from the first media type.

[0248] At 1710, multiple data blocks comprising corresponding single discrete media samples of the first media type are received. The single discrete media samples of the first media type are ordered in the stream. At 1715, a second data block is received. The second data block comprises a second sequence identifier for the second media type and a first single discrete sample of the second media type. The second data block does not include media samples of a media type different from the second media type. The second media type is a media type different from the first media type.

[0249] At 1720, it is determined that the consecutive number of discrete media samples of the first media type meets a threshold. Based on the determination that the consecutive number of discrete media samples of the first media type meets the threshold, at 1725, a first discrete sample of the second media type is inserted between consecutive media samples of the first media type in the stream. At 1730, a first single discrete media sample of the first media type and a second discrete media sample of the second media type are provided to the media player for presentation at the streaming client. The streaming client receives data blocks for a stream comprising multiple data blocks from multiple sources.

[0250] Figure 18 This is a flowchart of process 1800, which involves performing processing operations at a streaming client relative to a block of media samples to facilitate low-latency media streaming. Process 1800 includes, at 1805, receiving a first data block at a first time. The first data block includes a first sequence identifier for a first media type and a first single discrete media sample of the first media type. The first data block does not include media samples of media types different from the first media type.

[0251] At 1810, a second data block is received at a second time. This second time follows the first time. The second data block includes a second sequence identifier for the first media type and a second single discrete sample of the first media type. The second data block does not include media samples of media types different from the first media type.

[0252] At 1815, it is determined that the second sequence identifier is lower than the first sequence identifier. At 1820, the first single discrete media sample and the second single discrete media sample of the first media type are reordered such that the second single discrete media sample of the first media type is provided to the media player in the stream before the first single discrete media sample of the first media type. At 1825, the first single discrete media sample and the second discrete media sample of the first media type are provided to the media player for presentation at the streaming client. The streaming client receives data blocks for a stream comprising multiple data blocks from multiple sources.

[0253] Figure 19 This is a flowchart of process 1900, which describes the process of performing processing operations at the streaming client relative to a block of media samples to facilitate low-latency media streaming. At 1905, a first data block is received at a first time. The first data block includes a first sequence identifier for a first media type and a first single discrete media sample of the first media type. The first data block does not include media samples of media types different from the first media type.

[0254] At 1910, a second data block is received at a second time. This second time occurs after the first time. The second data block includes a second sequence identifier for the first media type and a second single discrete media sample of the second media type. The second data block does not include media samples of media types different from the first media type. At 1915, it is determined that the second sequence identifier is the same as the first sequence identifier. At 1920, the second data block is discarded, and the second single discrete media sample of the first media type to be presented at the client device is not provided. At 1925, the first single discrete media sample of the first media type is provided to the media player for presentation at the streaming client. The streaming client receives data blocks for a stream comprising multiple data blocks from multiple sources.

[0255] Figure 20 This is a flowchart of a process 2000 in which a streaming client performs operations in peer-to-peer delivery of streaming media content with low latency. At 2005, a first data block is sent to a first receiving computing system. The first data block includes a first sequence identifier for a first media type and a first single discrete media sample of the first media type. The first data block does not include media samples of media types different from the first media type.

[0256] At point 2010, a second data block is sent to the second receiving computing system. The second data block includes a second sequence identifier for a second media type and a first single discrete sample of the second media type. The second data block does not include media samples of media types different from the second media type. The second media type is either the first media type or a media type different from the first media type. The second receiving computing system is either the first receiving computing system or a computing system different from the first receiving computing system.

[0257] At 2015, a first single discrete media sample of a first media type and a second single discrete media sample of a second media type are presented. A first sequence identifier and a second sequence identifier are used to order the media samples for presentation at a first streaming client. The streaming client receives data blocks comprising multiple data blocks from multiple sources.

[0258] Figure 21 A flowchart is provided for a process 2100 that performs processing operations when streaming media content with low latency. A first data block is received at 2105. The first data block includes a first sequence identifier for a first media type and a first single discrete media sample of the first media type. The first data block does not include media samples of media types different from the first media type.

[0259] At 2110, a second data block is received. The second data block includes a second sequence identifier for a second media type and a first single discrete sample of the second media type. The second data block does not include media samples of a media type different from the second media type. The second media type is either the first media type or a media type different from the first media type. At 2115, the first single discrete media sample of the first media type and the second discrete media sample of the second media type are provided to the media player for presentation at the client device. The client device receives data blocks for a stream comprising the first data block and the second data block from multiple sources.

[0260] Figure 22 A flowchart of a process 2200 for delivering streaming media content with low latency is provided. At 2205, a first data block is sent to a first receiving computing system. The first data block includes a first sequence identifier for a first media type and a first single discrete media sample of the first media type. The first data block does not include media samples of media types different from the first media type.

[0261] At 2210, a second data block is sent to a second receiving computing system. The second data block includes a second sequence identifier for a second media type and a first single discrete sample of the second media type. The second data block does not include media samples of a media type different from the second media type, and the second media type is either the first media type or a media type different from the first media type. The second receiving computing system is either the first receiving computing system or a computing system different from the first receiving computing system. The first sequence identifier and the second sequence identifier are used to order the media samples for presentation at a first streaming client. The first streaming client receives data blocks from a stream comprising first and second data blocks from multiple sources.

[0262] Example 16 - Additional Implementation Example 1 is a computing system including at least one hardware processor and at least one memory coupled to said at least one hardware processor. The computing system also includes one or more computer-readable storage media storing computer-executable instructions that, when executed by said computing system, cause the computing system to perform processing operations at a streaming client relative to a block including media samples to facilitate low-latency media streaming.

[0263] The operation includes receiving a first data block. The first data block includes a first sequence identifier for a first media type and a first single discrete media sample of the first media type. The first data block does not include media samples of media types different from the first media type. Multiple data blocks including corresponding single discrete media samples of the first media type are received. The single discrete media samples of the first media type are ordered in the stream.

[0264] Receive a second data block. The second data block includes a second sequence identifier for a second media type and a first single discrete sample of the second media type. The second data block does not include media samples of a media type different from the second media type, and the second media type is a media type different from the first media type.

[0265] A threshold is determined to be met by identifying the number of consecutive discrete media samples of the first media type. Based on this threshold, a first discrete sample of the second media type is inserted between consecutive media samples of the first media type in the stream. The first single discrete media sample of the first media type and the second discrete media sample of the second media type are provided to the media player for presentation at the streaming client. The streaming client receives data blocks for a stream comprising multiple data blocks from multiple sources.

[0266] Example 2 includes the subject matter of Example 1. The first data block is received from a first source among the plurality of sources, and the second data block is received from a second source among the plurality of sources. The second source is different from the first source.

[0267] Example 3 includes the subject matter of either Example 1 or Example 2. In response to a specific request for either the first data block or the second data block, neither the first data block nor the second data block is received.

[0268] Example 4 includes the subject of any one of Examples 1-3. The first data block and the second data block do not include codec configuration information that the media player can use to render a first single discrete sample of the first media type or a first single discrete sample of the second media type.

[0269] Example 5 includes the subject matter of any one of Examples 1-4. The first data block includes a first value for a block type identifier. The first value indicates that the first data block includes a sample of the first media type. The second data block includes a second value for the block type identifier. The second value indicates that the second data block includes a sample of the second media type. The second value is different from the first value.

[0270] Example 6 includes the subject matter of any one of Examples 1-5. Example 6 also specifies receiving a third data block. The third data block includes a mapping from a GUID for the encryption key to an identifier for the local encryption key. The third data block does not include media samples.

[0271] Example 7 includes the subject of any one of Examples 1-6. The first data block includes a first timestamp for a first single discrete media sample of the first media type, and the second data block includes a second timestamp for a first single discrete media sample of the second media type. Example 7 also specifies that during rendering at the client device, the first timestamp and the second timestamp are used to synchronize the rendering of the first single discrete media sample of the first media type and the first single discrete media sample of the second media type.

[0272] Example 8 includes the subject of any one of Examples 1-7. A first single discrete media sample of the first media type is presented at the streaming client without decoding the first single discrete media sample of the first media type.

[0273] Example 9 is a method implemented in a computing system including at least one hardware processor and at least one memory coupled to the at least one hardware processor. The method is used to perform processing operations at a streaming client relative to a block comprising media samples to facilitate low-latency media streaming. The method includes receiving a first data block at a first time. The first data block includes a first sequence identifier for a first media type and a first single discrete media sample of the first media type. The first data block does not include media samples of media types different from the first media type.

[0274] A second data block is received at a second time. The second time is after the first time. The second data block includes a second sequence identifier for the first media type and a second single discrete sample of the first media type. The second data block does not include media samples of media types different from the first media type.

[0275] The second sequence identifier is determined to be lower than the first sequence identifier. The first single discrete media sample and the second single discrete media sample of the first media type are reordered such that the second single discrete media sample of the first media type is provided to the media player in the stream before the first single discrete media sample of the first media type. The first single discrete media sample and the second discrete media sample of the first media type are provided to the media player for presentation at the streaming client. The streaming client receives data blocks of a stream comprising multiple data blocks from multiple sources.

[0276] Example 10 includes the theme of Example 9. Determining that the second serial number is lower than the first serial number includes determining that the first serial number and the second serial number are not consecutive. The method further includes, in response to determining that a first single discrete media sample of the first media type and a second single discrete media sample of the first media type are not consecutive, providing the first single discrete media sample of the first media type and the second single discrete media sample of the first media type for presentation by the client device utilizing the gap between the first media sample of the first media type and the second media sample of the first media type.

[0277] Example 11 includes the subject matter of any of Examples 9 and 10. Determining that the second serial number is lower than the first serial number includes determining that the first serial number and the second serial number are not consecutive. The method further includes, in response to determining that a first single discrete media sample of the first media type and a second single discrete media sample of the first media type are not consecutive, generating a third single discrete media sample of the first media type to be placed in the gap between the first single discrete media sample of the first media type and the second single discrete media sample of the first media type.

[0278] Example 12 includes the subject of any one of Examples 9-11. The first data block includes a first value for a block type identifier. The first value indicates that the first data block includes a sample of the first media type. The second data block includes a first value for the block type identifier.

[0279] Example 13 includes the subject matter of any one of Examples 9-12. In response to a specific request for either the first or the second data block, neither the first nor the second data block is received.

[0280] Example 14 includes the subject of any one of Examples 9-13. The first data block and the second data block do not include codec configuration information that the media player can use to render a first single discrete sample of the first media type or a second single discrete sample of the first media type.

[0281] Example 15 includes the subject matter of any one of Examples 9-14. The first data block is received from a first source of the plurality of sources, and the second data block is received from a second source of the plurality of sources. The second source is different from the first source.

[0282] Example 16 is one or more computer-readable storage media comprising computer-executable instructions capable of performing processing operations at a streaming client relative to a block including media samples to facilitate low-latency media streaming. The one or more computer-readable storage media store computer-executable instructions that, when executed by a computing system including at least one hardware processor and at least one memory coupled to said at least one hardware processor, cause the computing system to perform various operations.

[0283] The operation includes receiving a first data block at a first time. The first data block includes a first sequence identifier for a first media type and a first single discrete media sample of the first media type. The first data block does not include media samples of media types different from the first media type.

[0284] A second data block is received at a second time. The second time is after the first time. The second data block includes a second sequence identifier for the first media type and a second single discrete media sample of the second media type. The second data block does not include media samples of media types different from the first media type.

[0285] The second sequence identifier is determined to be the same as the first sequence identifier. The second data block is discarded, and a second single discrete media sample of the first media type is not provided for presentation at the streaming client. A first single discrete media sample of the first media type is provided to the media player for presentation at the streaming client. The streaming client receives data blocks of a stream comprising multiple data blocks from multiple sources.

[0286] Example 17 includes the subject of Example 16. The first data block includes a first value for a block type identifier. The first value indicates that the first data block includes a sample of the first media type. The second data block includes a first value for the block type identifier.

[0287] Example 18 includes the subject matter of Example 16 or Example 17. In response to a specific request for the first data block and the second data block, the first data block and the second data block are not received.

[0288] Example 19 includes the subject of any one of Examples 16-18. The first data block and the second data block do not include codec configuration information that the media player can use to render a first single discrete sample of the first media type or a second single discrete sample of the first media type.

[0289] Example 20 includes the subject matter of any one of Examples 16-19. The first data block is received from a first source of the plurality of sources, and the second data block is received from a second source of the plurality of sources. The second source is different from the first source.

[0290] Example 21 is a computing system including at least one hardware processor and at least one memory coupled to said at least one hardware processor. The computing system also includes one or more computer-readable storage media storing computer-executable instructions that, when executed by said computing system, cause the computing system to perform operations as a streaming client in peer-to-peer delivery of streaming media content with low latency.

[0291] The operation includes sending a first data block to a first receiving computing system. The first data block includes a first sequence identifier for a first media type and a first single discrete media sample of the first media type. The first data block does not include media samples of media types different from the first media type.

[0292] A second data block is sent to a second receiving computing system. The second data block includes a second sequence identifier for a second media type and a first single discrete sample of the second media type. The second data block does not include media samples of a media type different from the second media type. The second media type is either the first media type or a media type different from the first media type. The second receiving computing system is either the first receiving computing system or a computing system different from the first receiving computing system. The first single discrete media sample of the first media type and the second single discrete media sample of the second media type are presented. The first sequence identifier and the second sequence identifier are used to order the media samples for presentation at a first streaming client. The streaming client receives data blocks for a stream comprising multiple data blocks from multiple sources.

[0293] Example 22 includes the subject matter of Example 21. The first receiving computing system is a second streaming client. The second streaming client is the first streaming client, or a streaming client different from the first streaming client.

[0294] Example 23 includes the subject matter of Example 21 or Example 22. Example 23 also specifies that the first data block is sent to at least one other receiving computing system.

[0295] Example 24 includes the subject of any one of Examples 21-23. The second media type is a media type that is different from the first media type.

[0296] Example 25 includes the subject of any one of Examples 21-23. The second media type is the first media type.

[0297] Example 26 includes the subject of any one of Examples 21-25. The first data block includes a first timestamp for a first single discrete media sample of the first media type, and the second data block includes a second timestamp for a first discrete media sample of the second media type. The first and second timestamps can be used to synchronize the presentation of media samples of different media types.

[0298] Example 27 includes the subject matter of any one of Examples 21-26. The computing system receives the first data block from the provider's computing system.

[0299] Example 28 includes the topics covered in Example 27. The provider's computing system is also a streaming client.

[0300] Example 29 includes the subject matter of any one of Examples 21-28. The second receiving computing system is the first receiving computing system.

[0301] Example 30 includes the subject matter of any one of Examples 21-28. The second receiving computing system is a computing system different from the first receiving computing system.

[0302] Example 31 includes the subject of any one of Examples 21-30. A first single discrete media sample of the first media type is a keyframe, and transmission is performed after determining that the first single discrete media sample of the first media type is a keyframe.

[0303] Example 32 includes the subject matter of any one of Examples 21-31. A first single discrete media sample of the first media type can be rendered at the first streaming client without decoding the first single discrete media sample of the first media type.

[0304] Example 33 includes the subject of any one of Examples 21-32. In response to a specific request for the first data block and the second data block, the first data block and the second data block are not sent.

[0305] Example 34 includes the subject of any one of Examples 21-33. The first data block and the second data block do not include codec configuration information that can be used by the media player of the first streaming client to render a first single discrete sample of the first media type or a first single discrete sample of the second media type.

[0306] Example 35 includes the subject matter of any one of Examples 21-24 or 26-34. The first data block includes a first value for a block type identifier, the first value indicating that the first data block includes a sample of the first media type. The second data block includes a second value for the block type identifier. The second value indicates that the second data block includes a sample of the second media type. The second media type is different from the first media type. The second value is different from the first value.

[0307] Example 36 includes the subject of any one of Examples 21-23 or 25-34. The second media type is the first media type, and the first data block includes a first value for a block type identifier. The first value indicates that the first data block includes a sample of the first media type. The second data block includes the first value for the block type identifier.

[0308] Example 37 includes the subject matter of any one of Examples 21-36. Example 37 also specifies sending a third data block to the first receiving computing system. The third data block includes a mapping from a GUID for the encryption key to an identifier for a local encryption key. The third data block does not include media samples.

[0309] Example 38 is a method implemented in a computing system including at least one hardware processor and at least one memory coupled to the at least one hardware processor. The method is used to perform processing operations in streaming media content with low latency. The method includes receiving a first data block. The first data block includes a first sequence identifier for a first media type and a first single discrete media sample of the first media type. The first data block does not include media samples of media types different from the first media type.

[0310] A second data block is received. The second data block includes a second sequence identifier for a second media type and a first single discrete sample of the second media type. The second data block does not include media samples of a media type different from the second media type. The second media type is the first media type or a media type different from the first media type. The first single discrete media sample of the first media type and the second discrete media sample of the second media type are provided to the media player for presentation at the client device. The client device receives data blocks for a stream including the first data block and the second data block from multiple sources.

[0311] Example 39 includes the subject of Example 38. The second media type differs from the first media type. Example 39 also specifies that the method includes determining that the consecutive number of discrete media samples of the first media type meets a threshold. Based on determining that the consecutive number of discrete media samples of the first media type meets the threshold, a first discrete sample of the second media type is inserted between consecutive media samples of the first media type in the stream. The first single discrete media sample of the first media type and the second discrete media sample of the second media type are provided to a media player for presentation at the streaming client.

[0312] Example 40 is one or more computer-readable storage media including computer-executable instructions capable of performing processing operations when delivering streaming media content with low latency. The one or more computer-readable storage media include computer-executable instructions that, when executed by a computing system including at least one hardware processor and at least one memory coupled to said at least one hardware processor, cause the computing system to perform various operations.

[0313] The operation includes sending a first data block to a first receiving computing system. The first data block includes a first sequence identifier for a first media type and a first single discrete media sample of the first media type. The first data block does not include media samples of a media type different from the first media type. A second data block is sent to a second receiving computing system. The second data block includes a second sequence identifier for a second media type and a first single discrete sample of the second media type. The second data block does not include media samples of a media type different from the second media type. The second media type is either the first media type or a media type different from the first media type. The second receiving computing system is either the first receiving computing system or a computing system different from the first receiving computing system. The first sequence identifier and the second sequence identifier are used to order the media samples for presentation at a first streaming client. The first streaming client receives data blocks from a stream including first and second data blocks from multiple sources.

[0314] Example 17 - Example Computing System Figure 23 A generalized example of a suitable computing system 2300 in which the described technology can be implemented is depicted. The computing system 2300 is not intended to impose any limitations on its scope of use or functionality, as the technology can be implemented in various general-purpose or special-purpose computing systems.

[0315] refer to Figure 23 The computing system 2300 includes one or more processing units 2310, 2315 and memories 2320, 2325. Figure 23 In this diagram, the basic configuration 2330 is included within the dashed lines. Processing units 2310 and 2315 execute computer-executable instructions. The processing units can be general-purpose central processing units (CPUs), processors in application-specific integrated circuits (ASICs), or any other type of processor. Processing units can also include multiple processors. In a multiprocessor system, multiple processing units execute computer-executable instructions to increase processing power. For example, Figure 23 A central processing unit 2310 and a graphics processing unit or coprocessor 2315 are shown. Physical memories 2320 and 2325 may be volatile memories (e.g., registers, caches, RAM), non-volatile memories (e.g., ROM, EEPROM, flash memory, etc.), or some combination of both accessible to the processing unit. Memories 2320 and 2325 store software 2380 implementing one or more of the technologies described herein in the form of computer-executable instructions suitable for execution by one or more processing units.

[0316] The computing system may have additional features. For example, computing system 2300 includes storage device 2340, one or more input devices 2350, one or more output devices 2360, and one or more communication connections 2370. Interconnect mechanisms (not shown), such as buses, controllers, or networks, interconnect the components of computing system 2300. Typically, operating system software (not shown) provides an operating environment for other software executing in computing system 2300 and coordinates the activities of the components of computing system 2300.

[0317] The physical storage device 2340 may be removable or non-removable and includes a magnetic disk, magnetic tape or tape cartridge, CD-ROM, DVD, or any other medium capable of storing information and accessible within the computing system 2300. The storage device 2340 stores instructions for implementing one or more of the technologies described herein.

[0318] One or more input devices 2350 may be touch input devices, such as a keyboard, mouse, pen or trackball, voice input device, scanning device, or another device that provides input to computing system 2300. For video encoding, one or more input devices 2350 may be a camera, video card, TV tuner card, or similar device that accepts video input in analog or digital form, or read video samples into a CD-ROM or CD-RW in computing system 2300. One or more output devices 2360 may be a monitor, printer, speaker, CD burner, or another device that provides output from computing system 2300.

[0319] One or more communication connections 2370 enable communication with another computing entity via a communication medium. The communication medium conveys information such as computer-executable instructions, audio or video input or output, or other data in a modulated data signal. A modulated data signal is a signal having one or more characteristics set or altered in a manner that encodes information in the signal. By way of example and not limitation, the communication medium can be electrical, optical, RF, or other carrier waves.

[0320] The technology is described in the general context of computer-executable instructions (such as those included in program modules) that can be executed on a computing system on a target real or virtual processor. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., that perform specific tasks or implement specific abstract data types. The functionality of program modules can be combined or split among program modules as needed in various embodiments. The computer-executable instructions used for program modules can execute within a local or distributed computing system.

[0321] The terms “system” and “device” are used interchangeably herein. Unless the context clearly indicates otherwise, the terms do not imply any limitation on the type of computing system or computing device. Generally, a computing system or computing device can be local or distributed and can include any combination of dedicated hardware and / or general-purpose hardware with software that implements the functions described herein.

[0322] For the sake of presentation, detailed descriptions use terms such as "determine" and "use" to describe computer operations in a computing system. These terms are high-level abstractions of operations performed by a computer and should not be confused with actions performed by humans. The actual computer operations corresponding to these terms vary depending on the implementation.

[0323] Example 18 - Environments Supported by Example Cloud Figure 24 The illustration depicts a generalized example of a suitable cloud-supported environment 2400 in which the described embodiments, technologies, and techniques can be implemented. In the example environment 2400, various types of services (e.g., computing services) are provided by a cloud 2410. For example, the cloud 2410 can include a collection of computing devices, which can be centrally or distributed, providing cloud-based services to various types of users and devices connected via a network such as the Internet. The implementation environment 2400 can be used in different ways to perform computing tasks. For example, some tasks (e.g., processing user input and presenting a user interface) can be performed on local computing devices (e.g., connected devices 2430, 2440, 2450), while other tasks (e.g., storage of data to be used in subsequent processing) can be performed in the cloud 2410.

[0324] In example environment 2400, cloud 2410 provides services to connected devices 2430, 2440, and 2450 with various screen capabilities. Connected device 2430 represents a device with a computer screen 2435 (e.g., a medium-sized screen). For example, connected device 2430 can be a personal computer, such as a desktop computer, laptop computer, notebook computer, netbook, etc. Connected device 2440 represents a device with a mobile device screen 2445 (e.g., a small-sized screen). For example, connected device 2440 can be a mobile phone, smartphone, personal digital assistant, tablet computer, etc. Connected device 2450 represents a device with a large screen 2455. For example, connected device 2450 can be a television screen (e.g., a smart TV) or another device connected to a television (e.g., a set-top box or game console). One or more of connected devices 2430, 2440, and 2450 can include touchscreen capabilities. Touchscreens can accept input in various ways. For example, a capacitive touchscreen detects touch input when an object (e.g., a fingertip or stylus) twists or interrupts the current extending across the surface. As another example, a touchscreen can use an optical sensor to detect touch input when the beam from the optical sensor is interrupted. Physical contact with the screen surface is not required for some inputs to be detected by a touchscreen. Devices without screen capabilities can also be used in example environment 2400. For example, cloud 2410 can provide services to one or more computers (e.g., server computers) that do not have a display.

[0325] The service can be provided by the cloud 2410 through the service provider 2420 or through other providers of online services (not shown). For example, the cloud service can be customized for the screen size, display capabilities, and / or touchscreen capabilities of specific connected devices (e.g., connected devices 2430, 2440, 2450).

[0326] In example environment 2400, cloud 2410 utilizes service provider 2420 at least in part to provide the technologies and solutions described herein to various connected devices 2430, 2440, and 2450. For example, service provider 2420 can provide centralized solutions for various cloud-based services. Service provider 2420 can manage service subscriptions for users and / or devices (e.g., for connected devices 2430, 2440, 2450, and / or their respective users).

[0327] Example 19 - Example Implementation Although some of the methods disclosed are described in a specific, sequential order for ease of presentation, it should be understood that this descriptive approach includes rearrangement unless a specific ordering is required by the particular language set forth below. For example, in some cases, the sequentially described operations may be rearranged or performed concurrently. Furthermore, for simplicity, the accompanying drawings may not show the various ways in which the disclosed methods can be combined with other methods.

[0328] Any of the disclosed methods can be implemented as computer-executable instructions or a computer program product stored on one or more computer-readable storage media and executed on a computing device (i.e., any available computing device, including smartphones or other mobile devices that include computing hardware). A computer-readable storage medium is a tangible medium accessible within a computing environment, such as one or more optical media discs like DVDs or CDs, volatile memory (such as DRAM or SRAM), or non-volatile memory (such as flash memory or hard disk drives). As an example and reference. Figure 23 Computer-readable storage media include memories 2320 and 2325 and storage device 2340. The term "computer-readable storage medium" does not include signals and carrier waves. Furthermore, the term "computer-readable storage medium" does not include communication connections, such as 2370.

[0329] Any computer-executable instructions used to implement the disclosed technology and any data created and used during the implementation of the disclosed embodiments may be stored on one or more computer-readable storage media. The computer-executable instructions may be a dedicated software application or part of a software application, accessible or downloaded, for example, via a web browser or other software application, such as a remote computing application. Such software may be executed, for example, on a single local computer (e.g., any suitable commercially available computer) or in a networked environment (e.g., via the Internet, a wide area network, a local area network, a client-server network (such as a cloud computing network), or using one or more networked computers.

[0330] For clarity, only selected aspects of the software-based implementation are described. Other details well-known in the art are omitted. For example, it should be understood that the disclosed techniques are not limited to any particular computer language or program. For instance, the disclosed techniques can be implemented using software written in C++, Java, Perl, or any other suitable programming language. Similarly, the disclosed techniques are not limited to any particular computer or type of hardware. Certain details of suitable computers and hardware are well-known and do not need to be described in detail in this disclosure.

[0331] Furthermore, any software-based implementation (including, for example, computer-executable instructions for causing a computer to perform any of the disclosed methods) can be uploaded, downloaded, or remotely accessed via suitable communication means. Such suitable communication means include, for example, the Internet, the World Wide Web, intranets, software applications, cables (including fiber optic cables), magnetic communications, electromagnetic communications (including RF, microwave, and infrared communications), electronic communications, or other such communication means.

[0332] The disclosed methods, apparatus, and systems should not be construed as limiting in any way. Rather, this disclosure is intended to highlight all novel and non-obvious features and aspects of the various disclosed embodiments, individually and in various combinations and sub-combinations. The disclosed methods, apparatus, and systems are not limited to any particular aspect or feature or combination thereof, nor are the disclosed embodiments required to have any one or more particular advantages or problems.

[0333] The techniques from any example can be combined with the techniques described in any one or more of the other examples. Given the many possible embodiments to which the principles of the disclosed techniques can be applied, it should be recognized that the illustrated embodiments are examples of the disclosed techniques and should not be considered as limiting the scope of the disclosed techniques.

Claims

1. A computing system, comprising: At least one hardware processor; At least one memory coupled to the at least one hardware processor; as well as One or more computer-readable storage media storing computer-executable instructions, which, when executed by the computing system, cause the computing system to perform operations as a streaming client in peer-to-peer delivery of streaming media content with low latency, the operations including: Send a first data block to a first receiving computing system. The first data block includes a first sequence identifier for a first media type and a first single discrete media sample of the first media type. The first data block does not include media samples of media types different from the first media type. A second data block is sent to a second receiving computing system. The second data block includes a second sequence identifier for a second media type and a first single discrete sample of the second media type. The second data block does not include media samples of media types different from the second media type, and the second media type is either the first media type or a media type different from the first media type. The second receiving computing system is either the first receiving computing system or a computing system different from the first receiving computing system. Presenting a first single discrete media sample of the first media type and a second single discrete media sample of the second media type; The first sequence identifier and the second sequence identifier are used to sort media samples for presentation at a first streaming client, and the first streaming client receives data blocks for a stream including the first data blocks and the second data blocks from multiple sources.

2. The computing system according to claim 1, wherein, The first receiving computing system is a second streaming client, which is either the first streaming client or a streaming client different from the first streaming client.

3. The computing system according to claim 1 or claim 2, wherein the operation further includes: The first data block is sent to at least another receiving computing system.

4. The computing system according to any one of claims 1-3, wherein, The second media type is a media type that is different from the first media type.

5. The computing system according to any one of claims 1-3, wherein, The second media type is the first media type.

6. The computing system according to any one of claims 1-5, wherein, The first data block includes a first timestamp for a first single discrete media sample of the first media type, and the second data block includes a second timestamp for a first discrete media sample of the second media type, wherein the first timestamp and the second timestamp can be used to synchronously present media samples of different media types.

7. The computing system according to any one of claims 1-6, wherein, The computing system receives the first data block from the provider's computing system.

8. The computing system according to claim 7, wherein, The provider's computing system is also a streaming client.

9. The computing system according to any one of claims 1-8, wherein, The second receiving computing system is the first receiving computing system.

10. The computing system according to any one of claims 1-8, wherein, The second receiving computing system is a computing system different from the first receiving computing system.

11. The computing system according to any one of claims 1-10, wherein, The first single discrete media sample of the first media type is a keyframe, and the transmission is performed after it has been determined that the first single discrete media sample of the first media type is a keyframe.

12. The computing system according to any one of claims 1-11, wherein, The first single discrete media sample of the first media type can be presented at the first streaming client without decoding the first single discrete media sample of the first media type.

13. The computing system according to any one of claims 1-12, wherein, The first data block and the second data block were not sent in response to a specific request for the first data block and the second data block.

14. The computing system according to any one of claims 1-13, wherein, The first data block and the second data block do not include codec configuration information, which can be used by the media player of the first streaming client to present a first single discrete sample of the first media type or a first single discrete sample of the second media type.

15. The computing system according to any one of claims 1-4 or 6-14, wherein, The first data block includes a first value for a block type identifier, the first value indicating that the first data block includes a sample of the first media type, and the second data block includes a second value for the block type identifier, the second value indicating that the second data block includes a sample of the second media type, wherein the second media type is different from the first media type and the second value is different from the first value.

16. The computing system according to any one of claims 1-3 or 5-14, wherein, The second media type is the first media type, and the first data block includes a first value for a block type identifier, the first value indicating that the first data block includes a sample of the first media type, and the second data block includes the first value for the block type identifier.

17. The computing system according to any one of claims 1-16, wherein the operation further comprises: A third data block is sent to the first receiving computing system. The third data block includes a mapping from a GUID of the encryption key to an identifier of the local encryption key, wherein the third data block does not include media samples.

18. A method implemented in a computing system, the computing system including at least one hardware processor and at least one memory coupled to the at least one hardware processor, for performing processing operations while streaming media content with low latency, the method comprising: Receive a first data block, the first data block including a first sequence identifier for a first media type and a first single discrete media sample of the first media type, wherein the first data block does not include media samples of media types different from the first media type; Receive a second data block, the second data block including a second sequence identifier for a second media type and a first single discrete sample of the second media type, wherein the second data block does not include media samples of media types different from the second media type, and the second media type is the first media type or a media type different from the first media type; and A first single discrete media sample of the first media type and a second discrete media sample of the second media type are provided to a media player for presentation at a client device, wherein the client device receives data blocks for a stream comprising the first data block and the second data block from multiple sources.

19. The method according to claim 18, wherein, The second media type is different from the first media type, and the method further includes: The number of consecutive discrete media samples of the first media type satisfies a threshold. Based on the determination that the consecutive number of discrete media samples of the first media type meets a threshold, a first discrete sample of the second media type is inserted between consecutive media samples of the first media type in the stream; and A first single discrete media sample of the first media type and a second discrete media sample of the second media type are provided to the media player for presentation at the streaming client.

20. One or more computer-readable storage media including computer-executable instructions, said computer-executable instructions being configured to perform processing operations in the delivery of streaming media content with low latency, said one or more computer-readable storage media comprising: When executed by a computing system including at least one hardware processor and at least one memory coupled to said at least one hardware processor, the computing system sends computer-executable instructions to a first receiving computing system to transmit a first data block, the first data block including a first sequence identifier for a first media type and a first single discrete media sample of the first media type, wherein the first data block does not include media samples of media types different from the first media type; and When executed by the computing system, the computing system sends a computer-executable instruction to the second receiving computing system to send a second data block, the second data block including a second sequence identifier for a second media type and a first single discrete sample of the second media type, wherein the second data block does not include media samples of a media type different from the second media type, and the second media type is the first media type or a media type different from the first media type, and the second receiving computing system is the first receiving computing system or a computing system different from the first receiving computing system; The first sequence identifier and the second sequence identifier are used to sort media samples for presentation at a first streaming client, and the first streaming client receives data blocks from a stream including the first data block and the second data block from multiple sources.