Video Streaming
Compressed domain compositing techniques combine multiple video streams into a single composite stream at an edge server, using segment identifiers and static symbol frequencies, addressing computational inefficiencies and enhancing network efficiency for seamless multi-stream display.
Patent Information
- Application Number
- JP2025519844
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-26
- Filing Date
- 2024-10-22
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2044-10-22
AI Technical Summary
Existing video streaming technologies often require multiple decoders to handle multiple video streams simultaneously, leading to computational inefficiencies and increased processing requirements, especially when combining different types of content like live streams and third-party content.
The method involves compressed domain compositing techniques that combine two or more video streams into a single composite stream at an edge server, using segment identifiers and static symbol frequencies to ensure compatibility, allowing decoding with a single decoder on the user device.
This approach reduces computational requirements at both the edge server and user device, enhances network bandwidth efficiency, and improves caching capacity by delivering customized composite streams without additional processing, enabling seamless display of multiple streams with minimal delay.
Smart Images

Figure 2025537071000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Application No. 18 / 495,405, filed October 26, 2023. The disclosure of the prior application is considered part of the disclosure of this application and is incorporated by reference into the disclosure of this application.
[0002] This specification relates to streaming video content. [Background technology]
[0003] Picture-in-picture (PiP) streaming is a multi-video display mode that allows a user to view two or more video streams presented in the same window on a user device. For example, PiP can be used to present two or more different video streams with different content in the same display window, allowing a user to experience multiple streams simultaneously. In some cases, a user device may include only a single decoder so that only one video stream at a time can be decoded for presentation on the user device. Summary of the Invention
[0004] The subject matter of this specification generally relates to compressed domain compositing of video streams.
[0005] More specifically, the subject matter herein relates to using compressed domain compositing techniques to provide a user with a fully personalized composite video stream that includes two or more video streams. The composite video stream may include at least two separate streams that are combined in the compressed domain of an edge server and provided to an end-user device, where the composite video stream is decodable by the end-user device using a single decoder. The composite stream may be a PiP or mosaic video stream that includes a primary video stream (e.g., third-party content video) and a secondary video stream (e.g., live-stream video content), which acts as a 50 / 50 split screen to the end-user device.
[0006] In general, one innovative aspect of the subject matter described herein can be embodied in a method including, by a local area server, acquiring compressed video streams, each compressed video stream including a plurality of frames, each frame including a video content region that is a portion (e.g., up to 50%) of the frame. Each frame of the compressed video stream is encoded using a segment identifier of pixels included in the video content region, and the frames of the compressed video stream are encoded using a set of static symbol frequencies. To acquire the compressed-domain composite video stream, the local area server receives a request for video stream content from a user device and, in response to the request from the user device, composites a first compressed video stream and a second compressed video stream of the compressed video streams. Each frame of the compressed-domain composite video stream includes a first video content region of the first compressed video stream and a second video content region of the second compressed video stream. Each frame of the compressed-domain composite video stream includes a respective first segment identifier for pixels corresponding to the first video content region and a second segment identifier for pixels corresponding to the second video content region. The first video content region and the second video content region occupy non-overlapping portions of the frame. The local area server provides the user device with packets containing sets of frames of the compressed domain composite video stream decodable by a single decoder.
[0007] Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.
[0008] The above and other embodiments may each include one or more of the following features, alone or in combination. Specifically, one embodiment includes all of the following features in combination. In some implementations, obtaining the compressed video stream includes receiving the video stream by a server; and generating each of the compressed video streams, the generating including, for each frame of a plurality of frames of the video stream, rendering a new frame including a video content region that includes a portion of the frame and corresponds to content of the frame; for the video content region, defining segment identifiers for pixels included in the video content region; and encoding the frame using a set of static symbol frequencies.
[0009] In some implementations, the rendering further includes rendering areas not included in the video content areas as null content.
[0010] In some implementations, generating the compressed video streams includes generating a primary video stream including context-responsive video content and generating at least one secondary video stream including live stream content, where the first compressed video stream is the primary video stream and the second compressed video stream is the secondary video stream. Generating the primary video streams may include pre-conditioning and storing the context-responsive video in a database for inclusion in the compressed domain composite video stream, where an appropriate subset of the context-responsive video content is pre-cached on a local area server.
[0011] In some implementations, generating the compressed domain composite video stream includes selecting, by a local area server, primary content videos from a database of primary video streams for inclusion in the compressed domain composite video stream in response to user-based selection criteria of the user device. In response to different user requests, the local area server selects different primary content videos for inclusion in the compressed domain composite video stream.
[0012] In some embodiments, generating the plurality of compressed video streams includes generating a video stream including at least one of context-responsive video content and live stream content, and the first compressed video stream and the second compressed video stream are generated from the video stream.
[0013] In some implementations, encoding the frame using a set of static symbol frequencies includes encoding at a default frequency value for each symbol probability of the frame.
[0014] In some implementations, defining, for a video content region, a segment identifier further includes assigning a first quantization level to segment identifiers of pixels included in the video content region.
[0015] In some implementations, the rendering further includes rendering a null region including a width within the video content region of the compressed video stream along an edge of the video content region, the null region of the video content region being located adjacent to a replacement region configured to be replaced with another compressed video stream in the composite video stream.
[0016] In some embodiments, generating the compressed domain composite video stream from the first compressed stream and the second compressed stream includes constructing an uncompressed header for each frame of the compressed domain composite video stream based on respective segment identifiers of the first compressed stream and the second compressed stream.
[0017] In some implementations, the frame comprises tiles, the video content region comprises one or more contiguous tiles of a tile-based composition of the frame, and the pixels comprised in the video content region comprise one or more contiguous tiles, and the one or more contiguous tiles comprising the video content region are half of a tile of the tile-based composition of the frame.
[0018] In some implementations, the method further includes obtaining, by a local area server, a second set of compressed video streams, each of the second set of compressed video streams including a plurality of frames. The local area server receives a second request for the video stream content from the user device and, in response to the second request, provides to the user device a second packet including a set of frames for one of the second set of compressed video streams.
[0019] The subject matter described herein can be implemented in certain embodiments to achieve one or more of the following advantages: An end device can receive, decode, and display a bitstream-level composite stream as a single stream, even though the composite stream may be from two completely unrelated video streams. As a result, the end device does not experience any additional delay or increased computational requirements to display a composite stream that includes two separate streams. The solution described herein involves back-end preconditioning, but is computationally lightweight in terms of real-time processing at the cache edge server and user device. The described solution can increase network bandwidth efficiency for providing multiple customized composite streams to different user devices because the server delivers only the live stream and N primary content streams (e.g., those that can be pre-cached at the cache edge server).
[0020] Additionally, the features described herein enable composite streams, allowing an end user device to receive a fully customized composite video stream that includes two or more video streams (e.g., a live stream and a third-party content stream) without substantial additional computational requirements at the end user device or edge server during real-time provisioning of the live stream. In other words, the techniques described herein can substantially improve the caching capacity of video streams at a caching edge server, where the caching edge server only needs to store N+M video streams instead of N×M video streams.
[0021] As described herein, compressed-domain compositing of video generates a single composite stream in the compressed domain from two or more input compressed streams by a caching edge server on a per-user basis, without requiring the edge server or user device to perform arithmetic decoding and / or arithmetic encoding of the compressed streams (e.g., performing concatenation of video metadata with the video inputs). This would otherwise be prohibitively expensive and computationally intensive to provide a custom stream for each user. The decoder of the user device effectively "sees" a single video. As a result, no additional processing requirements are required on the decoder to decode additional frames (e.g., hidden frames, alt-ref frames, etc.). By placing the majority of the processing in a preconditioning step before providing the compressed video to the caching edge server, reduced processing requirements are met by the caching edge server and end-user device.
[0022] The details of one or more embodiments of the subject matter herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. [Brief explanation of the drawings]
[0023] [Figure 1] FIG. 1 is a block diagram of an example environment for compressed domain composite video streaming. [Figure 2] FIG. 1 is a block diagram of an example environment for compressed domain composite video streaming. [Figure 3A] 1 shows a schematic example of tile-based compositing. [Figure 3B] 1 shows a schematic example of tile-based compositing. [Figure 4] FIG. 1 is a schematic diagram of an exemplary compressed domain synthesis process. [Figure 5A] 1 is a schematic diagram of an example of a VP9-based frame structure. [Figure 5B] 1 shows an example of a corrupted video stream. [Figure 6] 1 illustrates an example of a composite video frame containing multiple tiles. [Figure 7] 1 shows a schematic diagram of an example of reference frames between a primary stream, a secondary stream, and their composite video stream; [Figure 8] 1 shows a schematic diagram of an example of a gap inserted between two video streams of a composite video stream; [Figure 9A] An example of a motion vector containment solution is shown. [Figure 9B] An example of a motion vector containment solution is shown. [Figure 9C] An example of a motion vector containment solution is shown. [Figure 10A] 10 shows another example of a motion vector containment solution. [Figure 10B] 10 shows another example of a motion vector containment solution. [Figure 11] 1 is a flowchart of an exemplary process for compressed domain composite video streaming. [Figure 12] FIG. 1 is a block diagram of an exemplary computer. DETAILED DESCRIPTION OF THE INVENTION
[0024] Like reference symbols and designations in the various drawings indicate like elements.
[0025] As used herein, the terms "compressed-domain composited video stream" and "compressed-domain composite video processing" refer to the bitstream-level compositing of two or more encoded (i.e., compressed) video streams to form an output composited video stream for presentation incorporating all or a portion of the input video streams. In other words, a picture-in-picture or mosaic of two or more different video streams is presented in the frames of the composited video stream. Generating the composited video stream is done in the compressed domain, i.e., without decoding the composited video to generate each frame of the composited video and then re-encoding the frames of the composited video. Instead, as described in more detail below, frames of the compressed-domain composited video are generated from two or more compressed videos without having to decode / re-encode the combining videos.
[0026] As used herein, the term "preconditioning" refers to receiving the output of a general-purpose video encoder (e.g., a hardware encoder) and adapting the output to enable lightweight compositing of said content (e.g., by copying appropriate video metadata into appropriately constructed output frames). Several examples of codecs for encoding / compressing and decoding / decompressing video content are provided herein, but for clarity, the VP9 codec is used as the primary example of the process described.
[0027] 1 is a block diagram of an exemplary environment 100 for compressed domain composite video streaming. The exemplary environment 100 includes a network 102, such as a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof. The network 102 connects electronic document servers 104 ("electronic document servers"), user devices 106, and a digital component distribution system 110 (also referred to as DCDS 110). The exemplary environment 100 may include many different electronic document servers 104 and user devices 106.
[0028] A user device 106 is an electronic device capable of requesting and receiving resources (e.g., electronic documents) over the network 102. Exemplary user devices 106 include personal computers, mobile communication devices, and other devices capable of sending and receiving data over the network 102. A user device 106 typically includes a user application, such as a web browser, to facilitate sending and receiving data over the network 102, although native applications executed by the user device 106 can also facilitate sending and receiving data over the network 102.
[0029] The one or more third parties 130 include content providers, product designers, product manufacturers, and other parties involved in the design, development, marketing, or delivery of videos, products, and / or services.
[0030] An electronic document is data that presents a set of content on a user device 106. Examples of electronic documents include web pages, word processing documents, portable document format (PDF) documents, images, videos, search result pages, and feed sources. Native applications (e.g., “apps”), such as applications installed on a mobile, tablet, or desktop computing device, are also examples of electronic documents. Electronic documents 105 (“electronic documents”) can be provided to a user device 106 by an electronic document server 104. For example, the application server 104 can include a server that hosts a publisher's website. In this example, the user device 106 can initiate a request for a web page of a given publisher, and the electronic document server 104 that hosts the web page of the given publisher can respond to the request by sending machine hypertext markup language (HTML) code that initiates the presentation of the given web page on the user device 106.
[0031] The request 108 may include data specifying characteristics of an electronic document (e.g., video content) and a location where the digital content can be presented. For example, data specifying a reference to the electronic document (e.g., a composite video stream) in which the digital content will be presented, an available location of the electronic document available for presenting the digital content (e.g., a digital content area within a frame of a composite video), the size of the available location, and / or the position of the available location within a presentation of the electronic document may be provided to the DCDS 110. Similarly, keywords specified for selection of the electronic document ("document keywords") referenced by the electronic document or data specifying entities (e.g., people, places, or things) referenced by the electronic document may also be included in the request 108 (e.g., as payload data) and provided to the DCDS 110 to facilitate identification of digital content items eligible for presentation with the electronic document.
[0032] The request 108 may also include data related to other information, such as user-provided information, geographic information indicating the state or region in which the request is submitted, or other information that provides context about the environment in which the digital content is displayed (e.g., the type of device on which the digital content is displayed, such as a mobile device or tablet device). The user-provided information may include demographic data about the user of the user device 106. For example, demographic information may include age, gender, geographic location, education level, marital status, household income, occupation, hobbies, social media data, and whether the user owns a particular item, among other characteristics.
[0033] In situations where the systems described herein may collect or utilize personal information about a user, the user may be provided with the opportunity to control whether a program or feature collects personal information (e.g., information about the user's social network, social behavior or activities, occupation, user preferences, or the user's current geographic location) or whether and / or how to receive content from a content server that may be more relevant to the user. Additionally, certain data may be anonymized in one or more ways so that personally identifiable information is removed before it is stored or used. For example, a user's identifying information may be anonymized so that personally identifiable information about the user cannot be determined, or if location information is obtained, the user's geographic location may be generalized (such as to the city, zip code, or state level) so that the user's specific location cannot be determined. Thus, users may control how information about them is collected and used by content servers.
[0034] Data specifying characteristics of the user device 106 may also be provided in the request 108, such as information identifying the model of the user device 106, the configuration of the user device 106, or the size (e.g., physical size or resolution) of the electronic display (e.g., touchscreen or desktop monitor) on which the electronic document will be presented. The request 108 may be transmitted, for example, over a packetized network, and the request 108 itself may be formatted as packetized data having a header and payload data. The header may specify the destination of the packet, and the payload data may include any of the information described above.
[0035] In response to receiving request 108 and / or using information contained in request 108, DCDS 110 selects digital content to be presented with a given electronic document. For example, DCDS 110 selects a digital content item from digital component database 112, such as one of a repository of available pre-tuned videos, for generating a composite video stream with at least one other pre-tuned video.
[0036] In some embodiments, the DCDS 110 is implemented in a distributed computing system (or environment), which may include, for example, a server (e.g., server 200 of FIG. 2 ) and a set of computing devices (e.g., local area servers 202 of FIG. 2 ), interconnected to identify and deliver digital content in response to requests 108. For example, as described in further detail with reference to FIG. 2 , a cache edge server may receive requests to deliver digital content and generate a composite video stream including electronic document(s) for delivery to an end-user device. The set of computing devices works together to identify a set of digital content eligible for presentation with the electronic documents as a composite video stream from a corpus of millions or more of available digital content (e.g., pre-conditioned video content that can be used to generate a compressed-domain composite video stream). The millions or more of available digital content may be indexed, for example, in the digital component database 112. Each digital content index entry may reference the corresponding digital content and / or include delivery parameters (e.g., selection criteria) that condition the delivery of the corresponding digital content.
[0037] In some implementations, digital components from the digital component database 112 may include content provided by a third party 130. For example, the DCDS 110 can present video content stored in the digital component database 112. As described in further detail below, digital components (e.g., videos) stored in the digital component database 112 can be pre-conditioned and stored in the database 112 in anticipation of being included in a composite video stream for presentation at a user device. Additionally, the digital components can be real-time (e.g., live stream video), which can also be pre-conditioned by the DCDS 110 for inclusion of the live stream video in a composite video stream for presentation at a user device. As used herein, the term “primary stream” generally refers to one or more user-directed, pre-conditioned, context-responsive content stored in the digital component database 112 for later incorporation by the DCDS 110 into a customized composite video stream. As used herein, the term “secondary stream” refers to pre-conditioned live stream content that can be incorporated into a customized composite video stream by the DCDS 110. Although described herein as a composite video stream including a primary stream and a secondary stream, more than one video stream can be used to generate the composite video stream. For example, a mosaic composite video stream can include three or more video streams, and each video stream can be a primary stream or a secondary stream type of video content.
[0038] The identification of eligible digital content can be segmented into tasks, and the tasks can then be allocated among the computing devices in the set of computing devices. For example, different computing devices can each analyze different portions of the digital component database 112 to identify different digital components having distribution parameters that match the information included in the request 108.
[0039] The DCDS 110 aggregates the results received from the set of multiple computing devices and uses information associated with the aggregated results to select one or more instances of digital content to provide in response to the request 108. The DCDS 110 can then generate and transmit, over the network 102, response data 114 (e.g., digital data representing the response) that enables the user device 106 to integrate the selected set of digital content into a given electronic document. As a result, the selected set of digital content and the content of the electronic document are presented together on the display of the user device 106.
[0040] As described above, the DCDS 110 can be implemented in a distributed computing system (or environment) that includes, for example, a server (e.g., server 200) and a set of multiple computing local area servers 202 (e.g., cache edge devices) that are interconnected and that identify and deliver digital content in response to requests 108.
[0041] 2 is a block diagram of an example environment 201 for compressed-domain composite video streaming. Server 200 can generate, process, and / or store video content. The video content can be pre-processed by server 200 to generate encoded (e.g., compressed) video content. For example, server 200 can encode context-responsive video content and store the encoded video content in a repository (e.g., digital component database 112) for later incorporation into a customized composite video stream. In another example, server 200 can encode live stream video content in preparation for publishing the live stream video content to user devices.
[0042] Although described here as actions performed by server 200, in some implementations, some or all of the encoding of the video content (e.g., context-responsive video content and / or live stream video content) may be pre-arranged (e.g., encoded) by one or more third parties. For example, a third-party content publisher (e.g., third party 130) may pre-arrange the third-party-generated video content and provide the pre-arranged video content to server 200, e.g., over network 102.
[0043] Server 200 preconditions (i.e., preprocesses) the video content by encoding it using the VP9 codec, which is described in more detail below. For simplicity, the encoding of the video content is described with reference to the VP9 codec, but other codecs may be used, for example, using H.264, H.265, AV1, or other codecs configured to provide sufficient features to enable block replacement (e.g., tile replacement), as described below.
[0044] In (1), server 200 can receive, process, and store primary video content 205 (e.g., third-party digital content), which can be accessible for future use in generating user-customized composite video content. As described above, the primary video stream can include content customized for an end user device (e.g., selected to contextually respond to one or more user attributes), which can be specific to each end user device receiving the composite (e.g., PiP) video stream from DCDS 110. For example, the primary video stream can be selected based on contextual information, including one or more attributes, such as the user's demographics, geography, viewing history, or other user-based preferences or information. Server 200 can pre-generate and pre-tune a repository of primary content stream videos, which can include different content that can be selected for inclusion in the composite stream at a later time based in part on specific criteria (e.g., user preferences, contextual information obtained from or associated with the receiving user's device, etc.).
[0045] At (2), the server can receive and process secondary video content 207, e.g., live stream (e.g., real-time broadcast) video content that can be accessible for use in generating customized composite video content. The server 200 can expose at least two versions of the secondary content, including (A) the secondary stream alone, and (B) a pre-processed version of the secondary stream that can be combined with the primary stream(s) to generate a composite stream.
[0046] In some implementations, the server may pre-cache a repository of primary stream digital content at a local area server (e.g., local area servers 202a, 202b, 202c) selected based, for example, on the set of user devices served by the local area server, the geographic location of the local area server, etc. For example, server 200 may periodically pre-cache a set of primary stream digital content at local area servers 202a-c to maintain an updated repository of pre-tuned video content at the local area servers that may be combined by local area servers 202a-c to generate a composite video stream for presentation at user device 204. In some implementations, a different set of primary stream digital content may be cached at each local area server based, for example, in part, on the geographic location of the local area server and the geographic locations of the user devices served by the local area server.
[0047] At (4), the local area server 202a receives a request for content from a user device, e.g., user device 204a. The request for content may include a request to display a secondary stream, i.e., live stream video content (A). The local area server 202a can identify the possibility of providing the secondary stream within a composite stream, e.g., as a PiP within a primary stream containing third-party content. The selection process is described with reference to FIG. 1, but generally, the selection can be based on user information for the user device.
[0048] In some implementations, an auction can be conducted (e.g., at an auction server) for third-party content from a repository of pre-tuned primary stream videos, and the winning customized pre-tuned primary stream from the repository of third-party content can be returned to the edge server.
[0049] The local area server 202a performs compressed domain compositing of the selected primary and secondary streams to generate a composite video. The local area server 202a can provide the composite stream including the primary stream (B) and the secondary stream (A) to the end user device. At (5), the end user device receives and decodes the composite stream and presents the composite stream including the live stream (A) and the third-party content (B) on the user device's display within the user device's viewport.
[0050] The local area server 202a can generate two or more different composite streams to provide to each end user device. For example, in (6), the local area server 202a provides a different composite stream to the second user device 204b, the different composite stream including the same secondary content (A) accompanied by different primary content (C).
[0051] In some implementations, the local area server may perform this process continuously, continuously requesting and providing multiple (e.g., two or more) primary content streams with the encoded composite stream that includes the secondary streams. In some implementations, the local area server 202a may provide multiple (e.g., two or more) continuous secondary streams along with the encoded composite stream.
[0052] In some implementations, the request (e.g., request 108) may include an opt-out for displaying the primary content (e.g., non-live stream video content), whereby the local area server (e.g., local area server 202b) provides only the secondary stream (A) (e.g., the live stream) for presentation to the user device(s) 204c at (7). In some implementations, as described in further detail below with reference to FIG. 4, the request (e.g., request 108) may include a change between a request to display a composite stream and a request to display a live stream.
[0053] 2, the local area server 202a performs compressed-domain composition of selected primary and secondary streams to generate a composite video in response to a request for content from a user device. Thus, the local area server 202a can generate a different customized composite video for each device (or cluster of devices) with minimal decoding effects downstream of the user device. In effect, each user device 204 receives a single customized composite video from the local area server 202a that can be decoded by a single decoder.
[0054] In some cases, different end user devices can all receive the composite stream from the local area server 202a. In some cases, different end user devices can each receive either (i) the composite stream or (ii) only a secondary stream from the local area server 202a. In some cases, at least one end user device receives a composite stream that has a different primary stream than at least one other end user device.
[0055] Each frame can be divided into subportions, e.g., containing one or more tiles (e.g., in the VP9 codec) or one or more slices (e.g., in the H.264 codec), each of which is encoded independently of the others. Within the VP9 codec, a tile permutation process can be used to generate a composite-encoded stream using compressed-domain compositing techniques, e.g., at a local edge server. Within the VP9 framework, each tile is arithmetically encoded and decoded independently, and tiles can be encoded and / or decoded in parallel. Depending in part on the constraints of the VP9 specification, different numbers of tiles can be used. For example, a 1080p stream might contain four tiles, a 4k stream might contain eight tiles, a 480p stream might contain two tiles, etc.
[0056] For example, as shown in Figures 3A and 3B, each frame of the primary stream can be divided into multiple sub-portions, such as tiles 1, 2, 3, and 4 in Figure 3A, or sub-portions 5, 6, 7, 8, 9, 10, 11, and 12 in Figure 3B. Each sub-portion can be decodable independently of the other sub-portions, and the system 110 (e.g., local area server 202a) can replace one or more of the sub-portions with tile portions of the secondary stream during the compressed-domain composition process to generate the composite video.
[0057] As further described with reference to FIG. 4 , preconditioning of each video stream involves rendering each frame of the adjusted video stream to include video content areas and “null” or black areas, where the null or black portions of the preprocessed video stream can be replaced with video content areas from the other video stream. In such cases, the primary and secondary streams are each encoded such that tiles are extracted from each stream and combined to form a composite frame. For example, each video stream may be encoded as half-width content in a first portion of the frame, preconditioned as described herein, and a separate, second portion of the frame may be compressed domain (CD) combined at the encoder to generate full-width content. In other words, preconditioning involves encoding only tiles that contain content (e.g., non-null tiles). In another example, as shown in FIG. 3A, during pre-adjustment, frames of a video stream included as a primary content stream in a composite video stream can be rendered such that the video content area is within tiles 3 and 4, and frames of a video stream included as a secondary content stream in the composite video stream can be rendered such that the video content area of the secondary content stream is within tiles 1 and 2.
[0058] In some implementations, instead of (or in addition to) a "null" or black area, preconditioning may include rendering a border that encompasses the video content area. For example, the border may be a technical feature (e.g., a texture, a pattern, a graphic, etc.).
[0059] In some implementations, the primary content stream is preprocessed and stored in a digital component database to condition the content stream so that it can be later included in a composite stream for presentation on a user device. The condition includes rendering, for each frame of the primary content stream, a new frame in which the original content occupies no more than 50% of the frame (e.g., no more than 50% of the frame, including edge areas shared with the replaced content), with the remainder of the frame being rendered as "null" or black content. In other words, the original content can be rendered within the video content area of the new frame that forms the preprocessed primary content stream. The video content area of the primary content stream can be selected to have the same orientation relative to the frame, e.g., the first (left) portion of the frame. The secondary stream (e.g., live stream content) can be preprocessed to render the content of each frame of the secondary stream as no more than 50% of the new frame, with the remainder of the new frame being rendered as "null" or black content. The original content can be rendered within the video content area of the new frame that forms the preprocessed secondary content stream. The video content area of the secondary content stream can be selected to have the same orientation relative to the frame, for example, the second (right-hand) portion of the frame.
[0060] In some implementations, a local area server (e.g., local area server 202a) can switch between providing an encoded PiP composite stream, a live stream (e.g., only a secondary stream), and third-party content (e.g., only a primary stream) to a user device.
[0061] 4 is a schematic diagram of an exemplary compressed domain composition process. A local area server 402 (e.g., local area server 202a) publishes video streams (1A) (e.g., live stream content), (1B) (e.g., live stream content pre-conditioned for inclusion in a composite stream), and (2) (e.g., content pre-conditioned for inclusion in a composite stream). The local area server 402 can select whether to provide (i) video stream (1A) including only the live stream video content, or (ii) a composite video stream including video streams (1B) and (2). For example, the local area server 402 can select whether to provide video stream (1A) or video stream (1B) to a user device based in part on a request from a user device including user device information.
[0062] If only the live stream (1A) is provided to the user, the stream can be encoded according to standard processes (e.g., as if there were no PiP stream). In other words, this is done without any additional preconditioning steps used to prepare the video streams for inclusion in the composite video stream. The live stream (1A) is provided to the user device as a single video stream (3) that can be decoded by a single decoder on the user device.
[0063] When the live stream is provided to the user device in the composite video stream, the stream is conditioned (1B) to replace it as PiP with the preconditioned video stream (2). The local area server 402 generates and provides the composite video stream as a single video stream (3) from the video stream (1B) and the video stream (2) that can be decoded by a single decoder of the user device. The preconditioning described herein enables the local area server 402 to provide a unique composite video stream for each user device.
[0064] As previously mentioned, video content streams are pre-conditioned for inclusion in a composite video stream. Figure 5A is a schematic diagram of an example VP9-based frame structure, which serves as a high-level overview for the description below. Within the VP9 framework, each sub-portion (e.g., "tile") of a frame is arithmetically encoded separately.
[0065] For example, as shown schematically in FIG. 5A, the frame structure includes an uncompressed header containing information about how to decode the full frame (e.g., in the examples of VP9 and AV1), such that if the frame is a composite frame containing content from two or more different video streams, the uncompressed header contains information for decoding all of the composite video streams. For example, the uncompressed header of a frame of a composite stream containing a primary stream and a secondary stream contains information for decoding the primary stream portion and the secondary stream portion within the composite stream. Thus, the information in the uncompressed header of pre-conditioned video streams (before combining the compressed domains) is similarly structured; e.g., at least some aspects of the uncompressed header match across all pre-conditioned streams. Otherwise, structural differences may result in corrupted video files, as shown in FIG. 5B.
[0066] Thus, preconditioning a video stream using, for example, VP9 or AV1, generally involves aligning the arithmetic coding of the composite video stream to reduce the likelihood of corruption of the resulting composite stream. One modification to the encoding process of a video stream for inclusion in the composite video stream is to align the arithmetic coding used to represent the symbol frequencies of the symbols used to describe how to reconstruct the output at the decoder. In arithmetic decoding, probabilities are specified at the frame level, and typically the symbol frequency of one frame updates the probability of the next frame. Thus, subsequent frames can be transmitted using fewer bits.
[0067] When pre-conditioning video streams for inclusion in a composite video, a portion of each frame of the pre-conditioning video is rendered as "null" or black. As a result, when encoding the pre-conditioning video stream, a portion (e.g., half) of the frame will be encoded as "null" or black. If standard encoding techniques are used, the resulting composite video stream (replacing the null or black tiles with an alternative PiP video stream) will cause a decoder to see different symbol frequencies for the current frame and, consequently, use incorrect probabilities when decoding the next frame of the composite video stream. In other words, because the symbols describing the frames of each video stream forming the composite video stream are individually compressed using arithmetic coding to facilitate transmission to end devices using fewer bits, the symbol frequencies used to update the probabilities between successive frames of the video streams should match (e.g., perfectly match) during the encoding of the streams to ensure consistency between composite video streams (e.g., between the primary and secondary streams forming the composite video stream) so that the composite video can be decoded properly. This can be accomplished, for example, using error-resilient encoding or by explicitly signaling the probabilities.
[0068] In some implementations, ensuring consistency during decoding of a composite video stream includes using a static (e.g., default) set of known values in a standard to encode frames of pre-conditioned video streams. For example, within the VP9 codec, the default over_under.webmprobabilities feature can be used when encoding each frame in error-resilient mode. By using this feature (typically used for lossy, "error-prone" communication links) in a novel way, using the default probabilities for each frame (rather than updating them) can ensure that symbol probabilities match when a composite video stream containing two or more different video streams is decoded.
[0069] In some implementations, for example, if reduced bandwidth at the end user device justifies the increased processing load on the local area server, ensuring match during decoding of the composite video stream includes combining (e.g., concatenating) the symbol frequency probabilities of each composite frame to include a probability for each of the symbol frequencies of each of the video streams included in the composite video stream. For example, ensuring match during decoding may include undoing the arithmetic encoding of the two video streams and then re-arithmetically encoding the two videos together (i.e., updating the probabilities for each frame). While this approach may require significantly more processing at the local area server, it may still be significantly less expensive than using traditional methods that require full encoding of the video streams.
[0070] Another modification to the encoding of video streams for inclusion in the composite video stream is to utilize codec-defined frame segmentation map functionality to ensure consistency between the various qualities and compression levels of the encoded video streams included in the encoded composite video stream. The segmentation map functionality can be implemented, for example, using segment identifiers (IDs) to enable each of the video streams included in the composite video stream to utilize an independent encoding quality / compression level appropriate to the complexity of the content included in the video stream. In such a case, the decoding process then functions to present a decoded composite video with the characteristics of each of the constituent video streams. For example, a segmentation map within the VP9 framework can be applied to a frame to divide the frame into 8x8 pixel block regions corresponding to the different included video streams, and different segment identifiers (IDs) can be assigned to the pixel block regions. The uncompressed header incorporates, for each segment ID used, the values of one or more properties that adjust the decoding of that region in the composite video. In other words, different segment IDs can be used to enforce uncompressed headers that enable matching values of one or more properties for the decoding of the compressed video. For example, characteristics of the segmentation map can be used to ensure matching of one or more of the quantization level (Q), loop filter strength, reference frame, and skip mode between the source and synthesized video streams.
[0071] 6 shows an example of a composite video frame including multiple tiles. As shown, frame 600 is divided into four tiles A, B, C, and D. Tiles A and B contain video content areas of a first video stream (e.g., a secondary stream), and tiles C and D contain video content areas of a second video stream (e.g., a primary stream). The first video stream contained in tiles A and B uses a first segment ID for all pixels in tiles A and B during encoding of the first video, and the second video uses a second segment ID for all pixels in tiles C and D during encoding of the second video.
[0072] In some implementations, separate segment IDs can be used for different tiles of a frame to ensure consistent quantization levels (Q) between each of the video streams included in the composite video stream. In other words, each composite video in the composite video stream can be assigned an independent quantization level. Generally, encoding utilizes some constraints on bitrate and quality to reduce the number of transmission bits for a network or device, or to reduce the number of transmission bits without significantly improving quality. For example, high-complexity content results in lower quality (higher quantization level), while simpler content (e.g., up to a certain point) results in higher quality (lower quantization level). Using a single quantization level for all content can result in either too many bits for complex content or poor quality for simple content.
[0073] Briefly, a Discrete Cosine Transform (DCT) is implemented by the server to identify important or salient aspects in the structure of the image captured by the frame, to which a quantization factor can be applied to produce a lossy compression of the frame. The data at the DCT output can be quantized, resulting in lower quantization levels resulting in higher quality (more faithful to the original image) and higher quantization levels resulting in lower quality (less faithful to the original image).
[0074] In some implementations, adjusting the quantization level can be used to ensure rate control that the encoding (e.g., compression) of video content is within a target bit range, so that quality variations between different video content streams are less noticeable to a viewing user. An encoder can apply a distinct Q to the discrete cosine transform (DCT) of tiles of a frame that contain video content regions of a video stream from the DCT of tiles that contain "null" or black areas of the frame. For example, a lower quantization level can be used for tiles that contain video content regions, and a higher quantization level can be used for tiles that contain null or black areas.
[0075] Typically, Q can be chosen to be constant across frames. However, to ensure consistency of Q across composite video stream frames, an encoder can utilize a frame segmentation map to use a distinct segment ID for each of the frame's tiles during the encoding process. This allows the uncompressed header of a frame in the composite video stream to allow for distinct Qs for different regions identified by segment IDs. For example, when a local area server performs compressed-domain compositing of different video streams and generates a composite video stream by replacing one or more tiles in the primary stream with tiles from the secondary stream, the distinct segment IDs for each of the retained tiles in the primary stream and the inserted tiles in the secondary stream ensure that the Q of each of the tiles in the composite stream matches its original value.
[0076] In some implementations, segment IDs are used to increase pixel resolution of a region of interest (e.g., a feature) with a first segment ID having a first Q for the base content (e.g., high Q) and a second segment ID having a second Q for a smaller region to resolve any quality issues (e.g., low Q). In some examples, segment IDs can be used to boost blocks of pixels that have high visibility to the user (e.g., an object or region in the video content that appears throughout many frames of playback).
[0077] In some implementations, a separate segment ID can be used for each composite video in the composite video stream. For example, two segment IDs can be used for each composite video, where a first segment ID is used for most of the content from the composite video and uses the same Q value as the source video, and a second segment ID has the smallest possible Q value, which is used when motion vector blocks need to be decomposed into smaller blocks (e.g., in the case of hard motion vector correction described in FIG. 10B). In this case, residuals for blocks used to correct any differences between the copied content and the desired content often need to be re-encoded for smaller blocks. If the original Q value were used, large errors could occur, so using a small Q value minimizes these errors.
[0078] 7 shows a schematic diagram of an example 700 of reference frames between a primary stream, a secondary stream, and their composite video stream. Within the VP9 codec, a motion vector within a frame can reference up to three different frames from a set of eight reference buffers. During decoding of the composite video stream, the reference frames used in each of the primary and secondary video streams must match so that the reference frames used in the composite video frame are the same as those in the primary and secondary streams.
[0079] In some implementations, after a frame is decoded and reconstructed for presentation, one or more post-processing steps are performed at the user device. One such post-processing step includes a looping / deblocking filter that can be applied to the reconstructed frame by the user device to smooth edges across adjacent pixel blocks. Edge artifacts between adjacent pixel blocks can occur due to the independent DCT encoding of each block.
[0080] In some implementations, a gap is included between the video streams included in the composite stream to prevent the loop filter from smoothing between the two different streams that make up the composite stream. FIG. 8 shows a schematic diagram of an example gap inserted between two video streams in the composite video stream. The gap 802 located between the primary stream 804 and the secondary stream 806 is at least a threshold number of pixels wide to prevent smoothing / blending, e.g., separated by at least 16 pixels from the edge of the video content area of the primary stream to the edge of the video content area of the secondary stream. In some cases, portions of the gap 802 between the primary stream 804 and the secondary stream 806 are inserted into each of the primary stream 804 and the secondary stream 806 during pre-conditioning of the primary stream 804 and the secondary stream 806, respectively. For example, 8 pixels are added to the edge of the video content area of one video stream adjacent to the other video content area of the other video stream when the two video streams are included in the composite video stream.
[0081] Generally, a significant portion of the compression of video content during the encoding process is the result of using motion vectors to track actual changes between successive frames, and using residuals (e.g., represented by DCT) to encode the difference between the actual and predicted frames. Motion vectors can be represented by different numbers of bits and can have high or low precision, but they need to be consistent across frames to ensure that the reconstructed composite video stream reflects the content of the original primary and secondary streams. In some implementations, the motion vector precision can be set to a default value during the encoding process of the video streams included in the composite video stream, e.g., a value of 1.
[0082] In some implementations, due to limitations of hardware encoders (e.g., not having keep-out regions), the server may pull content from outside the frames of the video stream into frames during the encoding process. As a result, during the decoding process of the composited video stream by the user device, the decoder may use references from outside the frames, which may be occupied by video content regions of different video streams in the composited video stream after compositing. Vectors referencing regions during encoding that may cause decoding problems can be updated as described with reference to Figures 9A, 9B, 9C, and 10A, 10B.
[0083] 9A-9C illustrate examples of motion vector containment solutions. As shown in FIG. 9A, a motion vector 900 references a "null" or black region 902 from outside a video content region 904 of a frame 906. For example, as shown in FIG. 9B, the reference region 902 is located within the area occupied by another video content region 908 when a composite video stream frame is generated. As shown in FIG. 9C, the motion vector 900 moves to an alternative "safe" region 910 of "null" or black content above or below the first video content region 904. In some implementations, all vectors moving from the black content moved into the "safe" region can copy content from the same source in the previous frame (i.e., can reference the same safe block).
[0084] 10A-10B show another example of a motion vector containment solution. As shown in FIG. 10A, a motion vector 1000 references a region 1002 located between two video content regions 1001, 1003 of a composite video stream 1005. As shown in FIG. 10B, the referenced region 1002 is divided into sub-blocks 1004a and 1004b, for example, using partitioning. Sub-block 1004a contains only the retained content (of the video content region in question) or the black content between the two video content regions. Sub-block 1004b contains the area occupied by the other video content region (possibly also black content) and is moved to a "safe" region, as described with reference to FIGS. 9A-9C. A separate motion vector 1006a, 1006b is generated for each of the sub-blocks 1004a, 1004b, and an updated residual is generated for each of the sub-blocks 1004a, 1004b. Generating the updated residual involves decoding (e.g., using an inverse DCT) and re-encoding (e.g., performing the DCT again) the size of the sub-blocks 1004a, 1004b, where the re-encoded values are substantially the same as the original residual for the region 1002. For example, if the original region is 32x32 pixels, each sub-block may have a size of four 16x16 pixels.
[0085] In some implementations, a separate segment ID can be used for the sub-block and Q can be set to a very low value (compared to the rest of the tile) to ensure at least a threshold match between the original and re-encoded residual values of the sub-block, resulting in better resolution, higher quality, and greater fidelity than the originally encoded content. In other words, blocks with problematic motion vectors can be replaced with "intra-coded" blocks (and with segment IDs with low Q values).
[0086] In some implementations, for efficiency, the server can encode only half of a frame containing video content areas (e.g., containing content from a video stream) and pad (e.g., fill) the remaining frame with black to generate a full frame. Often, motion vectors are encoded relative to existing motion vectors from the current frame or a previous frame. Performing this process results in a motion vector being encoded relative to other reference motion vectors that extend from the edge of the frame (for new blocks). If the reference motion vector is too far from the edge of the image (e.g., >32 pixels), it is clipped, and the clipped version is used as the base for a new motion vector to encode the motion vector's difference from the reference motion vector (if any). However, when a reduced-width frame is expanded to full width in compressed-domain synthesis, some of these previously clipped reference motion vectors may be de-clipped (because they no longer extend beyond the edge of the frame), which may cause motion vectors that depend on them to shift. To solve this problem, the server can perform a de-clipping process on the frame's motion vectors during the pre-conditioning process, where any de-clipped reference motion vectors are identified and the dependent motion vector differentials are updated based on the new de-clipped base motion vectors.
[0087] 11 is a flowchart of an exemplary process 1100 for compressed domain composite video streaming. For convenience, process 1100 is described as being performed by one or more computer systems located at one or more locations and suitably programmed in accordance with this specification. For example, a digital component distribution system, such as digital component distribution system 110 in FIG. 1 (e.g., including server 200 and local area server 202 in FIG. 2), can perform process 1100.
[0088] A system (e.g., local area server 202) obtains a compressed video stream, each of which includes multiple frames (1102). Each frame of the compressed video includes a video content region occupying any number of contiguous tiles / slices. Each frame is encoded with a first segment identifier for pixels included in the video content region and a second segment identifier for pixels not included in the video content region. The frames of the compressed video are encoded using a set of static symbol probabilities.
[0089] In some implementations, the compressed video stream is encoded by a server (e.g., server 200) in data communication with local area server 202. The compressed video stream may be generated, for example, from a video stream obtained from a third-party content generator. Generating the compressed video stream may include rendering, by the server, a new frame for each frame of the video stream, the new frame including a video content region, the video content region being any number of contiguous tiles within the frame and corresponding to the content of the original frame. The video content region may include, for example, one or more tiles of a tile-based composite of the frame, as shown in FIG. 3A.
[0090] The new frame may further include black or "null" regions outside the video content areas. The frame can then be encoded (e.g., using the VP9 codec), where the symbol frequency probabilities are set using default values. In other words, the probabilities are not updated between frames (e.g., using error resilient mode in the VP9 framework).
[0091] In some implementations, generating the compressed video further includes generating a segmentation map for the frame and defining a first segment identifier for pixels included in the video content area of the new frame and a second segment identifier for pixels not included in the video content area (i.e., black or "null" areas).
[0092] In some implementations, the system may assign one or more parameter values (e.g., quantization level (Q), reference frame, loop filter value, motion vector value, etc.) to each of the segment identifiers. For example, a first segment identifier may be assigned a first Q, and a second segment identifier may be assigned a second, different Q.
[0093] In some implementations, generating the compressed video further includes rendering gaps (e.g., those illustrated in FIG. 8) within the video content regions of the compressed video streams along edges of the video content regions. The gaps may be at least 8 pixels wide, and when the compressed videos are included in the composite video stream, null regions of the video content regions of the compressed videos are adjacent to other video content regions of other compressed videos. The edges at which the null regions are encoded during the encoding process may depend in part on whether the compressed videos are included in the composite video stream as a primary stream or a secondary stream. For example, null regions are included on the right side of the video content regions of the secondary streams (e.g., 806) and on the left side of the video content regions of the primary streams (e.g., 804).
[0094] In some implementations, generating the compressed video stream includes generating a primary video stream, e.g., pre-conditioned and stored video content, and generating at least one secondary video stream (e.g., live stream content). For example, multiple video content streams can be pre-conditioned and stored for inclusion in a composite video stream that includes the live stream content, and different users can receive different composites of the primary video stream along with the same live stream video.
[0095] In some implementations, the local area server 202 obtains the second set of compressed video streams, where the frames of the second set of compressed video streams include the rendered video stream in more than 50% of the frames, e.g., all or nearly all of the frames. In other words, the local area server 202 can receive encoded public live stream video content from the server 200, which can be provided to user devices in response to requests for non-composite video stream content.
[0096] A system (e.g., a local area server 202) receives a request for video stream content from a user device (1104). The request (e.g., request 108) may include a request for live stream video content (e.g., a secondary stream) and user information that can be used to select a primary stream for inclusion in a composite video stream with the secondary stream for presentation on the user device. The local area server 202 can receive multiple requests from different user devices and use compressed domain composition to generate each unique composite video stream for presentation on each of the requesting user devices based in part on the user information.
[0097] In some implementations, the local area server 202 can pre-cache a set of primary streams at the local area server 202 in anticipation of generating a composite video stream(s) with one or more secondary streams (e.g., live streams).
[0098] The system generates, by a local area server, in response to a request from a user device, a compressed domain composite video stream from a first compressed stream and a second compressed stream of the plurality of compressed video streams (1106), wherein each frame of the compressed domain composite video stream includes a first video content region of the first compressed stream and a second video content region of the second compressed stream, the first video content region and the second video content region occupying non-overlapping regions of the frame corresponding to a complete tile / slice in the encoded video.
[0099] In some implementations, the system generates the composite video stream by replacing one or more tiles of the primary stream with one or more tiles of the secondary stream. For example, by replacing null tiles of the primary stream with video content area tiles of the secondary stream. The system also generates an updated uncompressed header from the uncompressed header information for each of the composite videos. For example, the system incorporates segment identifiers of the video content area of the primary stream and segment identifiers of the video content area of the secondary stream into the uncompressed header of the composite video stream using tile replacement and while still in the compressed domain.
[0100] The system provides (1108) to the user device a packet containing a set of frames of the compressed domain composite video stream that is decodable by a single decoder.
[0101] In some implementations, a user request for video content (e.g., for live stream content) includes a request for a non-composite video stream. In other words, the user request may be to view only the secondary stream. In such a case, the local area server can provide a packet containing a set of frames of a third video stream that includes only live stream content decodable by a single decoder.
[0102] 12 is a block diagram of an exemplary computer system 1200 that can be used to perform the operations described above. The system 1200 includes a processor 1210, a memory 1220, a storage device 1230, and an input / output device 1240. Each of the components 1210, 1220, 1230, and 1240 can be interconnected using, for example, a system bus 1250. The processor 1210 is capable of processing instructions for execution within the system 1200. In one embodiment, the processor 1210 is a single-threaded processor. In another embodiment, the processor 1210 is a multi-threaded processor. The processor 1210 is capable of processing instructions stored in the memory 1220 or on the storage device 1230.
[0103] Memory 1220 stores information internal to system 1200. In one embodiment, memory 1220 is a computer-readable medium. In one embodiment, memory 1220 is a volatile memory unit. In other embodiments, memory 1220 is a non-volatile memory unit.
[0104] Storage device 1230 is capable of providing mass storage for system 1200. In one implementation, storage device 1230 is a computer-readable medium. In various different implementations, storage device 1230 may include, for example, a hard disk device, an optical disk device, a storage device shared over a network by multiple computing devices (e.g., a cloud storage device), or some other mass storage device.
[0105] Input / output device(s) 1240 provide input / output operations to system 1200. In one embodiment, input / output device(s) 1240 may include one or more of a network interface device, such as an Ethernet card, a serial communication device (e.g., an RS-232 port), and / or a wireless interface device (e.g., an 802.11 card). In other embodiments, input / output device(s) may include driver devices configured to receive input data and send output data to other devices, such as keyboards, printers, displays, and other peripheral devices 1260. However, other embodiments, such as mobile computing devices, mobile communication devices, set-top boxes, television client devices, etc., may also be used.
[0106] Although an exemplary processing system has been described with reference to FIG. 12, implementations of the subject matter and functional operations described herein can be implemented in other types of digital electronic circuitry, or computer software, firmware, or hardware, including the structures disclosed herein and structural equivalents thereof, or one or more combinations thereof.
[0107] The subject matter and actions and operations described herein can be implemented in digital electronic circuitry, tangibly embodied computer software or firmware, computer hardware, or a combination of one or more of them, including the structures disclosed herein and structural equivalents thereof. The subject matter and actions and operations described herein can be implemented as or in one or more computer programs (e.g., one or more modules of computer program instructions encoded on a computer program carrier for execution by or to control the operation of a data processing apparatus). The carrier may be a tangible, non-transitory computer storage medium. Alternatively, or additionally, the carrier may be an artificially generated propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, generated to encode information for transmission to a suitable receiving apparatus for execution by a data processing apparatus. The computer storage medium may be, or may be part of, a machine-readable storage device, a machine-readable storage substrate, a random access memory device, or a serial access memory device, or a combination of one or more of these. The computer storage medium is not a propagated signal.
[0108] The term "data processing apparatus" encompasses all kinds of apparatus, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. A data processing apparatus may include special-purpose logic circuitry, such as an FPGA (field-programmable gate array), an ASIC (application-specific integrated circuit), or a GPU (graphics processing unit). An apparatus also optionally includes, in addition to hardware, code that creates an execution environment for a computer program (e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these).
[0109] A computer program can be written in any type of programming language, including compiled or interpreted, or declarative or procedural languages, and the computer program can be deployed in any form, such as a standalone program, e.g., an app, or as a module, component, engine, subroutine, or other unit suitable for execution in a computing environment, which may include one or more computers interconnected by a data communications network in one or more locations.
[0110] A computer program may, but need not, correspond to a file in a file system: a computer program can be stored in part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinate files (e.g., files that store one or more modules, subprograms, or portions of code).
[0111] The processes and logic flows described herein may be performed by one or more computers executing one or more computer programs to perform operations by manipulating input data to generate output. The processes and logic flows may also be performed by special purpose logic circuitry (e.g., FPGAs, ASICs, or GPUs), or a combination of special purpose logic circuitry and one or more programmed computers.
[0112] A computer suitable for executing a computer program may be based on a general-purpose or special-purpose microprocessor, or both, or any other type of central processing unit. Generally, the central processing unit receives instructions and data from a read-only memory or a random-access memory, or both. The essential components of a computer are a central processing unit for executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry.
[0113] Generally, a computer also includes or is operably coupled to one or more mass storage devices and is configured to receive data from or transfer data to the mass storage devices. The mass storage devices may be, for example, magnetic, magneto-optical, or optical disks, or solid-state drives. However, a computer need not have such devices. Furthermore, a computer may be incorporated into other devices, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive, to name just a few.
[0114] To provide interaction with a user, the subject matter described herein can be implemented on one or more computers, which can have or be configured to communicate with a display device (e.g., an LCD (liquid crystal display) monitor, or a virtual reality (VR) or augmented reality (AR) display) for displaying information to the user and an input device (e.g., a keyboard and a pointing device, such as a mouse, trackball, or touchpad) through which the user can provide input to the computer. Similarly, other types of devices can be used to provide interaction with the user. For example, feedback and responses provided to the user can be any form of sensory feedback (e.g., visual, auditory, audio, or tactile), and input from the user can be received in any form, including acoustic, audio, or tactile input, including touch actions or gestures, or motor actions or gestures, or directional actions or gestures. Additionally, the computer can interact with the user by sending documents to and receiving documents from devices used by the user. For example, a computer may interact with a user by sending a web page to a web browser on the user's device in response to a request received from the web browser, or by interacting with an app running on the user's device (e.g., a smartphone or electronic tablet). A computer may also interact with a user by sending a text message or another form of message to a personal device (e.g., a smartphone running a messaging application) and receiving a response message from the user in return.
[0115] This specification uses the term "configured to" in connection with systems, devices, and computer program components. To say that one or more computer systems are configured to perform particular operations or actions means that the system has installed thereon software, firmware, hardware, or a combination thereof that, when in operation, causes the system to perform those operations or actions. To say that one or more computer programs are configured to perform particular operations or actions means that the one or more programs contain instructions that, when executed by a data processing device, cause the device to perform the operation or action. To say that special purpose logic circuitry is configured to perform a particular operation or action means that the circuitry has electronic logic that performs the operation or action.
[0116] The subject matter described herein can be implemented in a computing system that includes a back-end component (e.g., acting as a data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an embodiment of the subject matter described herein), or includes any combination of one or more of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.
[0117] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some implementations, a server transmits data, e.g., HTML pages, to a user device for the purpose of displaying the data to and receiving user input from a user interacting with the device acting as a client. Data generated at the user device, e.g., results of user interaction, can be received from the device by the server.
[0118] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of what is claimed, which is defined by the claims themselves, but rather as descriptions of features that may be unique to particular embodiments of a particular invention. Certain features described herein in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can be implemented separately in any suitable subcombination or in multiple embodiments. Furthermore, even if features may be described above as functioning in a particular combination and originally claimed as such, one or more features from the claimed combination may, in some cases, be deleted from the combination, and the claims may cover subcombinations or variations of the subcombination.
[0119] Similarly, while operations are illustrated in the figures and described in the claims in a particular order, this in itself should not be understood as requiring that such operations be performed in the particular order or sequential order shown, or that all of the operations shown be performed, to achieve desirable results. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the program components and systems described generally can be integrated together in a single software product or packaged in multiple software products.
[0120] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. By way of example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. 1. A computer-implemented method comprising: obtaining, by a local area server, a plurality of compressed video streams, each compressed video stream including a plurality of frames, each frame including a video content region including a portion of said frame; Each frame is encoded using a segment identifier for pixels contained in said video content region; obtaining the plurality of frames, wherein the plurality of frames are encoded using a set of static symbol frequencies; receiving, by the local area server, a request for video stream content from a user device; combining, by the local area server in response to the request from the user device, a first compressed video stream and a second compressed video stream of the plurality of compressed video streams to obtain a compressed domain composite video stream; each frame of the compressed domain composite video stream includes a first video content region of the first compressed video stream and a second video content region of the second compressed video stream; each frame including a respective first segment identifier for pixels corresponding to the first video content region and a second segment identifier for pixels corresponding to the second video content region; combining the first video content region and the second video content region, wherein the first video content region and the second video content region occupy non-overlapping portions of the frame; providing packets to the user device containing sets of frames of the compressed domain composite video stream decodable by a single decoder; A method comprising:
2. obtaining the plurality of compressed video streams receiving, by a server, a plurality of video streams; generating the plurality of compressed video streams, wherein for each of the compressed video streams: for each frame of the plurality of frames of the video stream, rendering a new frame that includes the portion of the frame and includes a video content area that corresponds to content of the frame; defining, for the video content region, the segment identifiers for pixels included in the video content region; encoding the frame using a set of static symbol frequencies; and generating the signal comprising:
3. The method of claim 2 , wherein the rendering further comprises rendering regions not included in the video content region as null content.
4. generating the plurality of compressed video streams, generating a primary video stream including context-responsive video content; generating at least one secondary video stream including the live stream content; 3. The method of claim 2, wherein the first compressed video stream comprises a primary video stream and the second compressed video stream comprises a secondary video stream.
5. 5. The method of claim 4, wherein said generating said primary video stream includes pre-conditioning and storing in a database context-responsive video for inclusion in said compressed domain composite video stream, and an appropriate subset of said context-responsive video content is pre-cached on said local area server.
6. Obtaining the compressed domain composite video stream includes: selecting, by the local area server, a primary content video from the database of primary video streams for inclusion in the compressed domain composite video stream in response to user-based selection criteria of the user device; The method of any one of claims 1 to 5, wherein in response to different user requests, the local area server selects different primary content videos for inclusion in the compressed domain composite video stream.
7. generating the plurality of compressed video streams, generating a video stream comprising at least one of context-responsive video content and live stream content; The method of claim 2 , wherein the first compressed video stream and the second compressed video stream are generated from the video stream.
8. encoding the frame using a set of static symbol frequencies 3. The method of claim 2, comprising encoding each symbol probability of the frame with a default frequency value.
9. Defining the segment identifier for the video content region includes: The method of claim 2 , further comprising assigning a first quantization level to the segment identifiers of pixels included in the video content region.
10. Rendering is rendering a null region including a width within the video content region of the compressed video stream along an edge of the video content region; The method of claim 2 , wherein the null regions of the video content regions are located alongside replacement regions configured to be replaced with other compressed video streams in a composite video stream.
11. generating the compressed domain composite video stream from a first compressed stream and a second compressed stream; 2. The method of claim 1, comprising constructing an uncompressed header for each frame of the compressed domain composite video stream based on the segment identifiers of the first compressed stream and the second compressed stream, respectively.
12. 2. The method of claim 1 , wherein the frame comprises tiles, the video content region comprises one or more contiguous tiles of a tile-based composite of the frame, and the pixels included in the video content region comprise the one or more contiguous tiles.
13. The method of claim 12 , wherein the one or more contiguous tiles that include the video content area include half of the tiles of the tile-based composite of the frame.
14. acquiring, by the local area server, a second plurality of compressed video streams, each of the second plurality of compressed video streams including a plurality of frames; receiving, by the local area server, a second request for video stream content from the user device; providing, by the local area server to the user device in response to the second request, a second packet including a set of frames for one of the second plurality of compressed video streams; The method of any one of claims 1 to 5, further comprising:
15. One or more non-transitory computer storage media encoded with computer program instructions that, when executed by one or more computers, cause the one or more computers to: obtaining, by a local area server, a plurality of compressed video streams, each compressed video stream including a plurality of frames, each frame including a video content region including a portion of said frame; Each frame is encoded using a segment identifier for pixels contained in said video content region; obtaining the plurality of frames, wherein the plurality of frames are encoded using a set of static symbol frequencies; receiving, by the local area server, a request for video stream content from a user device; combining, by the local area server in response to the request from the user device, a first compressed video stream and a second compressed video stream of the plurality of compressed video streams to obtain a compressed domain composite video stream; each frame of the compressed domain composite video stream includes a first video content region of the first compressed video stream and a second video content region of the second compressed video stream; each frame including a respective first segment identifier for pixels corresponding to the first video content region and a second segment identifier for pixels corresponding to the second video content region; combining the first video content region and the second video content region, wherein the first video content region and the second video content region occupy non-overlapping portions of the frame; providing packets to the user device containing sets of frames of the compressed domain composite video stream decodable by a single decoder; One or more non-transitory computer storage media that cause operations to be performed, including:
16. obtaining the plurality of compressed video streams receiving, by a server, a plurality of video streams; generating the plurality of compressed video streams, wherein for each of the compressed video streams: for each frame of the plurality of frames of the video stream, rendering a new frame that includes the portion of the frame and includes a video content area that corresponds to content of the frame; defining, for the video content region, the segment identifiers for pixels included in the video content region; encoding the frame using a set of static symbol frequencies; 16. The computer storage medium of claim 15, wherein generating includes:
17. generating the plurality of compressed video streams, generating a primary video stream including context-responsive video content; generating at least one secondary video stream including the live stream content; 17. The computer storage medium of claim 15 or 16, wherein the first compressed video stream comprises a primary video stream and the second compressed video stream comprises a secondary video stream.
18. 20. The computer storage medium of claim 17, wherein generating the primary video stream includes pre-conditioning and storing in a database context-responsive video for inclusion in the compressed domain composite video stream, and an appropriate subset of the context-responsive video content is pre-cached on the local area server.
19. Obtaining the compressed domain composite video stream includes: selecting, by the local area server, a primary content video from the database of primary video streams for inclusion in the compressed domain composite video stream in response to user-based selection criteria of the user device; 20. The computer storage medium of claim 18, wherein in response to different user requests, the local area server selects different primary content videos for inclusion in the compressed domain composite video stream.
20. 1. A system comprising one or more computers and one or more storage devices having operable instructions stored thereon, the instructions, when executed by the one or more computers, causing the one or more computers to: obtaining a plurality of compressed video streams, each compressed video stream including a plurality of frames, each frame including a video content region including a portion of said frame; Each frame is encoded using a segment identifier for pixels contained in said video content region; obtaining the plurality of frames, wherein the plurality of frames are encoded using a set of static symbol frequencies; receiving a request for video stream content from a user device; combining a first compressed video stream and a second compressed video stream of the plurality of compressed video streams in response to the request from the user device to obtain a compressed domain composite video stream; each frame of the compressed domain composite video stream includes a first video content region of the first compressed video stream and a second video content region of the second compressed video stream; each frame including a respective first segment identifier for pixels corresponding to the first video content region and a second segment identifier for pixels corresponding to the second video content region; combining the first video content region and the second video content region, wherein the first video content region and the second video content region occupy non-overlapping portions of the frame; providing packets to the user device containing sets of frames of the compressed domain composite video stream decodable by a single decoder; A system that causes an operation including
Citation Information
Patent Citations
Image processor, image recorder and image reproducer
JP2002027393A
Encoder, decoder, coding method, and decoding method
JP2019033456A
Method, device and computer program for improving rendering display during streaming of timed media data
JP2022066370A
Optimized reduced bitrate encoding for titles and credits in video content
US11496738B1
Multi layer display apparatus
US20140184669A1