Video streaming

Compressed domain synthesis generates a single decodable video stream from multiple streams, addressing the challenge of simultaneous video display with a single decoder, enhancing efficiency and reducing computational load.

JP7860341B2Active Publication Date: 2026-05-15GOOGLE LLC
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
GOOGLE LLC
Filing Date
2024-10-22
Publication Date
2026-05-15

Smart Images

  • Figure 0007860341000001
    Figure 0007860341000001
  • Figure 0007860341000002
    Figure 0007860341000002
  • Figure 0007860341000003
    Figure 0007860341000003
Patent Text Reader

Abstract

A method, system, and apparatus, including a computer program encoded on a computer storage medium, for compressed-domain combining video streams, wherein a server obtains compressed video streams, each compressed video stream including a plurality of frames, each frame including a video content region that is a portion of the frame and encoded using a segment identifier of pixels included in the video content region and encoded using a set of static symbol frequencies, the server receives a request for video stream content from a user device, combines a first compressed video stream and a second compressed video stream to obtain a compressed-domain combined video stream including a first video content region of the first compressed video stream and a second video content region of the second compressed video stream, and provides packets to the user device including a set of frames of the compressed-domain combined video stream that are decodable by a single decoder.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority to U.S. Application No. 18 / 495,405, filed on October 26, 2023. The disclosure of the prior application is considered part of the disclosure of this application and is incorporated into the disclosure of this application by reference.

[0002] This specification relates to streaming video content.

Background Art

[0003] Picture - in - Picture (PiP) streaming is a multi - video display mode that enables a user to view two or more video streams presented in the same window of a user device. For example, using PiP, two or more different video streams with different content can be presented within the same display window, so that a user can experience multiple streams simultaneously. In some cases, the user device may include only a single decoder so that it can decode only one video stream at a time for presentation on the user device.

Summary of the Invention

[0004] The subject matter of this specification generally relates to compressed - domain synthesis of video streams.

[0005] More specifically, the subject of this specification concerns the use of compressed domain synthesis technology to provide a user with a fully personalized synthesized video stream containing two or more video streams. The synthesized video stream may contain at least two separate streams synthesized within a compressed domain on an edge server and provided to an end-user device, and the synthesized video stream is decodeable by the end-user device using a single decoder. The synthesized stream may be a PiP or mosaic video stream containing a primary video stream (e.g., third-party content video) and a secondary video stream (e.g., live stream video content), which function as a 50 / 50 split screen for the end-user device.

[0006] In general, one innovative aspect of the subject matter described herein can be embodied in a manner that includes an action by a local area server to acquire a compressed video stream, each compressed video stream containing multiple frames, each frame containing a video content region that is a portion (e.g., up to 50%) of the frame. Each frame of the compressed video stream is encoded using segment identifiers of the pixels contained in the video content region, and the frames of the compressed video stream are encoded using a set of static symbol frequencies. To acquire a compressed domain composite video stream, the local area server receives a request from a user device for video stream content and, in response to the request from the user device, synthesizes a first compressed video stream and a second compressed video stream. Each frame of the compressed domain composite video stream contains a first video content region of the first compressed video stream and a second video content region of the second compressed video stream. Each frame of the compressed domain composite video stream contains a first segment identifier for each pixel corresponding to the first video content region and a second segment identifier for the pixels corresponding to the second video content region. The first and second video content regions occupy the non-overlapping portions of the frame. The local area server provides the user device with a packet containing a set of frames of a compressed domain-composite video stream that can be decoded by a single decoder.

[0007] Other embodiments of this model include a corresponding computer system, apparatus, and a computer program recorded on one or more computer storage devices, each configured to perform the actions of this method.

[0008] The above and other embodiments may each include one or more of the following features, either individually or in combination. Specifically, one embodiment includes all of the following features in combination. In some embodiments, obtaining a compressed video stream includes receiving a video stream from a server and generating each of the compressed video streams, which includes rendering a new frame for each of the multiple frames of the video stream that includes a portion of the frame and a video content area corresponding to the content of the frame, defining a segment identifier for the pixels contained in the video content area, and encoding the frame using a set of static symbol frequencies.

[0009] In some embodiments, rendering further includes rendering areas not included in the video content area as null content.

[0010] In some embodiments, generating a compressed video stream involves generating a primary video stream containing context-responsive video content and generating at least one secondary video stream containing live stream content, where the first compressed video stream is the primary video stream and the second compressed video stream is the secondary video stream. Generating the primary video stream may involve pre-caching and storing context-responsive video in a database for inclusion in a compressed domain composite video stream, and an appropriate subset of the context-responsive video content is pre-cached on a local area server.

[0011] In some embodiments, generating a compressed domain composite video stream involves the local area server selecting primary content videos from a database of primary video streams to include in the compressed domain composite video stream in response to user-based selection criteria of the user device. In response to different user requests, the local area server selects different primary content videos to include in the compressed domain composite video stream.

[0012] In some embodiments, generating multiple compressed video streams involves generating a video stream containing at least one of context-responsive video content and live stream content, wherein the first and second compressed video streams are generated from the video stream.

[0013] In some embodiments, encoding a frame using a static set of symbol frequencies involves encoding each symbol probability in the frame with a default frequency value.

[0014] In some embodiments, defining a segment identifier for a video content region further includes assigning a first quantization level to the segment identifier of pixels contained within the video content region.

[0015] In some embodiments, rendering further includes rendering a null region, including its width, along the edge of the video content region, within the video content region of the compressed video stream, where the null region of the video content region is located next to a replacement region configured to be replaced by another compressed video stream in the composite video stream.

[0016] In some embodiments, generating a compressed domain composite video stream from a first compressed stream and a second compressed stream involves constructing an uncompressed header for each frame of the compressed domain composite video stream based on the respective segment identifiers of the first and second compressed streams.

[0017] In some embodiments, a frame includes tiles, a video content area includes one or more consecutive tiles of a tile-based composite of frames, and pixels included in the video content area include one or more consecutive tiles. One or more consecutive tiles including the video content area are half of the tiles of a tile-based composite of frames.

[0018] In some embodiments, the method further includes obtaining a second set of compressed video streams by a local area server, each of which contains multiple frames. The local area server receives a second request from the user device for video stream content and, in response to the second request, provides the user device with a second packet containing a set of frames for one of the second sets of compressed video streams.

[0019] The subject matter described herein can be implemented in specific embodiments to achieve one or more of the following advantages: End devices can receive, decode, and display a bitstream-level synthetic stream as a single stream, even if the synthetic stream could be from two completely unrelated video streams. As a result, end devices do not experience any additional latency or increased computational requirements to display a synthetic stream containing two separate streams. The solutions described herein are lightweight in terms of computation in real-time processing at the cache edge server and user devices, although they involve backend pre-tuning. The solutions described can increase network bandwidth efficiency for providing multiple customized synthetic streams to different user devices, because the server delivers only the live stream and N primary content streams (e.g., those that can be pre-cached to the cache edge server).

[0020] Furthermore, the features described herein enable synthetic streams, allowing end-user devices to receive fully customized synthetic video streams containing two or more video streams (e.g., a live stream and a third-party content stream) without substantial additional computing requirements on the end-user device or edge server during real-time provisioning of live streams. In other words, the technology described herein substantially improves the video stream caching capacity of cache edge servers, which only need to store N+M video streams instead of N×M video streams.

[0021] As described herein, video compression domain synthesis processing generates, for each user by a cache edge server, a single synthesized stream within the compression domain from two or more input compressed streams, which is done without requiring the edge server or user device to perform arithmetic decoding and / or arithmetic encoding of the compressed streams (e.g., concatenating video metadata with video input). Otherwise, this can be prohibitively expensive and computationally intensive to provide a custom stream for each user. The decoder of the user device effectively "views" a single video. As a result, no additional processing requirements are needed from the decoder to decode additional frames (e.g., invisible frames, alt-ref frames, etc.). By placing most of the processing in a preconditioning step before providing the compressed video to the cache edge server, the reduced processing requirements are met by the cache edge server and the end user device.

[0022] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.

Brief Description of the Drawings

[0023] [Figure 1] FIG. 1 is a block diagram of an exemplary environment for compressed domain synthetic video streaming. [Figure 2] FIG. 1 is a block diagram of an exemplary environment for compressed domain synthetic video streaming. [Figure 3A] FIG. 2 shows a schematic example of tile-based synthesis. [Figure 3B] FIG. 2 shows a schematic example of tile-based synthesis. [Figure 4] FIG. 3 is a schematic diagram of an exemplary compressed domain synthesis process. [Figure 5A] FIG. 4 is a schematic diagram of an example of a VP9-based frame structure. [Figure 5B] FIG. 5 shows an example of a corrupted video stream. [Figure 6] An example of a composite video frame including a plurality of tiles is shown. [Figure 7] A schematic diagram of an example of reference frames among a primary stream, a secondary stream, and their composite video stream is shown. [Figure 8] A schematic diagram of an example of a gap inserted between two video streams of a composite video stream is shown. [Figure 9A] An example of a motion vector containment solution is shown. [Figure 9B] An example of a motion vector containment solution is shown. [Figure 9C] An example of a motion vector containment solution is shown. [Figure 10A] Another example of a motion vector containment solution is shown. [Figure 10B] Another example of a motion vector containment solution is shown. [Figure 11] A flowchart of an exemplary process for compressed domain composite video streaming. [Figure 12] An exemplary block diagram of a computer.

MODE FOR CARRYING OUT THE INVENTION

[0024] Like reference symbols and designations in the various drawings refer to like elements.

[0025] As used herein, the terms “compressed domain composite video stream” and “compressed domain composite video processing” refer to the bitstream-level synthesis of two or more encoded (i.e., compressed) video streams to form an output composite video stream in order to incorporate and present all or part of an input video stream. In other words, a picture-in-picture or mosaic of two or more different video streams is presented in frames of the composite video stream. The generation of the composite video stream is performed in the compressed domain, that is, without decoding the composite video to generate each frame of the composite video, and without subsequently re-encoding the frames of the composite video. Instead, as will be described in more detail below, the frames of the compressed domain composite video are generated from two or more compressed videos without the need to decode / re-encode the videos to be synthesized.

[0026] As used herein, the term “pre-adjustment” refers to receiving the output of a general-purpose video encoder (e.g., a hardware encoder) and adapting that output to enable lightweight synthesis of the aforementioned content (e.g., by copying the appropriate video metadata into a properly constructed output frame). Several examples of codecs for encoding / compressing and decoding / decompressing video content are provided herein, but for clarity, the VP9 codec is used as the primary example of the process described.

[0027] Figure 1 is a block diagram of an exemplary environment 100 for compressed domain composite video streaming. The exemplary environment 100 includes a network 102 such as a local area network (LAN), wide area network (WAN), the internet, or a combination thereof. Network 102 connects an electronic document server 104 ("electronic document server"), user devices 106, and a digital component distribution system 110 (also known as DCDS 110). The exemplary environment 100 may include many different electronic document servers 104 and user devices 106.

[0028] User device 106 is an electronic device capable of requesting and receiving resources (e.g., electronic documents) via network 102. Exemplary user device 106 includes personal computers, mobile communication devices, and other devices capable of sending and receiving data via network 102. User device 106 typically includes user applications such as web browsers to facilitate data transmission and reception via network 102, but native applications running on user device 106 can also facilitate data transmission and reception via network 102.

[0029] One or more third parties 130 include content providers, product designers, product manufacturers, and other parties involved in the design, development, marketing, or distribution of videos, products, and / or services.

[0030] An electronic document is data that presents a set of content on a user device 106. Examples of electronic documents include web pages, word processing documents, portable document format (PDF) documents, images, videos, search results pages, and feed sources. Native applications (e.g., “apps”), such as applications installed on mobile, tablet, or desktop computing devices, are also examples of electronic documents. An electronic document 105 (“electronic document”) can be provided to the user device 106 by an electronic document server 104. For example, the application server 104 may include a server that hosts a publisher’s website. In this example, the user device 106 can initiate a request for a web page of a given publisher, and the electronic document server 104 hosting the web page of a given publisher can respond to the request by sending Machine Hypertext Markup Language (HTML) code that initiates the presentation of the given web page on the user device 106.

[0031] Request 108 may include data specifying an electronic document (e.g., video content) and the characteristics of a location where the digital content can be presented. For example, data may be provided to DCDS110 specifying a reference to the electronic document (e.g., a composite video stream) on which the digital content is presented, an available location in the electronic document available for presenting the digital content (e.g., the digital content area within a frame of the composite video), the size of the available location, and / or the position of the available location within the presentation of the electronic document. Similarly, data specifying keywords ("document keywords") or entities (e.g., people, places, or things) referenced by the electronic document may also be included in Request 108 (e.g., as payload data) and provided to DCDS110 to facilitate the identification of digital content items that are eligible for presentation with the electronic document.

[0032] Request 108 may also include data related to other information, such as information provided by the user, geographical information indicating the state or region from which the request was submitted, or other information that provides context about the environment in which the digital content is displayed (e.g., the type of device on which the digital content is displayed, such as a mobile device or tablet device). Information provided by the user may include demographic data about the user of user device 106. For example, demographic information may include, among other characteristics, age, gender, geographical location, education level, marital status, household income, occupation, hobbies, social media data, and whether the user owns a particular item.

[0033] Where the systems described herein may collect or use personal information about a user, the user may be given the opportunity to control whether the program or feature collects personal information (e.g., information about the user's social networks, social behavior or activities, occupation, user preferences, or current geographical location), or whether and / or how it receives content that may be more relevant to the user from the content server. Furthermore, certain data may be anonymized in one or more ways so that personally identifiable information is removed before it is stored or used. For example, user identification information may be anonymized so that personally identifiable information cannot be determined, or, if location information is obtained, the user's geographical location may be generalized (to the city, zip code, or state level, etc.) so that the user's specific location cannot be determined. Thus, the user may control how information is collected about them and how it is used by the content server.

[0034] Request 108 may also provide data specifying the characteristics of user device 106, such as information identifying the model of user device 106, the configuration of user device 106, or the size (e.g., physical size or resolution) of the electronic display (e.g., touchscreen or desktop monitor) on which the electronic document is presented. Request 108 may be transmitted, for example, over a packetized network, and request 108 itself may be formatted as packetized data having a header and payload data. The header may specify the destination of the packet, and the payload data may contain any of the above information.

[0035] DCDS110 selects digital content to be presented with a given electronic document in response to receiving request 108 and / or using the information contained in request 108. For example, DCDS110 selects a digital content item from the digital component database 112, e.g., one of the repositories of pre-tuned videos available to generate a composite video stream with at least one other pre-tuned video.

[0036] In some embodiments, DCDS110 is implemented in a distributed computing system (or environment), which includes, for example, a server (e.g., Server 200 in Figure 2) and a set of multiple computing devices (e.g., Local Area Server 202 in Figure 2), which are interconnected and identify and deliver digital content in response to requests 108. For example, as further described with reference to Figure 2, a cache edge server can receive a request to deliver digital content and generate a composite video stream containing the electronic document(s) to be delivered to an end-user device. The set of multiple computing devices works together to identify a set of digital content from a corpus of millions or more of available digital content (e.g., pre-arranged video content that can be used to generate a composite video stream of a compressed domain) that is suitable to be presented with the electronic document as a composite video stream. The millions or more of available digital content can be indexed in, for example, a digital component database 112. Each digital content index entry may include delivery parameters (e.g., selection criteria) that refer to the corresponding digital content and / or condition the delivery of the corresponding digital content.

[0037] In some embodiments, digital components from the digital component database 112 may include content provided by a third party 130. For example, DCDS 110 can present video content stored in the digital component database 112. As will be described in more detail below, digital components (e.g., video) stored in the digital component database 112 can be pre-tuned and stored in the database 112 in anticipation of being included in a composite video stream for presentation on a user device. Furthermore, digital components may be real-time (e.g., live stream video), which can also be pre-tuned by DCDS 110 to include live stream video within a composite video stream for presentation on a user device. As used herein, the term “primary stream” generally refers to pre-tuned context-responsive content for one or more users stored in the digital component database 112 for later incorporation by DCDS 110 into a customized composite video stream. As used herein, the term “secondary stream” refers to pre-tuned live stream content that can be incorporated into a customized composite video stream by DCDS 110. In this specification, a composite video stream is described as including a primary stream and a secondary stream, but a composite video stream can be generated using two or more video streams. For example, a mosaic composite video stream may include three or more video streams, each of which may be primary or secondary stream video content.

[0038] The identification of eligible digital content can be segmented into multiple tasks, which are then assigned among computing devices within a set of multiple computing devices. For example, different computing devices can each analyze different parts of the digital component database 112 to identify various digital components with distribution parameters that match the information contained in request 108.

[0039] DCDS110 aggregates the results received from a set of multiple computing devices and uses the information associated with the aggregated results to select one or more instances of digital content to be provided in response to request 108. DCDS110 can then generate and transmit response data 114 (e.g., digital data representing the response) over the network 102, enabling the user device 106 to integrate the selected set of digital content into a given electronic document. As a result, the selected set of digital content and electronic document content are presented together on the display of the user device 106.

[0040] As described above, DCDS110 can be implemented in a distributed computing system (or environment), which includes, for example, a server (e.g., server 200) and a set of multiple computing local area servers 202 (e.g., cache edge devices), which are interconnected and identify and deliver digital content in response to requests 108.

[0041] Figure 2 is a block diagram of an exemplary environment 201 for compressed domain synthetic video streaming. Server 200 can generate, process, and / or store video content. Video content can be preprocessed by Server 200 to generate encoded (e.g., compressed) video content. For example, Server 200 can encode context-responsive video content and store the encoded video content in a repository (e.g., digital component database 112) for later incorporation into a customized synthetic video stream. In another example, Server 200 can encode live stream video content in preparation for exposing the live stream video content to user devices.

[0042] Although described here as an action performed by server 200, in some embodiments, some or all of the encoding of the video content (e.g., context-responsive video content and / or live-stream video content) can be pre-coordinated (e.g., encoded) by one or more third parties. For example, a third-party content publisher (e.g., third party 130) can pre-coordinate video content generated by the third party and provide the pre-coordinated video content to server 200, for example, via network 102.

[0043] Server 200 pre-processes (i.e., pre-sets) the video content by encoding it using the VP9 codec, the details of which are described below. For simplicity, the encoding of the video content will be described with reference to the VP9 codec, but other codecs may be used, for example, H.264, H.265, AV1, or other codecs configured to provide sufficient features to enable block substitution (e.g., tile substitution), as described below.

[0044] In (1), the server 200 can receive, process, and store primary video content 205 (e.g., third-party digital content) which may be accessible for future use when generating user-customized synthetic video content. As described above, the primary video stream may include content customized for an end-user device (e.g., selected to respond contextually to one or more user attributes), which may be unique to each end-user device receiving a synthetic (e.g., PiP) video stream from the DCDS 110. For example, the primary video stream may be selected based on contextual information including one or more attributes, such as the user's demographic attributes, geography, viewing history, or other user-based preferences or information. The server 200 can pre-generate and pre-configure a repository of primary content stream videos, which may include different content that can be selected to be included in the synthetic stream at a later point in time, partly based on specific conditions (e.g., user preferences, contextual information obtained from or associated with the receiving user's device, etc.).

[0045] (2) The server can receive and process secondary video content 207, for example, live stream (e.g., real-time broadcast) video content that may be accessible for use in generating customized composite video content. The server 200 can publish at least two versions of the secondary content, including (A) only the secondary stream and (B) a pre-processed version of the secondary stream that can be combined with primary streams to generate a composite stream.

[0046] In some embodiments, the server may pre-cach a repository of primary stream digital content at local area servers (e.g., local area servers 202a, 202b, 202c) selected based on a set of user devices served by the local area servers, for example, the geographical location of the local area servers. For example, server 200 can periodically pre-cach a set of primary stream digital content at local area servers 202a-c to maintain an updated repository of pre-coordinated video content at local area servers that may be combined by local area servers 202a-c, and generate a composite video stream for presentation at user device 204. In some embodiments, different sets of primary stream digital content can be cached at each local area server, for example, partly based on the geographical location of the local area servers and the geographical location of the user devices served by the local area servers.

[0047] In (4), the local area server 202a receives a content request from a user device, for example, user device 204a. The content request may include a request to display a secondary stream, i.e., live stream video content (A). The local area server 202a can identify the possibility of providing the secondary stream in the composite stream, for example, as a PiP in the primary stream containing third-party content. The selection process is illustrated with reference to Figure 1, but can generally be selected based on user information of the user device.

[0048] In some embodiments, the auction can be conducted (e.g., on the auction server) for third-party content from a repository of pre-tuned primary stream videos, and the winning customized pre-tuned primary stream from the third-party content repository can be returned to the edge server.

[0049] The local area server 202a performs compressed domain synthesis of the selected primary and secondary streams to generate a composite video. The local area server 202a can provide the end-user device with a composite stream containing the primary stream (B) and the secondary stream (A). In (5), the end-user device receives and decodes the composite stream and presents the composite stream containing the live stream (A) and third-party content (B) on the user device's display within the user device's viewport.

[0050] The local area server 202a can generate two or more different composite streams and provide them to each end-user device. For example, in (6), the local area server 202a provides the second user device 204b with different composite streams containing the same secondary content (A) with different primary content (C).

[0051] In some embodiments, the local area server can perform this process sequentially and request and provide multiple (e.g., two or more) primary content streams sequentially with an encoded composite stream that includes a secondary stream. In some embodiments, the local area server 202a can provide multiple (e.g., two or more) sequential secondary streams along with the encoded composite stream.

[0052] In some embodiments, a request (e.g., request 108) may include an opt-out for displaying primary content (e.g., non-livestream video content), thereby allowing the local area server (e.g., local area server 202b) to provide only secondary streams (A) (e.g., livestreams) to be presented to the user device(s) 204c in (7). In some embodiments, as will be described in more detail below with reference to Figure 4, a request (e.g., request 108) may include a change between a request to display a composite stream and a request to display a livestream.

[0053] As shown in Figure 2, the local area server 202a performs compressed domain synthesis of selected primary and secondary streams to generate a composite video in response to content requests from user devices. Thus, the local area server 202a can generate different customized composite videos for each device (or cluster of devices) while minimizing the decoding effect downstream of the user devices. In practice, each user device 204 receives a single customized composite video from the local area server 202a that can be decoded by a single decoder.

[0054] In some cases, all different end-user devices may receive a composite stream from the local area server 202a. In some cases, each different end-user device may receive either (i) a composite stream or (ii) only a secondary stream from the local area server 202a. In some cases, at least one end-user device may receive a composite stream that has a primary stream different from at least one other end-user device.

[0055] Each frame can be divided into sub-parts, for example, sub-parts containing one or more tiles (e.g., VP9 codec) or one or more slices (e.g., H.264 codec), each sub-part being encoded independently of the others. Within the VP9 codec, a tile substitution process can be used to generate a synthetically encoded stream using compressed domain synthesis techniques, for example, on a local edge server. Within the VP9 framework, each tile is independently arithmetically encoded and decoded, and tiles can be encoded and / or decoded in parallel. Depending in part to the constraints of the VP9 specification, different numbers of tiles can be used. For example, a 1080p stream may contain 4 tiles, a 4k stream 8 tiles, and a 480p stream 2 tiles.

[0056] For example, as shown in Figures 3A and 3B, each frame of the primary stream can be divided into multiple sub-parts, for example, into tiles 1, 2, 3, and 4 in Figure 3A, or into sub-parts 5, 6, 7, 8, 9, 10, 11, and 12 in Figure 3B. Each sub-part can be decoded independently of the other sub-parts, and the system 110 (e.g., local area server 202a) can replace one or more of the sub-parts with tile parts of the secondary stream during the compressed domain synthesis process to generate a composite video.

[0057] As further illustrated with reference to Figure 4, pre-adjusting each video stream involves rendering each frame of the adjusted video stream so that it includes video content areas and "null" or black areas, and the null or black areas of the pre-processed video stream can be replaced with video content areas from other video streams. In such cases, the primary and secondary streams are encoded so that tiles are drawn from each stream and combined to form a composite frame. For example, each video stream is encoded as half-width content in the first part of the frame and pre-adjusted as described herein, and another second part of the frame is a compressed domain (CD) synthesized by the encoder to produce full-width content. In other words, pre-adjustment involves encoding only the tiles containing content (e.g., non-null tiles). In another example, as shown in Figure 3A, during pre-adjustment, frames of video streams included as primary content streams in the composite video stream can be rendered so that the video content regions are inside tiles 3 and 4, and frames of video streams included as secondary content streams in the composite video stream can be rendered so that the video content regions of the secondary content streams are inside tiles 1 and 2.

[0058] In some embodiments, instead of (or in addition to) a "null" or black area, pre-processing may include rendering a boundary that encloses the video content area. For example, the boundary may be a technical feature (e.g., texture, pattern, graphic, etc.).

[0059] In some embodiments, a primary content stream is preprocessed and stored in a digital component database to condition the content stream, and as a result, the primary content stream may be included at a later point in a composite stream for presentation on the user device. The condition involves rendering a new frame for each frame of the primary content stream such that the original content occupies no more than 50% of the frame (e.g., no more than 50% of the frame including edge regions shared with replaced content), with the remainder of the frame being rendered as "null" content or black content. In other words, the original content can be rendered within the video content area of ​​the new frame that forms the preprocessed primary content stream. The video content area of ​​the primary content stream can be selected to be oriented in the same direction as the frame, for example, the first (left) portion of the frame. A secondary stream (e.g., live stream content) can be preprocessed to render the content of each frame of the secondary stream so that it occupies no more than 50% of the new frame, with the remainder of the new frame being rendered as "null" content or black content. The original content can be rendered within the video content area of ​​the new frame that forms the preprocessed secondary content stream. The video content area of ​​the secondary content stream can be selected to have the same orientation as the frame, for example, the second (right-hand side) portion of the frame.

[0060] In some embodiments, the local area server (e.g., local area server 202a) can switch between providing the user device with encoded PiP composite streams, live streams (e.g., only secondary streams), and third-party content (e.g., only primary streams).

[0061] Figure 4 is a schematic diagram of an exemplary compressed domain synthesis process. Local Area Server 402 (e.g., Local Area Server 202a) issues video streams (1A) (e.g., live stream content), 1B (e.g., live stream content pre-adjusted for inclusion in the synthesis stream), and (2) (e.g., pre-adjusted content for inclusion in the synthesis stream). Local Area Server 402 can choose to (i) provide video stream (1A) containing only live stream video content, or (ii) provide a synthesis video stream containing video streams (1B) and (2). For example, Local Area Server 402 can choose to provide video stream (1A) or video stream (1B) to the user device, partly based on a request from a user device containing user device information.

[0062] If only a live stream (1A) is provided to the user, the stream can be encoded according to a standard process (for example, as if there were no PiP stream). In other words, it is done without the additional pre-adjustment steps used to prepare the video stream for inclusion in a composite video stream. The live stream (1A) is provided to the user device as a single video stream (3) that can be decoded by a single decoder on the user device.

[0063] When a live stream is provided to a user device within a composite video stream, the stream is conditioned to be replaced as a PiP by a pre-tuned video stream (2) (1B). The local area server 402 generates and provides a composite video stream as a single video stream (3) from the video stream (1B) and video stream (2), which can be decoded by a single decoder on the user device. With the pre-tune described herein, the local area server 402 can provide a unique composite video stream for each user device.

[0064] As mentioned earlier, the video content stream is pre-adjusted for inclusion in the composite video stream. Figure 5A is a schematic diagram of an example of a VP9-based frame structure, serving as a high-level overview for the following explanation. Within the VP9 framework, each sub-part of a frame (e.g., a "tile") is individually arithmetically encoded.

[0065] For example, as schematically shown in Figure 5A, the frame structure includes an uncompressed header containing information about how to decode the full frame (e.g., in the VP9 and AV1 examples), so that if the frame is a composite frame containing content from two or more different video streams, the uncompressed header contains information for decoding all the video streams that make up the composite. For example, the uncompressed header of a frame in a composite stream containing a primary stream and a secondary stream contains information for decoding the primary stream portion and the secondary stream portion within the composite stream. Therefore, the information in the uncompressed header of pre-tuned video streams (before compositing the compressed domains) is similarly structured, so that, for example, at least some aspects of the uncompressed header match across all pre-tuned streams. Otherwise, as shown in Figure 5B, structural differences can corrupt the video file.

[0066] Therefore, pre-tuning of video streams using, for example, VP9 or AV1, generally involves aligning the arithmetic coding of the composite video stream to reduce the possibility of corruption in the resulting composite stream. One modification to the video stream encoding process to include in a composite video stream is to align the arithmetic coding used to represent the symbol frequencies of the symbols used to describe how the output is reconstructed in the decoder. In arithmetic decoding, probabilities are specified at the frame level, and typically the symbol frequency of one frame updates the probability of the next frame. Thus, subsequent frames can be transmitted using fewer bits.

[0067] When video streams are pre-tuned for inclusion in a composite video, a portion of each frame of the pre-tuned video is rendered as a "null" or black. Consequently, when encoding the pre-tuned video stream, a portion of the frame (e.g., half) will be encoded as a "null" or black. If standard encoding techniques are used, the resulting composite video stream (replacing null or black tiles with alternative PiP video streams) will cause the decoder to check different symbol frequencies for the current frame, resulting in the use of incorrect probabilities when decoding the next frame of the composite video stream. In other words, since the symbols describing the frames of each video stream forming the composite video stream are individually compressed using arithmetic coding to facilitate transmission to the end device with fewer bits, the symbol frequencies used to update the probabilities between consecutive frames of the video streams should match (e.g., perfectly match) at the time of encoding the streams to ensure a match between the composite video streams (e.g., between the primary and secondary streams forming the composite video stream) so that the composite video can be properly decoded. This can be achieved, for example, by using error-resilient encoding or by explicitly signaling the probabilities.

[0068] In some embodiments, ensuring a match during the decoding of a composite video stream involves using a static (e.g., default) set of known values ​​in the standard to encode frames of a pre-coordinated video stream. For example, within the VP9 codec, the default over_under.webmprobabilities feature can be used when encoding each frame in error-resilient mode. By using this feature in a novel way (generally used for lossy, "error-prone" communication links), and by using the default probabilities for each frame (rather than updating them), it is possible to ensure that symbol probabilities match when a composite video stream containing two or more different video streams is decoded.

[0069] In some embodiments, for example, where an increased processing load on the local area server is justified by reducing bandwidth on the end-user device, ensuring a match during decoding of the composite video stream involves combining (e.g., concatenating) the symbol frequency probabilities of each composite frame so that the probability of each symbol frequency of each video stream contained in the composite video stream is included. For example, ensuring a match during decoding may involve undoing the arithmetic encoding of the two video streams and then re-arithmetic encoding the two videos together (i.e., updating the probabilities frame by frame). This approach may require significantly more processing on the local area server, but it may still be significantly cheaper than using conventional methods that require full encoding of the video streams.

[0070] Another modification to video stream encoding for inclusion in a composite video stream is to leverage the functionality of a frame segmentation map defined by the codec to ensure a match between the various quality and compression levels of the encoded video streams included in the encoded composite video stream. The functionality of the segmentation map can be implemented, for example, using segment identifiers (IDs), so that each video stream included in the composite video stream can utilize an independent encoding quality / compression level appropriate to the complexity of the content it contains. In such a case, the decoding process then functions to present a decoded composite video with the characteristics of each of the constituent video streams. For example, a segmentation map within the VP9 framework can be applied to a frame to divide that frame into 8x8 pixel block regions corresponding to differently included video streams, assigning different segment identifiers (IDs) to the pixel block regions, and the uncompressed header incorporates the value of one or more properties that adjust the decoding of that region in the composite video for each segment ID used. In other words, different segment IDs can be used to enforce an uncompressed header that allows for matching the values ​​of one or more properties of the decoding of the compressed video. For example, the properties of a segmentation map can be used to ensure a match between one or more of the following: quantization level (Q), loop filter intensity, reference frame, and skip mode between the source video stream and the composite video stream.

[0071] Figure 6 shows an example of a composite video frame containing multiple tiles. As shown, frame 600 is divided into four tiles A, B, C, and D. Tiles A and B contain the video content areas of a first video stream (e.g., a secondary stream), and tiles C and D contain the video content areas of a second video stream (e.g., a primary stream). The first video stream contained in tiles A and B uses a first segment ID for all pixels in tiles A and B during the encoding of the first video, and the second video uses a second segment ID for all pixels in tiles C and D during the encoding of the second video.

[0072] In some embodiments, separate segment IDs can be used for different tiles of a frame to ensure matching of the quantization level (Q) between each video stream contained in a composite video stream. In other words, each composite video in a composite video stream can be assigned an independent quantization level. Generally, encodings utilize several constraints on bitrate and quality to reduce the number of bits transmitted to a network or device, or to reduce the number of bits transmitted without significantly improving quality. For example, highly complex content results in lower quality (higher quantization level), and the simpler the content (e.g., up to a certain point), the higher the resulting quality (lower quantization level). Using a single quantization level for all content can result in either too many bits for complex content or poor quality for simple content.

[0073] Simply put, the server implements a Discrete Cosine Transform (DCT) to identify important or prominent features within the structure of the image captured by the frame, and then applies quantization coefficients to generate lossy compression of the frame. The data from the DCT output can be quantized, and consequently, a lower quantization level results in higher quality (higher fidelity to the original image), while a higher quantization level results in lower quality (lower fidelity to the original image).

[0074] In some embodiments, adjusting the quantization level can be used to ensure rate control that the encoding (e.g., compression) of video content is within the target bit range, thereby making the quality changes between different video content streams less noticeable to the viewing user. The encoder can apply separate Qs to the discrete cosine transform (DCT) of tiles containing the video content region of the video stream, from the DCT of tiles containing the "null" or black areas of the frame. For example, a lower quantization level can be used for tiles containing the video content region, and a higher quantization level can be used for tiles containing the null or black areas.

[0075] Typically, Q can be selected to be constant across the entire frame. However, to ensure Q consistency across the entire frame of a composite video stream, the encoder can utilize a frame segmentation map, allowing each tile in the frame to have a distinct segment ID during the encoding process. This enables the uncompressed header of the frame in the composite video stream to have separate Qs for different regions identified by the segment IDs. For example, when a local area server performs compressed domain synthesis of different video streams and generates a composite video stream by replacing one or more tiles in the primary stream with tiles from the secondary stream, the distinct segment IDs for the retained tiles in the primary stream and the inserted tiles in the secondary stream ensure that the Q of each tile in the composite stream matches its original value.

[0076] In some embodiments, segment IDs are used to enhance the pixel resolution of a region of interest (e.g., a feature), which is accompanied by a first segment ID having a first Q for base content (e.g., high Q) and a second segment ID having a second Q for smaller regions to address any quality issues (e.g., low Q). In some examples, segment IDs can be used to boost blocks of pixels that have high visibility to the user (e.g., objects or regions in video content that appear throughout many frames of playback).

[0077] In some embodiments, a separate segment ID can be used for each composite video in a composite video stream. For example, two segment IDs can be used for each composite video, in which case the first segment ID is used for the majority of the content of the composite video and uses the same Q value as the source video, while the second segment ID has the smallest possible Q value and is used when it is necessary to break down a block of motion vectors into smaller blocks (for example, in the case of hard motion vector correction described in Figure 10B). In this case, the residuals for the blocks used to correct any difference between the copied content and the desired content often need to be re-encoded into smaller blocks. If the original Q value is used, large errors can occur, so using a smaller Q value minimizes these errors.

[0078] Figure 7 shows a schematic diagram of example 700 reference frames between the primary stream, the secondary stream, and their combined video stream. Within the limits of the VP9 codec, motion vectors within a frame can reference up to three different frames from a set of eight reference buffers. During decoding of the combined video stream, the reference frames used in the primary and secondary video streams, respectively, must match so that the reference frames used in the combined video frame are the same as those in the primary and secondary streams.

[0079] In some embodiments, after the frame has been decoded and reconstructed for presentation, one or more post-processing steps are performed on the user device. One such post-processing step includes a loop / deblocking filter, which is applied to the frame reconstructed by the user device to smooth the edges across adjacent pixel blocks. Edge artifacts between adjacent pixel blocks may arise due to the independently performed DCT encoding of each block.

[0080] In some embodiments, a gap is included between the video streams contained in the composite stream to prevent the loop filter from smoothing between the two different streams that make up the composite stream. Figure 8 shows a schematic diagram of an example of a gap inserted between two video streams in a composite video stream. The gap 802 located between the primary stream 804 and the secondary stream 806 is at least a threshold number of pixels wide to prevent smoothing / mixing, for example, separated by at least 16 pixels from the edge of the video content area of ​​the primary stream to the edge of the video content area of ​​the secondary stream. In some cases, each portion of the gap 802 between the primary stream 804 and the secondary stream 806 is inserted into each of the primary stream 804 and the secondary stream 806 during their respective pre-adjustments. For example, 8 pixels are each added to the edge of the video content area of ​​a video stream adjacent to the other video content area of ​​the other video stream when the two video streams are contained in the composite video stream.

[0081] Generally, a significant portion of the compression of video content during the encoding process is the result of using motion vectors to track the actual changes between consecutive frames, and using residuals (e.g., represented as DCT) to encode the difference between the actual frame and the predicted frame. Motion vectors can be represented with different numbers of bits and may have high or low precision, but they must match across frames to ensure that the reconstructed synthetic video stream reflects the content of the original primary and secondary streams. In some embodiments, the motion vector precision can be set to a default value during the encoding process of the video streams included in the synthetic video stream, for example, a value of 1.

[0082] In some embodiments, due to limitations of the hardware encoder (e.g., lack of keep-out regions), the server may pull content from outside the frames of a video stream into the frames during the encoding process. As a result, during the decoding process of the composite video stream by the user device, the decoder may use references from outside the frames, and those frames may be occupied by video content regions of different video streams in the composite video stream after synthesis. Vectors that reference regions during encoding, which may cause decoding problems, can be updated as described with reference to Figures 9A, 9B, 9C and 10A, 10B.

[0083] Figures 9A and 9C illustrate examples of motion vector containment solutions. As shown in Figure 9A, motion vector 900 references a “null” or black area 902 from outside the video content area 904 of frame 906. For example, as shown in Figure 9B, the reference area 902 is located within an area occupied by other video content areas 908 when the composite video stream frame is generated. As shown in Figure 9C, motion vector 900 moves to an alternative “safe” area 910 above and below the “null” or black content of the first video content area 904. In some embodiments, all vectors moved from black content that has moved to the “safe” area can copy content from the same source in the previous frame (i.e., they can reference the same safe block).

[0084] Figures 10A and 10B illustrate another example of a motion vector containment solution. As shown in Figure 10A, the motion vector 1000 refers to a region 1002 located between two video content regions 1001 and 1003 of the composite video stream 1005. As shown in Figure 10B, the referenced region 1002 is divided into subblocks 1004a and 1004b, for example, using partitioning. Subblock 1004a contains only the retained content (of the video content region in question) or black content between the two video content regions. Subblock 1004b contains the area occupied by the other video content region (and possibly black content as well) and moves to a “safe” region, as described with reference to Figures 9A and 9C. Separate motion vectors 1006a and 1006b are generated for subblocks 1004a and 1004b, respectively, and updated residuals are generated for subblocks 1004a and 1004b, respectively. Generating updated residuals involves decoding them as the size of subblocks 1004a and 1004b (e.g., using inverse DCT) and re-encoding them (e.g., performing DCT again), where the re-encoded values ​​are substantially the same as the original residuals of region 1002. For example, if the original region has 32x32 pixels, each subblock may have a size of four 16x16 pixels.

[0085] In some embodiments, to ensure at least a threshold match between the original residual value of the subblock and the re-encoded residual value of the subblock, a separate segment ID can be used for the subblock, and Q can be set to a very low value (compared to the rest of the tile), resulting in improved resolution, quality, and fidelity compared to the original encoded content. In other words, blocks with problematic motion vectors can be replaced with "intra-coded" blocks (including segment IDs with low Q values).

[0086] In some embodiments, for efficiency, the server may encode only half of a frame containing the video content area (e.g., content from a video stream), padding the remaining frame with black (e.g., filling it) to produce a full frame. Often, motion vectors are encoded against existing motion vectors from the current or previous frame. As a result of performing this process, motion vectors can be encoded against other reference motion vectors extending from the frame's edge (for new blocks). If a reference motion vector is too far from the image's edge (e.g., >32 pixels), it is clipped, and the clipped version is used as the base for a new motion vector to encode the difference of the motion vector from the reference motion vector (if any). However, when a frame with reduced width is expanded to full width in a compressed domain composition, some of these previously clipped reference motion vectors may be declipped, potentially causing motion vectors that depend on them to shift (because they no longer extend beyond the frame's edge). To solve this problem, the server can perform a declipsing process on the frame's motion vectors during the pre-adjustment process, where an arbitrary declipsed base motion vector is identified, and the difference of the dependent motion vectors is updated based on the new declipsed base motion vector.

[0087] Figure 11 is a flowchart of an exemplary process 1100 for compressed domain synthesis video streaming. For convenience, the process 1100 is described as being performed by one or more computer systems located in one or more locations and appropriately programmed according to this specification. For example, a digital component distribution system, e.g., the digital component distribution system 110 in Figure 1 (e.g., including server 200 and local area server 202 in Figure 2), can perform the process 1100.

[0088] The system (e.g., local area server 202) acquires a compressed video stream, each compressed video stream containing multiple frames (1102). Each frame of the compressed video contains a video content region occupying any number of consecutive tiles / slice. Each frame is encoded with a first segment identifier for pixels contained within the video content region and a second segment identifier for pixels not contained within the video content region. The frames of the compressed video are encoded using a set of static symbol probabilities.

[0089] In some embodiments, the compressed video stream is encoded by a server (e.g., server 200) that communicates data with the local area server 202. The compressed video stream may be generated from a video stream obtained, for example, from a third-party content generator. Generating a compressed video stream may involve the server rendering a new frame for each frame of the video stream, the new frame containing a video content area, the video content area being any number of consecutive tiles within the frame, corresponding to the content of the original frame. The video content area may contain one or more tiles of a tile-based composite of frames, for example, as shown in Figure 3A.

[0090] The new frame may also contain black or "null" areas outside the video content area. The frame can then be encoded (for example, using the VP9 codec), where the symbol frequency probabilities are set using default values. In other words, the probabilities are not updated between frames (for example, using error-resilient mode in the VP9 framework).

[0091] In some embodiments, generating compressed video further includes generating a frame segmentation map and defining a first segment identifier for pixels that are included in the video content area of ​​a new frame and a second segment identifier (i.e., black areas or "null" areas) for pixels that are not included in the video content area.

[0092] In some embodiments, the system can assign one or more parameter values ​​(e.g., quantization level (Q), reference frame, loop filter value, motion vector value, etc.) to each of the segment identifiers. For example, a first segment identifier can be assigned a first Q, and a second segment identifier can be assigned a second different Q.

[0093] In some embodiments, generating compressed video further includes rendering gaps (e.g., as described in Figure 8) along the edges of the video content regions within the video content regions of the compressed video stream. The gaps may have a width of at least eight pixels, and when the compressed video is included in a composite video stream, the null regions of the video content regions of the compressed video are adjacent to other video content regions of the other compressed video. The edges on which the null regions are encoded during the encoding process may depend in part on whether the compressed video is included in the composite video stream as a primary or secondary stream. For example, the null regions may be included to the right of the video content region of the secondary stream (e.g., 806) and to the left of the video content region of the primary stream (e.g., 804).

[0094] In some embodiments, generating a compressed video stream involves generating a primary video stream, e.g., pre-arranged and stored video content, and generating at least one secondary video stream (e.g., live stream content). For example, multiple video content streams can be pre-arranged and stored for inclusion in a composite video stream containing live stream content, and different users can receive different composites of the primary video streams along with the same live stream video.

[0095] In some embodiments, the local area server 202 receives a second set of compressed video streams, where more than 50% of the frames, for example, all or nearly all of the frames, are rendered video streams. In other words, the local area server 202 can receive encoded public live stream video content from server 200, which can be provided to user devices in response to requests for uncomposite video stream content.

[0096] The system (e.g., local area server 202) receives requests for video stream content from a user device (1104). A request (e.g., request 108) may include a request for live stream video content (e.g., a secondary stream) and user information that can be used to select a primary stream to include in a composite video stream with the secondary stream for presentation on the user device. The local area server 202 can receive multiple requests from different user devices and use compressed domain synthesis to generate each unique composite video stream for presentation on each requesting user device, partially based on the user information.

[0097] In some embodiments, the local area server 202 can pre-cached a set of primary streams, anticipating that it will generate a composite video stream(s) with one or more secondary streams (e.g., live streams).

[0098] The system, via a local area server, generates a compressed domain composite video stream from a first and second compressed video stream of multiple compressed video streams in response to a request from a user device (1106). Each frame of the compressed domain composite video stream includes a first video content region from the first compressed stream and a second video content region from the second compressed stream, where the first and second video content regions occupy non-overlapping regions of frames corresponding to complete tiles / slice in the encoded video.

[0099] In some embodiments, the system generates a composite video stream by replacing one or more tiles of the primary stream with one or more tiles of the secondary stream. For example, a null tile in the primary stream may be replaced with a video content area tile in the secondary stream. The system further generates an updated uncompressed header from the uncompressed header information for each of the composite videos. For example, the system incorporates the segment identifiers of the video content area in the primary stream and the segment identifiers of the video content area in the secondary stream into the uncompressed header of the composite video stream using tile replacement and while within the compression domain.

[0100] The system provides the user device with a packet containing a set of frames of a compressed domain-composite video stream that can be decoded by a single decoder (1108).

[0101] In some embodiments, a user request for video content (e.g., for live stream content) includes a request for a non-composite video stream. In other words, the user request may only want to view a secondary stream. In such cases, the local area server may provide a packet containing a set of frames of a third video stream that contains only live stream content decodeable by a single decoder.

[0102] Figure 12 is a block diagram of an exemplary computer system 1200 that can be used to perform the operations described above. System 1200 includes a processor 1210, memory 1220, storage device 1230, and input / output device 1240. Each of components 1210, 1220, 1230, and 1240 can be interconnected using, for example, a system bus 1250. Processor 1210 is capable of processing instructions to be executed within system 1200. In one embodiment, processor 1210 is a single-threaded processor. In other embodiments, processor 1210 is a multi-threaded processor. Processor 1210 is capable of processing instructions stored in memory 1220 or on storage device 1230.

[0103] The memory 1220 stores information within the system 1200. In one embodiment, the memory 1220 is a computer-readable medium. In one embodiment, the memory 1220 is a volatile memory unit. In other embodiments, the memory 1220 is a non-volatile memory unit.

[0104] The storage device 1230 can provide mass storage to the system 1200. In one embodiment, the storage device 1230 is a computer-readable medium. In various different embodiments, the storage device 1230 may include, for example, a hard disk device, an optical disk device, a storage device shared over a network by multiple computing devices (e.g., a cloud storage device), or any other mass storage device.

[0105] The input / output device 1240 provides input / output operation to the system 1200. In one embodiment, the input / output device 1240 may include one or more network interface devices, such as an Ethernet card, a serial communication device (e.g., an RS-232 port), and / or wireless interface devices (e.g., an 802.11 card). In other embodiments, the input / output device may include a driver device configured to receive input data and transmit output data to other devices, such as a keyboard, printer, display, and other peripheral devices 1260. However, other embodiments such as mobile computing devices, mobile communication devices, and set-top box television client devices may also be used.

[0106] While Figure 12 illustrates an exemplary processing system, the subjective and functional embodiments of operation described herein can be implemented in other types of digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or any combination thereof.

[0107] The subject matter, actions, and operations described herein can be implemented in digital electronic circuits, tangibly embodied computer software or firmware, computer hardware, or a combination thereof, including structures disclosed herein and their structural equivalents. The subject matter, actions, and operations described herein can be implemented as, or in, one or more computer programs (e.g., one or more modules of computer program instructions encoded on a computer program carrier for execution by a data processing device or for controlling the operation of a data processing device). The carrier may be a tangible, non-transient computer storage medium. Alternatively or additionally, the carrier may be an artificially generated propagating signal, such as a machine-generated electrical signal, optical signal, or electromagnetic signal, generated to encode information for transmission to a suitable receiving device for execution by a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable memory board, a random-access memory device, or a serial-access memory device, or a combination thereof, or a part thereof. The computer storage medium is not a propagating signal.

[0108] The term "data processing device" encompasses all kinds of devices, machines, and equipment for processing data, including, for example, programmable processors, computers, or multiple processors or computers. A data processing device may include dedicated logic circuits, such as FPGAs (Field-Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), or GPUs (Graphics Processing Units). The device may also optionally include, in addition to hardware, code that generates an execution environment for computer programs (e.g., code that constitutes processor firmware, protocol stacks, database management systems, operating systems, or a combination of one or more of these).

[0109] Computer programs can be written in any form of programming language, including compiled languages, interpreted languages, declarative languages, or procedural languages. Computer programs can also be deployed in any form, such as standalone programs, for example, as apps, or as modules, components, engines, subroutines, or other units suitable for execution in a computing environment, which may include one or more computers interconnected by a data communication network in one or more locations.

[0110] Computer programs may, but do not necessarily, correspond to files in a file system. They can be stored in a single file dedicated to the program itself, or in multiple coordinate files (e.g., files storing one or more modules, subprograms, or parts of code), as part of a file that also holds other programs or data (e.g., one or more scripts stored in a markup language document).

[0111] The processes and logic flows described herein can be performed by one or more computers running one or more computer programs, by manipulating input data to produce outputs. Furthermore, the processes and logic flows can be performed by dedicated logic circuits (e.g., FPGAs, ASICs, or GPUs), or by a combination of dedicated logic circuits and one or more programmed computers.

[0112] A computer suitable for running computer programs may be based on a general-purpose microprocessor, a dedicated microprocessor, both, or any other type of central processing unit. Generally, the central processing unit receives instructions and data from read-only memory, random-access memory, or both. The basic components of a computer are the central processing unit for executing instructions and one or more memory devices for storing instructions and data. The central processing unit and memory can be complemented by or incorporated into dedicated logic circuits.

[0113] Generally, a computer is also configured to include, or be operably coupled to, one or more mass storage devices, and to receive or transfer data to or from a mass storage device. Mass storage devices can be, for example, magnetic, magneto-optical, or optical discs, or solid-state drives. However, a computer is not required to have such devices. Furthermore, a computer can be integrated into other devices, to name just a few, such as mobile phones, personal digital assistants (PDAs), mobile audio or video players, game consoles, Global Positioning System (GPS) receivers, or portable storage devices, such as Universal Serial Bus (USB) flash drives.

[0114] To provide user interaction, the subject matter described herein can be implemented on one or more computers, which have, or can be configured to communicate with, display devices for displaying information to the user (e.g., an LCD monitor, or a virtual reality (VR) display or an augmented reality (AR) display) and input devices that allow the user to input into the computer (e.g., a keyboard and a pointing device, e.g., a mouse, trackball, or touchpad). Similarly, user interaction can be provided using other types of devices. For example, feedback and responses provided to the user may be any form of sensory feedback (e.g., visual, auditory, vocal, or tactile), and input from the user may be received in any form, including acoustic input, voice input, or haptic input, including touch actions or touch gestures, or motor actions or motor gestures, or directional actions or directional gestures. Furthermore, the computer can interact with the user by sending documents to and receiving documents from the devices used by the user. For example, a computer can interact with a user by sending a web page to a web browser on the user's device in response to a request received from the web browser, or by interacting with an application running on the user's device (e.g., a smartphone or electronic tablet). A computer can also interact with a user by sending a text message or another form of message to a personal device (e.g., a smartphone running a messaging application) and receiving a response message from the user in return.

[0115] This specification uses the term “configured to” in relation to systems, devices, and components of computer programs. When one or more computer systems are configured to perform a particular operation or action, it means that software, firmware, hardware, or a combination thereof is installed on the system that causes the system to perform that operation or action during operation. When one or more computer programs are configured to perform a particular operation or action, it means that one or more programs, when executed by a data processing device, contain instructions that cause the device to perform that operation or action. When a dedicated logic circuit is configured to perform a particular operation or action, it means that the circuit has the electronic logic to perform that operation or action.

[0116] The subject matter described herein can be implemented in a computing system which includes a backend component (e.g., functioning as a data server), or a middleware component (e.g., an application server), or a frontend component (e.g., a client computer having a graphical user interface, a web browser, or an application that allows a user to interact with embodiments of the subject matter described herein), or one or any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication (e.g., a communication network) in any form or medium. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.

[0117] A computing system may include clients and servers. Clients and servers are generally remote to each other and typically interact via a communication network. The relationship between the client and server arises from computer programs running on each computer that have a client-server relationship with each other. In some embodiments, the server transmits data, such as an HTML page, to a user device for the purpose of displaying data to the user interacting with the device acting as a client and receiving user input from that user. Data generated on the user device, such as the results of user interactions, can be received by the server from the device.

[0118] While this specification includes details of many specific embodiments, these should not be interpreted as limitations on the scope of what is claimed as defined by the claims themselves, but rather as descriptions of features that may be specific to a particular embodiment of a particular invention. Certain features described herein in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can be implemented separately or in multiple embodiments in any suitable secondary combination. Furthermore, even if features are described above as functioning in a particular combination and were initially claimed as such, one or more features from the claimed combination may be removed from the combination, and the claims may cover secondary combinations or variations of secondary combinations.

[0119] Similarly, while operations are shown in the drawings and described in a specific order in the claims, this should not be understood as requiring that such operations be performed in a specific or sequential order, or that all operations be performed, in order to obtain the desired results. In certain circumstances, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together into a single software product or packaged into multiple software products.

[0120] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions described in the claims may be performed in a different order to achieve more desirable results. As an example, the process shown in the accompanying diagram does not necessarily require to be performed in a specific or sequential order to obtain the desired results. In some cases, multitasking and parallel processing may be advantageous.

Claims

1. A method performed by a computer, The local area server obtains multiple compressed video streams, each compressed video stream containing multiple frames, and each frame containing a video content area that includes a portion of the frame. Each frame is encoded using the segment identifier of the pixels contained in the video content area. The aforementioned multiple frames are encoded using a set of static symbol frequencies, and are obtained. The local area server receives requests for video stream content from user devices, In order to obtain a compressed domain composite video stream, the local area server, in response to the request from the user device, composites a first compressed video stream and a second compressed video stream of the plurality of compressed video streams, Each frame of the compressed domain composite video stream includes a first video content region of the first compressed video stream and a second video content region of the second compressed video stream. Each frame includes a first segment identifier for each pixel corresponding to the first video content area and a second segment identifier for each pixel corresponding to the second video content area. The first video content area and the second video content area occupy the non-overlapping portion of the frame, and are combined. To provide the user device with a packet containing a set of frames of the compressed domain composite video stream that can be decoded by a single decoder, Methods that include...

2. Obtaining the aforementioned multiple compressed video streams is, The server receives multiple video streams, The process involves generating the aforementioned plurality of compressed video streams, and for each of the compressed video streams, For each of the plurality of frames of the video stream, a new frame is rendered that includes a portion of the frame and includes a video content area corresponding to the content of the frame. With respect to the video content area, define the segment identifier of the pixels included in the video content area, Encoding the frame using a set of static symbol frequencies, The method according to claim 1, comprising generating, including

3. The method according to claim 2, wherein the rendering further includes rendering areas not included in the video content area as null content.

4. The generation of the aforementioned multiple compressed video streams is To generate a primary video stream containing context-responsive video content, This includes generating at least one secondary video stream containing live stream content, The method according to claim 2, wherein the first compressed video stream includes a primary video stream, and the second compressed video stream includes a secondary video stream.

5. The method according to claim 4, wherein generating the primary video stream includes pre-caching and storing context-responsive video in a database for inclusion in the compressed domain composite video stream, and a suitable subset of the context-responsive video content is pre-cached on the local area server.

6. Obtaining the aforementioned compressed domain composite video stream is: The local area server includes selecting primary content videos from a database of primary video streams to include in the compressed domain composite video stream in response to user-based selection criteria of the user device. The method according to any one of claims 1 to 5, wherein, in response to different user requests, the local area server selects different primary content videos to be included in the compressed domain composite video stream.

7. The generation of the aforementioned multiple compressed video streams is This includes generating a video stream that includes at least one of context-responsive video content and live stream content, The method according to claim 2, wherein the first compressed video stream and the second compressed video stream are generated from the video stream.

8. Encoding the frame using a set of static symbol frequencies is The method according to claim 2, comprising encoding each symbol probability of the frame with a default frequency value.

9. With respect to the aforementioned video content area, defining the segment identifier means that The method according to claim 2, further comprising assigning a first quantization level to the segment identifier of a pixel included in the video content area.

10. Rendering is Further comprising rendering a null region, including its width, within the video content region of the compressed video stream, along the edge of the video content region, The method according to claim 2, wherein the null region of the video content region is located next to a replacement region configured to be replaced by another compressed video stream in the composite video stream.

11. The process of generating the compressed domain composite video stream from the first compressed stream and the second compressed stream is as follows: The method according to claim 1, comprising constructing an uncompressed header for each frame of the compressed domain composite video stream based on the segment identifiers of the first and second compressed streams, respectively.

12. The method according to claim 1, wherein the frame includes tiles, the video content area includes one or more consecutive tiles of a tile-based composite of the frame, and the pixels included in the video content area include one or more consecutive tiles.

13. The method according to claim 12, wherein the one or more consecutive tiles including the video content area include half of the tiles in the tile-based composite of the frame.

14. The local area server acquires a second plurality of compressed video streams, each of which contains a plurality of frames. The local area server receives a second request for video stream content from the user device, In response to the second request, the local area server provides the user device with a second packet containing a set of frames for one of the second plurality of compressed video streams, The method according to any one of claims 1 to 5, further comprising:

15. One or more computer storage media encoded with computer program instructions, wherein when the computer program instructions are executed by one or more computers, the one or more computers will receive The local area server obtains multiple compressed video streams, each compressed video stream containing multiple frames, and each frame containing a video content area that includes a portion of the frame. Each frame is encoded using the segment identifier of the pixels contained in the video content area. The aforementioned multiple frames are encoded using a set of static symbol frequencies, and are obtained. The local area server receives requests for video stream content from user devices, In order to obtain a compressed domain composite video stream, the local area server, in response to the request from the user device, composites a first compressed video stream and a second compressed video stream of the plurality of compressed video streams, Each frame of the compressed domain composite video stream includes a first video content region of the first compressed video stream and a second video content region of the second compressed video stream. Each frame includes a first segment identifier for each pixel corresponding to the first video content area and a second segment identifier for each pixel corresponding to the second video content area. The first video content area and the second video content area occupy the non-overlapping portion of the frame, and are combined. To provide the user device with a packet containing a set of frames of the compressed domain composite video stream that can be decoded by a single decoder, One or more computer storage media that perform operations including those mentioned above.

16. Obtaining the aforementioned multiple compressed video streams is, The server receives multiple video streams, The process involves generating the aforementioned plurality of compressed video streams, and for each of the compressed video streams, For each of the plurality of frames of the video stream, a new frame is rendered that includes a portion of the frame and includes a video content area corresponding to the content of the frame. With respect to the video content area, define the segment identifier of the pixels included in the video content area, Encoding the frame using a set of static symbol frequencies, A computer storage medium according to claim 15, comprising generating, including

17. The generation of the aforementioned multiple compressed video streams is To generate a primary video stream containing context-responsive video content, This includes generating at least one secondary video stream containing live stream content, The computer storage medium according to claim 15 or 16, wherein the first compressed video stream includes a primary video stream, and the second compressed video stream includes a secondary video stream.

18. The computer storage medium according to claim 17, wherein generating the primary video stream includes pre-caching and storing context-responsive video in a database for inclusion in the compressed domain composite video stream, and a suitable subset of the context-responsive video content is pre-cached on the local area server.

19. Obtaining the aforementioned compressed domain composite video stream is: The local area server includes selecting primary content videos from the primary video stream database to include in the compressed domain composite video stream in response to user-based selection criteria of the user device. The computer storage medium according to claim 18, wherein, in response to different user requests, the local area server selects different primary content videos to be included in the compressed domain composite video stream.

20. A system comprising one or more computers and one or more storage devices that store operable instructions, wherein when an instruction is executed by one or more computers, the one or more computers... The process involves obtaining multiple compressed video streams, each compressed video stream containing multiple frames, and each frame containing a video content region that includes a portion of the frame. Each frame is encoded using the segment identifier of the pixels contained in the video content area. The aforementioned multiple frames are encoded using a set of static symbol frequencies, and are obtained. Receiving requests for video stream content from user devices, In order to obtain a compressed domain composite video stream, the first compressed video stream and the second compressed video stream of the plurality of compressed video streams are composited in response to the request from the user device, Each frame of the compressed domain composite video stream includes a first video content region of the first compressed video stream and a second video content region of the second compressed video stream. Each frame includes a first segment identifier for each pixel corresponding to the first video content area and a second segment identifier for each pixel corresponding to the second video content area. The first video content area and the second video content area occupy the non-overlapping portion of the frame, and are combined. To provide the user device with a packet containing a set of frames of the compressed domain composite video stream that can be decoded by a single decoder, A system that performs actions including those mentioned above.