Video track switching technology during media title playback
A comprehensive media package with metadata-driven track switching addresses redundant processing and limited presentation issues, reducing resource consumption and enhancing user-tailored playback experiences.
Patent Information
- Application Number
- JP2025544961
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-02-15
- Filing Date
- 2024-02-08
- Publication Date
- 2026-02-20
AI Technical Summary
Existing media streaming technologies redundantly process audio and timed text source files for each video-only package, leading to unnecessary storage and processing resource consumption, and limit the number of recommended presentations available to users based on a subset of stream sets, resulting in subpar viewing experiences when user preferences differ.
Generate a single comprehensive media package that includes all video, audio, and timed text streams from all source files, with metadata describing stream characteristics and compatibility, allowing instant selection based on user preferences for seamless track switching during playback.
Reduces storage and processing resources by avoiding redundant processing, and enables more tailored playback experiences by allowing instant selection of tracks from the entire set of streams, enhancing user experience.
Smart Images

Figure 2026505985000001_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority to U.S. Patent Application No. 18 / 169,783, filed February 15, 2023, which is incorporated herein by reference. [Technical Field]
[0002] Various embodiments relate generally to computer science and media streaming technologies, and more particularly to video track switching during playback of a media title. [Background technology]
[0003] In a process known as "video localization," an original video source file associated with a given media title is modified to generate a localized video source file that meets the individual preferences of a target audience. For example, English text appearing in the original video source file may be replaced with similarly meaning French text to generate a localized video source file that better meets the preferences of a French-speaking target audience. In another example, hockey scenes in the original video source file may be replaced with soccer scenes to reflect the sports preferences of a target audience living in Brazil.
[0004] A typical media title is streamed and played through one or more presentations, each of which may contain video tracks, audio tracks, and optional timed text tracks, each derived from a different video source file, a different audio source file, and a different optional timed text source file. Each track specifies either a single stream or multiple streams, usually corresponding to different bit rates, where a given stream constitutes an encoded version of the source file.
[0005] To enable a client device to stream and play different presentations of a media title, a media processing pipeline is used to independently generate a different video-only package for each video source file associated with the media file. To generate a video-only package for a given video source file, one or more target presentations are defined based on target audience preferences and compatibility with the corresponding video source file. Audio source files and any timed text source files necessary to generate the target presentation are identified. For each single video source file, one or more identified audio source files, and zero or more identified timed text source files, the media processing pipeline generates a different set of one or more corresponding streams, referred to as a "stream set." The media processing pipeline then generates a video-only package containing the stream set, stores the video-only package on an origin server device, and then deploys the video-only package from the origin server device to a content delivery network.
[0006] To play a media title, an endpoint application running on a client device requests a manifest file for the media title from the cloud-based media service. In response to the request for the manifest file, the manifest application selects an edge server in the content delivery network that is close to the client device. The manifest application then selects a video-only package for the media title based on one or more characteristics associated with the user (e.g., profile language). The profile language for the user specifies the language of the user interface that the streaming media provider uses to interact with the user. The manifest application then generates at least one recommended presentation based on a single set of video streams, a set of audio streams, and an optional set of timed text streams included in the selected video-only package. Each recommended presentation includes the same video track that specifies a subset of the set of video streams, and a different combination of audio tracks that specify a subset of the set of audio streams and optional timed text tracks that specify a subset of the set of timed text streams. The manifest application generates and sends the manifest file to the endpoint application, which enables the endpoint application to determine an edge server and select tracks for any of the recommended presentations. To play each portion or "segment" of a media title, the endpoint application selects a corresponding segment from each selected track stream and requests the selected segment from the selected edge server device. In some use cases, the endpoint application may switch between two or more recommended presentations during playback of a media title. Summary of the Invention [Problem to be solved by the invention]
[0007] One drawback of the above approach is that each audio source file and each timed text source file used to generate more than one video-only package is processed redundantly by the media processing pipeline. As a result, unnecessary storage resources, processing resources, and time are consumed to generate and deploy the different video-only packages to the content delivery network. For example, if a Spanish language audio source file is included in five different video-only packages using the above approach, the Spanish language audio source file would then be encoded five times at one or more different bit rates to generate five separate audio stream sets, all of which may contain the same audio stream. Furthermore, the five audio stream sets would be evaluated separately for sound quality and deployed separately to the content delivery network. In such a scenario, unnecessary storage resources, processing resources, and time would be consumed to process and deploy the audio streams associated with four of the five video-only packages.
[0008] Another drawback of the above approach is that it unnecessarily limits the number of different recommended presentations of a media title available to each user. In this regard, the recommended presentations specified in a given manifest file are generated based on a subset of the stream sets associated with the media title actually included in a single video-only package. In other words, the recommended presentations are generated based on the single video stream set, one or more audio stream sets, and zero or more timed text stream sets included in the selected video-only package. Therefore, if a given user's preferences differ from the preferences of the target audience associated with the selected video-only package, the given user's viewing experience may be subpar. For example, a video-only package for a media title may include a Spanish video stream set, an English original audio stream set, and a Spanish dubbed audio stream set, but omit a Polish dubbed audio stream set. A manifest file generated for a user associated with a profile language of Spanish and an audio preference of Polish may include one recommended presentation that identifies both Spanish localized video tracks and Spanish dubbed audio tracks, and another recommended presentation that identifies both Spanish localized video tracks and English original audio tracks. The manifest file would likely degrade the user's viewing experience because it would not enable the user to access the user's preferred set of Polish dubbed audio streams.
[0009] As can be seen, there is a need in the art for more effective techniques for tailoring media viewing experiences to user preferences. [Means for solving the problem]
[0010] One embodiment describes a computer-implemented method for switching video tracks during playback of a media title, the method including: selecting one or more video streams for inclusion in a first video track, the one or more video streams derived from a first video source file associated with the media title and included in a media package generated for the media title; generating a manifest file based on the first video track, the manifest file itemizing the first video track and indicating that a first alternate video track may be made available; transmitting the manifest file to a client device; receiving a request from the client device to generate an alternate manifest file itemizing the first alternate video track; selecting one or more video streams for inclusion in the first alternate video track, the one or more video streams derived from a second video source file associated with the media title and included in the media package; generating the alternate manifest file based on the first alternate video track; and transmitting the alternate manifest file to the client device.
[0011] At least one technical advantage of the disclosed technology over the prior art is that, when using the disclosed technology, source files are not processed redundantly by a media processing pipeline compared to a conventional video-only package when a single, all-encompassing media package that more comprehensively represents a media title is generated and deployed to a content delivery network. Therefore, the amount of storage resources, processing resources, and time consumed to generate and deploy a media title may be significantly reduced compared to prior art approaches. Another technical advantage of the disclosed technology is that, unlike prior art approaches, identified tracks in a manifest file generated using a single all-encompassing media package may be instantly selected based on predetermined user preferences from the entire set of streams derived from all source files associated with the corresponding media title. Therefore, the disclosed technology can generate tracks for playback that are more closely tailored to a user's preferences than may be possible using prior art approaches in which a manifest file is generated based on a subset of the streams included in a single video-only package. These technical advantages provide one or more technical advances over prior art approaches.
[0012] In order that the features of the various embodiments may be appreciated in detail, a briefly summarized inventive concept will now be described with reference to various embodiments, some of which are illustrated in the drawings. However, it should be noted that the attached drawings illustrate only exemplary embodiments of the inventive concept and should not be considered in any way as limiting the scope of the invention, and further that there may be other equally effective embodiments. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is a conceptual diagram of a system configured to implement one or more aspects of various embodiments. [Figure 2] 2 is a more detailed diagram of the comprehensive media package of FIG. 1 in accordance with various embodiments. [Figure 3]2 is a more detailed diagram of the manifest file of FIG. 1 in accordance with various embodiments. [Figure 4] 2 is a more detailed diagram of an alternative manifest file of FIG. 1 in accordance with various embodiments. [Figure 5] 1 is a flow diagram of method steps according to various embodiments for generating and deploying streams associated with media titles to a content delivery network. [Figure 6] 1 is a flow diagram of method steps according to various embodiments for generating a manifest file that itemizes streams associated with a media title. [Figure 7] 1 is a flow diagram of method steps according to various embodiments for generating a manifest file that enables video track switching during playback of a media title. DETAILED DESCRIPTION OF THE INVENTION
[0014] In the following description, numerous specific details are set forth to provide a more thorough understanding of various embodiments. However, it will be apparent to one skilled in the art that the inventive concepts may be practiced without these specific details.
[0015] In video localization, an original video source file is modified to generate one or more localized video source files that better suit a corresponding target audience. For example, lip re-animation techniques are used to modify the lip movements of speakers in an English-language movie to match those in 30 language dubs, generating 30 localized versions of the movie. To enable client devices to stream and play a media title, a typical media streaming service independently generates and delivers a different video-specific package for each different video source file associated with the media title.
[0016] To generate a video-only package for a given video source file associated with a given media title, a subset of audio source files associated with the media title and a subset of timed text source files associated with the media title are selected based on compatibility with the target audience and the video source file. A set of video streams is generated based on the video source file, a different set of audio streams is generated for each selected audio source file, and a different set of timed text streams is generated for each selected timed text source file. The set of video streams, the set of audio streams, and the set of timed text streams are assembled into a video-only package. The video-only package is deployed on a content delivery network that provides the streams on demand to client devices.
[0017] To play a media title, an endpoint application running on a client device requests a manifest file for the media title from the cloud-based media service. The cloud-based media service then selects a video-only package for the media title based on one or more characteristics associated with the user (e.g., profile language). The cloud-based media service generates one or more recommended presentations based on the selected video-only package. Each recommended presentation includes the same video track that identifies a subset of the set of video streams, and a different combination of audio tracks that identify a subset of the set of audio streams and optional timed text tracks that identify a subset of the set of timed text streams. The cloud-based media service generates and sends the manifest file to the endpoint application, which enables the endpoint application to stream the media title from a content delivery network and play it using any recommended presentation.
[0018] One drawback of the above approach is that each audio source file and each timed text source file used to generate more than one video-only package is processed redundantly by the media processing pipeline. As a result, unnecessary amounts of storage resources, processing resources, and time are consumed to generate and deploy different video-only packages to a CDN. Another drawback of the above approach is that the recommended presentations specified in a given manifest file are limited to a single video track and combinations of other tracks derived from the set of streams actually included in the single selected video package. If a user's preferences favor a set of streams omitted from the single selected video package, then the quality of the viewing experience for that user may be subpar.
[0019] For example, suppose 20 video-only packages are each generated based on a different video source file, three audio source files selected from a total of 20 different audio source files, and nine timed text source files selected from a total of 60 different timed text source files. Generating and deploying 40 redundant audio stream sets and 120 redundant timed text stream sets may consume unnecessary amounts of storage resources, processing resources, and time. Furthermore, if a user's preferences favor one of the 17 different audio stream sets and / or one of the 51 different timed text stream sets omitted from the single selected video package, the viewing experience for that user may be degraded.
[0020] However, using the techniques of this disclosure, the media processing pipeline generates a single comprehensive media package for a media title based on all video, audio, and timed text source files associated with the media title. The comprehensive media package includes all sets of video, audio, and timed text streams derived from all source files associated with the media title. The media processing pipeline also generates metadata that describes relevant characteristics of the associated stream sets, identifies the stream sets, and defines compatibility between them. The media processing pipeline deploys the comprehensive media package to a content delivery network that provides the streams on demand to client devices. The media processing pipeline stores the metadata associated with the comprehensive media package in cloud-based memory available to a manifest customization application included in the cloud-based media service.
[0021] Upon receiving a request for a manifest file from an endpoint application, the manifest customization application uses metadata associated with the comprehensive media package to select one of a set of video streams in the comprehensive media package based on associated user preferences. The manifest customization application generates one or more recommended presentations based on video tracks derived from the selected video streams and the user preferences. Each recommended presentation identifies a different combination of the same video track, audio track, and optional timed text track. The manifest customization application generates a manifest file that describes each recommended presentation, the video track, each audio track included in at least one recommended presentation, and each timed text track included in at least one recommended presentation, and identifies one or more alternative video tracks that may be made available. Each alternative video track is associated with a different non-selected video track. The manifest customization application sends the manifest file to the endpoint application.
[0022] The manifest file enables the endpoint application to play any portion of the media title according to any recommended presentation or any other combination of tracks specified in the manifest file. Furthermore, the manifest file enables the endpoint application to select one of the alternative video tracks specified in the manifest file based on a user's preference. Before or during playback of the media title, the endpoint application may send a request for an alternative manifest file to the manifest customization application based on the selected alternative video track. In response to the request for an alternative manifest file based on the identified alternative video track, the manifest customization application generates an alternative manifest file that describes one or more recommended presentations, each of which includes an alternative video track, the alternative video track, each audio track included in at least one recommended presentation, and each timed text track included in at least one recommended presentation, and identifies zero or more other video tracks that may be made available. The manifest customization application sends the alternative manifest file to the endpoint application. Upon receiving the alternative manifest file, the endpoint application may switch to the selected alternative video track while continuing to play the media title.
[0023] At least one technical advantage of the techniques of the present disclosure over the prior art is that, using the techniques of the present disclosure, a media processing pipeline does not redundantly process source files when generating and deploying a single comprehensive media package to a content delivery network. Therefore, the amount of storage resources, processing resources, and time consumed to generate and deploy a media title may be significantly reduced compared to generating and deploying multiple conventional video-only packages. Another technical advantage of the techniques of the present disclosure is that, unlike prior art approaches, the manifest customization application instantly selects a set of streams from the entire set of streams derived from all source files associated with the corresponding media title based on predetermined user preferences to generate tracks for inclusion in a manifest file. Therefore, the techniques of the present disclosure can generate tracks for playback that are more closely tailored to a user's preferences than is possible with prior art approaches in which a manifest file is generated based on a subset of the streams included in a single video-only package. These technical advantages provide one or more technical advances over prior art approaches.
[0024] System Overview 1 is a conceptual diagram of a system 100 configured to implement one or more aspects of various embodiments. As shown, in some embodiments, system 100 includes, but is not limited to, computation instance 110(1), computation instance 110(2), display device 192, audio device 194, input device 196, origin server device 140, content delivery network (CDN) edge server device 150, and cloud-based media service 160. For purposes of explanation herein, computation instance 110(1), computation instance 110(2), and computation instance 110(3) included in cloud-based media service 160 will be referred to individually as "computation instance 110" and collectively as "computation instance 110."
[0025] In some other embodiments, system 100 may omit one or more of compute instance 110, origin server device 140, CDN edge server device 150, cloud-based media service 160, input device 196, or any combination thereof. In the same or other embodiments, system 100 may further include, without limitation, one or more other compute instances, any number of other display devices, any number of other audio devices, any number and / or types of other cloud-based services, any number of other CDN edge server devices, or any combination thereof.
[0026] Components of system 100 may be implemented in any combination of one or more cloud computing environments (i.e., encapsulating shared resources, software, data, etc.) across any number of shared geographic locations and / or distributed across any number of different geographic locations.
[0027] As shown, compute instance 110(1) includes, but is not limited to, processing unit 112(1) and memory 116(1); compute instance 110(2) includes, but is not limited to, processing unit 112(2) and memory 116(2); and compute instance 110(3) includes, but is not limited to, processing unit 112(3) and memory 116(3). For purposes of explanation, herein, processors 112(1) through 112(3) are individually referred to as "processors 112" and may be collectively referred to as "processors 112." Herein, memories 116(1) through 116(4) are individually referred to as "memories 116" and may be collectively referred to as "memories 116." Each compute instance (including compute instance 110) may be implemented in a cloud computing environment, as part of any other distributed computing environment, or standalone.
[0028] Processing unit 112 may be any instruction execution system, apparatus, or device capable of executing instructions. For example, processing unit 112 may include a central processing unit, a graphics processing unit, a controller, a microcontroller, a state machine, or any combination thereof. Memory 116 of computation instance 110 stores content such as software applications and data for use by processing unit 112 of computation instance 110. Memory 116 may be readily available memory such as one or more of random access memory, read-only memory, a floppy disk, a hard disk, or any other form of digital storage, local or remote.
[0029] In some other embodiments, any number of computing instances may include any number of processing units and any number of memories in any combination. In particular, any number of computing instances (including zero or more computing instances 110) may provide multiple processing environments in any technically feasible manner.
[0030] In some embodiments, a storage unit (not shown) may supplement or replace the memory 116 of the computing instance 110. The storage unit may include any number and type of external memory accessible to the processing unit 112 of the computing instance 110. By way of example and not limitation, the storage unit may include a secure digital card, external flash memory, portable compact disc read-only memory, optical storage device, magnetic storage device, or any suitable combination thereof.
[0031] As indicated in italics, in some embodiments, computation instance 110(2) is a client device. Some examples of client devices include, but are not limited to, a desktop computer, a laptop, a smartphone, a smart TV, a game console, a tablet, etc. Computation instance 110(2) may stream media titles to CDN edge server device 150 via a network connection (not shown). Computation instance 110(2) may play media titles via display device 192 and audio device 194 and receive input from one or more associated users in any technically feasible manner. Audio device 194 may be any type of device (e.g., a speaker) that may be configured to generate sound based on any amount and / or type of audio data in any technically feasible manner. Input device 196 may be any type of device that may be configured to receive user input in any technically feasible manner. Some examples of input devices are a keyboard, a mouse, and a microphone configured to receive user input in any technically feasible manner.
[0032] In some other embodiments, display device 192, audio device 194, input device 196, or any combination thereof, may be replaced with any number and / or types of output devices, any number and / or types of input devices, any number and / or types of input / output devices, or any combination thereof. In the same or other embodiments, computational instance 110(2), any number of other computational instances, display device 192, audio device 194, any number and / or types of other output devices, input device 196, any number and / or types of input devices, any number and / or types of input / output devices, or any combination thereof, are incorporated into a client device.
[0033] In some embodiments, system 100 includes multiple client devices (including computing instance 110(2)) that can each access a media title library provided by a media streaming service. Origin server device 140, any number of other origin server devices, a CDN (not shown) including CDN edge server device 150, and cloud-based media service 160 are included in the delivery infrastructure for the media streaming service.
[0034] Origin server device 140 and zero or more other origin server devices (not shown) collectively store at least one copy of any number of each stream for each media title in a library of media titles for streaming to client devices over one or more content delivery networks (not explicitly shown). As referred to herein, a "stream" constitutes an encoded version of a source file. In particular, a video stream constitutes an encoded version of a video source file, an audio stream constitutes an encoded version of an audio file, and a timed text stream constitutes an encoded version of a timed text file. Each video stream, audio stream, and timed text stream associated with a given media title includes a series of one or more discrete segments that correspond (in a playback timeline) to, but are not limited to, a series of one or more discrete segments of a baseline video source file for the media title. In some embodiments, the baseline video source file for a media title is the original video source file for that media title.
[0035] As used herein, a "video source file" for a media title includes any amount and / or type of primary video content for that media title. Some examples of video source file types are original video source files and localized video source files. A localized video source file for a media title is a version of the original source file for that media title after it has been modified to better suit the personal preferences of the associated target audience. For example, text appearing in the original video source file may be replaced with text of a similar meaning in another language to generate a localized video source file. In another example, one or more scenes in the original video source file may be replaced with scenes more suitable for a different target audience to generate a localized video source file. For example, hockey scenes may be replaced with soccer scenes.
[0036] As used herein, an "audio source file" for a media title includes any amount and / or type of primary audio content for that media title. Some examples of audio source file types are original audio source files, dubbed audio source files, and description audio source files. An original audio source file for a media title provides the original soundtrack for that media title. A dubbed audio source file for a media title provides a translated soundtrack for that media title. An audio description source file provides either the original soundtrack or the translated soundtrack for that media title, plus a narration that verbally describes what is happening on screen (e.g., physical actions, facial expressions, costumes, settings, scene changes).
[0037] As used herein, "timed text" refers to text content and associated playback timing data that enables the text content to be presented simultaneously with other media content, such as video and / or audio content. Some examples of timed text types are closed captions, subtitles, and forced narrative. A subtitle source file provides a transcription or translation of the spoken dialogue of a media title. A closed caption source file provides a transcription or translation of the spoken dialogue of a media title and describes associated non-verbal sounds and various audio cues (e.g., speaker identification). A forced narrative source file provides a text overlay that reveals the communication or alternative language intended to be understood by the viewer (e.g., user). Forced narrative is also commonly referred to as "forced narrative subtitles" and forced "subtitles."
[0038] Upon receiving a request for a stream segment from a client device (e.g., compute instance 110(2)), CDN edge server device 150 locates the closest copy of the requested stream segment and transmits the copy to the client device. More specifically, if CDN edge server device 150 has a copy of the requested stream segment stored in its associated cache memory, a "cache hit" occurs, and CDN edge server device 150 transmits the copy of the requested stream segment. If the requested stream segment is not stored in its associated cache memory, a "cache miss" occurs. If a cache miss occurs, CDN edge server device 150 retrieves the requested stream segment from an intermediate cache memory included in the CDN, origin server device 140, or another origin server device before transmitting the requested stream segment to the client device.
[0039] Cloud-based media services 160 include, but are not limited to, any number and / or types of microservices, databases, and storage for activities and content associated with streaming media services that are not assigned to any origin server device, CDN, or client device. Some examples of functionality that cloud-based media services 160 may provide include, but are not limited to, login and billing, logging, personalized media title recommendations, video transcoding, server and network connection health monitoring, and client-specific CDN guidance.
[0040] Generally, each computing instance (including each computing instance 110) is configured to implement one or more software applications. For purposes of explanation only, each software application is described as residing in the memory 116 of a single computing instance and executing on the processing unit 112 of the same computing instance. However, in some embodiments, the functionality of each software application may be distributed across any number of other software applications resident in the memory of any number of computing instances and executing on the processing units of any number or combination of computing instances. Furthermore, the functionality of any number of software applications may be combined into a single software application.
[0041] In particular, computing instance 110(1) is configured to encode media source files associated with a media title to generate streams and then deploy the streams to a CDN. The media source files include, but are not limited to, an original video source file, zero or more localized video files, an original audio source file, any number of other audio source files, and zero or more timed text source files, and any number and / or type of other media source files (e.g., trick play source files).
[0042] As described in more detail herein, in conventional approaches to generating and deploying streams to a CDN, a distinct single video package is generated independently for each video source file. To generate a video-only package for a given video source file, one or more audio source files and zero or more any number of other types of media source files are selected based on their compatibility with the video source file and the target audience associated with the video-only package. The video source file and each selected source file are encoded to generate distinct sets of streams, or "stream sets," each including at least one corresponding stream. The media processing pipeline then generates a video-only package including each stream set, stores the video-only package on an origin server device, and then deploys the video-only package from the origin server device to a CDN.
[0043] To play a media title, a conventional endpoint application running on a client device requests a manifest file for the media title from a conventional cloud-based media service. In response to the request for the manifest file, the conventional manifest application selects an edge server and a video-only package from a CDN. The conventional manifest application generates one or more recommended presentations based on the video-only package. Each recommended presentation identifies a different combination of the same video track, one or more audio tracks compatible with the video track, and zero or more other tracks, each track including a subset of streams from a different stream set included in the video-only package. The manifest application generates and transmits a manifest file that enables the client device to determine the selected edge server, select a recommended presentation, and further request stream segments for the tracks in the selected recommended presentation.
[0044] One drawback of the above approach is that each audio source file and each timed text source file used to generate more than one video-only package is processed redundantly by the media processing pipeline. As a result, unnecessary amounts of storage resources, processing resources, and time are consumed to generate and deploy different video-only packages to a CDN. Another drawback of the above approach is that the recommended presentations specified in a given manifest file are limited to a single video track and combinations of other tracks derived from the set of streams actually included in the single selected video package. If a user's preferences favor a set of streams omitted from the single selected video package, then the quality of the viewing experience for that user may be subpar.
[0045] Creating a comprehensive media package for a media title To solve the above problem, system 100 includes, but is not limited to, a media processing pipeline 120 that generates a comprehensive media package 130 and a package metadata set 132 based on media source files 102 for a media title. Media source files 102 include, but are not limited to, any number and / or types of media source files for a media title. For purposes of illustration, media source files 102 as described herein include an original video source file, zero or more localized video source files, zero or more other video source files, an original audio source file, zero or more other audio source files, and zero or more timed text source files, but do not include any other types of media source files (e.g., trick play source files). Comprehensive media package 130 is also referred to herein as a “media package.”
[0046] Although not shown in FIG. 1 , the comprehensive media package 130 includes, but is not limited to, a title metadata set that describes the media title and both a media metadata set and a stream set for each media source file associated with the media title. The media metadata set describes relevant characteristics of the associated stream sets, identifies those stream sets, and defines compatibility between them. Some examples of media metadata sets are a video metadata set, an audio metadata set, and a timed text metadata set. Each stream set includes one or more streams and a different stream metadata set for each stream. Some examples of stream sets are a video stream set, an audio stream set, and a timed text stream set. The stream metadata set included in a given stream set describes relevant characteristics of the associated streams and identifies those streams. Referring now to FIG. 2 , an example comprehensive media package 130 is described in more detail. The package metadata set 132 includes a title metadata set, a media metadata set, and a stream metadata set.
[0047] As shown, in some embodiments, media processing pipeline 120 resides in memory 116(1) of compute instance 110(1) and executes on processing unit 112 of compute instance 110. In some other embodiments, any number of portions, including all, of the functionality described herein with respect to media processing pipeline 120 may be distributed across any number of compute instances in any technically feasible manner.
[0048] As shown, media processing pipeline 120 includes, but is not limited to, a comprehensive media package 130, an ingest application 122, an encoding application 124, and a deployment application 126. Ingest application 122 initializes comprehensive media package 130 as an empty package. Ingest application 122 then generates a title metadata set based on a baseline video source file (not shown) and adds it to the title metadata set and comprehensive media package 130.
[0049] The baseline video source file defines the runtime of the media title and is used to establish the segments of the media title. In some embodiments, the ingest application 122 then verifies that any other video source files included in the media source files 102 match (e.g., have the same runtime as) the baseline video source file. In some embodiments, the ingest application 122 defaults to the original video source file.
[0050] A title metadata set may specify any amount and / or type of information related to the playback of a media title. In some embodiments, the title metadata set specifies, without limitation, a title identifier (ID), a baseline video ID, a fallback video ID, a runtime, and a segment definition. The title ID, baseline video ID, and fallback video ID identify the media title, baseline video source file, and fallback video source file, respectively, in any technically feasible manner.
[0051] If none of the video stream sets in the final version of the comprehensive media package 130 are associated with one or more preferred languages for the user, the fallback video source file defines the video stream set sources to use in generating the manifest file. As described in more detail below, the preferred languages for the user are either explicitly identified (e.g., via metadata) or inferred (e.g., determined based on location). Herein, the fallback video source file and the fallback video ID are referred to as the "default source file" and the "default video identifier," respectively.
[0052] Individually, the segments correspond to different, non-overlapping time periods and collectively span runtime in a connected manner. The ingest application 122 may generate the segment definitions in any technically feasible manner. In some embodiments, the ingest application 122 divides the baseline source video file into shots, where each shot is captured sequentially from a single camera or a virtual representation of a camera (e.g., in the case of computer-generated motion video). The ingest application 122 then selects one or more sequences of shots based on the target segment length and defines a segment for each selected sequence of shots.
[0053] For each video source file included in the media source files 102, the ingest application 122 generates an associated video metadata set and adds the video metadata set to the comprehensive media package 130. Each video metadata set specifies a different set of values for a set of video features. In some embodiments, the set of video features includes language and video type (e.g., original, localized), and optionally any other characteristics of the video content relevant to streaming the media title, such as aspect ratio.
[0054] For each audio source file included in the media source files 102, the ingest application 122 generates an associated audio metadata set and adds the audio metadata set to the comprehensive media package 130. Each audio metadata set specifies a different set of values for a set of audio features. In some embodiments, the set of audio features includes language, audio type, and zero or more other characteristics of the audio content associated with streaming the media title. Some examples of audio types are original audio, dubbed audio, and audio description.
[0055] For each timed text source file included in the media source files 102, the ingest application 122 generates an associated timed text metadata set and adds the timed text metadata set to the comprehensive media package 130. Each timed text metadata set specifies a different set of values for a set of timed text features. In some embodiments, the set of timed text features includes language, timed text type, optional timed text subtypes, and zero or more other characteristics of the audio content associated with streaming a media title. Some examples of timed text types are subtitles, closed captions, and forced narrative. The optional timed text subtypes can be either full or partial. As described in more detail below, the optional timed text subtypes are used to prevent the display of overlapping text during playback of a media title.
[0056] The Full Timed Text subtype indicates that the corresponding timed text source file contains a translation of text included in the video content of the media title. Thus, the Full Timed Text subtype indicates that the associated timed text source file and corresponding timed text stream are compatible with each video source file and corresponding video stream associated with a different language. The Partial Timed Text subtype indicates that the corresponding timed text source file does not contain a translation of text included in the video content of the media title. Thus, the Partial Timed Text subtype indicates that the associated timed text source file and corresponding timed text stream are compatible with each video source file and corresponding video stream associated with the same language.
[0057] The encoding application 124 generates a stream set for each media source file 102 and adds each stream set to the comprehensive media package 130. Within the comprehensive media package 130, the stream set for a given media source file 102 is associated with the media metadata set for the media source file 102 in any technically feasible manner (e.g., by index, ID, etc.).
[0058] Each stream set includes, but is not limited to, at least one stream and a different stream metadata set for each stream. As previously described herein, a stream is an encoded version of a source file. As will be appreciated by those skilled in the art, the number of video streams included in a video stream set is typically significantly greater than the number of streams included in any other type of stream set (e.g., an audio stream set). The encoding application 124 may generate the streams in any technically feasible manner that allows each segment of the stream to be decoded independently of other segments of that stream. Ensuring that each stream segment is decoded independently allows endpoint applications to perform stream switching at segment boundaries during playback of a media title.
[0059] In some embodiments, the encoding application 124 reduces the video source file to multiple different resolutions to generate reduced-resolution video source files. The encoding application 124 then runs a shot-based encoder on the video source file and each reduced-resolution video source file across different sets of one or more encoding profiles to generate video streams having different combinations of resolution and bitrate. The encoding application 124 propagates any amount of metadata contained in each media metadata set to each stream in the associated stream set.
[0060] The stream metadata sets included in a given stream set describe relevant characteristics of and identify the associated streams, for example, in some embodiments, a video stream metadata set specifies the size in bytes, bit rate, resolution, encoding profile, frame rate, and quality score for the associated stream.
[0061] The deployment application 126 adds any amount (including none) and / or type of deployment-related metadata (not shown) to the comprehensive media package 130. One example of deployment-related metadata is digital rights management (DRM) metadata. The deployment application 126 extracts a package metadata set 132 from the comprehensive media package 130. As described in more detail below, the package metadata set 132 is associated with the comprehensive media package 130 on which both the manifest file 174 and the alternate manifest files 178 are based. In particular, the manifest file 174 and the alternate manifest files 178 itemize different video tracks derived from different sets of video streams included in the comprehensive media package 130l.
[0062] In some embodiments, package metadata set 132 includes, but is not limited to, a title metadata set, a video metadata set, a stream metadata set (contained in a stream set), and any deployment-related metadata. In some other embodiments, package metadata set 132 includes a subset of the title metadata set, a subset of the video metadata set, a subset of the stream metadata set, any amount of deployment-related metadata, or any combination thereof, related to generating a manifest file, as described in more detail below.
[0063] As shown, deployment application 126 stores package metadata set 132 in a media service database 162 included in cloud-based media service 160. As shown, deployment application 126 transmits comprehensive media package 130 to origin server device 140. Origin server device 140 stores comprehensive media package 130 in memory associated with origin server device 140. Any number and / or types of software applications executing on any number and / or types of server devices included in the CDN may transmit requests for comprehensive media package 130 to origin server device 140 at any time. In response to each request, the software application executing on origin server device 140 transmits comprehensive media package 130 to the requesting software application. In this manner, deployment application 126 delivers a playback stream of the media title via comprehensive media package 130.
[0064] In particular, a software application executing on CDN edge server device 150 sends a request for comprehensive media package 130 to origin server device 140, either directly or indirectly through any number of intermediate server devices in the CDN. After receiving comprehensive media package 130, the software application executing on CDN edge server device 150 stores comprehensive media package 130 in a cache memory associated with CDN edge server device 150. Referring now to Figure 2, an example comprehensive media package 130 will be described in more detail.
[0065] As shown, cloud-based media service 160 includes, but is not limited to, media service database 162 and compute instance 110(3). Although not shown, cloud-based media service 160 generates, maintains, and stores in media service database 162 CDN metadata set 134, a different client device metadata set for each client device, and a different user metadata set for each user associated with the streaming media service.
[0066] To generate and maintain CDN metadata set 134, cloud-based media service 160 monitors the health of CDN edge server device 150 and any number of other CDN edge server devices, compute instances 110(2), any number of other client devices, and associated network connections within any number of CDNs. CDN metadata set 134 identifies any quantity and / or type of characteristics of CDN edge server devices and network connections. In some embodiments, CDN metadata set 134 identifies, without limitation, routing distance, throughout, and latency for any number of connections between CDN edge server devices and client devices.
[0067] Each client device metadata set identifies any number and / or type of characteristics of the associated client device. For example, in some embodiments, each client device metadata set identifies, without limitation, the location and display resolution of the associated client device. Each user metadata set identifies any number and / or type of characteristics and preferences of the associated user. In some embodiments, each user metadata set identifies, without limitation, profile language, audio preferences, and timed text preferences. The profile language for a user identifies the language of the user interface that the streaming media provider uses to interact with the user. The audio preferences identify the language and, optionally, the audio type. The timed text preferences identify the language and the timed text type.
[0068] Advantageously, the media processing pipeline 120 does not redundantly process any of the media source files 102 when generating and deploying the comprehensive media package 130. Therefore, the amount of storage resources, processing resources, and time consumed to generate and deploy a media title may be significantly reduced compared to conventional techniques that represent a media title using multiple video-only packages. Furthermore, as described in more detail below, the comprehensive media package 130 and package metadata set 132 enable the manifest customization application 170 to generate recommended presentations of media titles and associated playback tracks that are more closely tailored to a user's preferences than may be possible using multiple video-only packages. For that matter, the comprehensive media package 130 includes at least one encoded version (i.e., stream) of each media source file 102. As a result, the manifest customization application 170 can instantly generate tracks to include in the manifest file from the entire set of video streams, audio streams, and any set of timed text streams derived from all media source files 102, based on the package metadata set 132, as well as any number and / or type of other associated metadata, predetermined user preferences, etc.
[0069] Generating a manifest file based on a comprehensive media package As shown, manifest customization application 170 resides in memory 116(23) of compute instance 110(3) and executes on processing unit 112(23) of compute instance 110(3). Manifest customization application 170 generates manifest files personalized for different users in response to manifest requests received from client devices. Each manifest request identifies a user and a media title.
[0070] For purposes of explanation, the functionality of manifest customization application 170 is described herein in the context of a sequence of events designated by circled numbers 1 through 11. Additionally, for purposes of explanation, comprehensive media package 130 as described herein includes video, audio, and timed text stream sets, but does not include any other type of stream set (e.g., trick play stream sets).
[0071] As indicated by circled number 1, manifest customization application 170 receives manifest request 172 from endpoint application 180. Endpoint application 180 resides in memory 116(3) of compute instance 110(2) (client device) and executes on processing unit 112(3) of compute instance 110(2). Manifest request 172 identifies compute instance 110(2), the user, and the media title associated with media source file 102.
[0072] In response to manifest request 172, manifest customization application 170 retrieves package metadata set 132, CDN metadata set 134, client device metadata set 136, and user metadata set 138 (denoted by the circled number 2) from media services database 162. Client device metadata set 136 and user metadata set 138 are associated with compute instance 110(2) and the user of compute instance 110(2), respectively.
[0073] Manifest customization application 170 designates each CDN edge server device 150, and optionally one or more other CDN edge server devices, as a "recommended" CDN edge server device based on CDN metadata set 134 and client device metadata set 136. Typically, manifest customization application 170 designates at least one CDN edge server device that is close to the client device as a recommended server device.
[0074] The manifest customization application 170 selects one of the video stream sets included in the comprehensive media package 130 based on the package metadata set 132, the user metadata set 138, and optionally the client device metadata set 136. In some embodiments, the manifest customization application 170 determines whether the language and / or aspect ratio specified in any video metadata set in the package metadata set 132 matches a preferred language and / or preferred aspect ratio specified in the user metadata set 138 and / or the client device metadata set 136. If the manifest customization application 170 finds a matching video stream set and determines that rights to access the matching video stream set are not restricted for the location associated with the client device, then the manifest customization application 170 selects the matching video stream set. Otherwise, the manifest customization application 170 selects the video stream set associated with the default video identifier specified in the title metadata set included in the package metadata set 132.
[0075] If audio preferences are specified in user metadata set 138, then manifest customization application 170 determines whether one of the audio stream sets included in comprehensive media package 130 matches the audio preferences based on the audio metadata set in package metadata set 132. If manifest customization application 170 finds a matching audio stream set, then manifest customization application 170 selects the matching audio stream set. Otherwise, manifest customization application 170 performs any number and / or types of heuristics to select an audio stream set based on user metadata set 138 and package metadata set 132, optionally based on client device metadata set 136, and optionally based on the selected video stream set.
[0076] If a timed text preference is specified in user metadata set 138, then manifest customization application 170 determines whether any timed text stream sets included in comprehensive media package 130 match the timed text preference based on the timed text metadata sets in package metadata set 132. If manifest customization application 170 finds a single matching timed text stream set, then manifest customization application 170 selects the matching timed text stream set.
[0077] However, if the manifest customization application 170 finds two different matching sets of timed text streams, then the manifest customization application 170 selects one of the matching timed text streams based on its compatibility with the selected set of video streams. In some embodiments, if the language specified in the timed text preference matches the language associated with the selected set of video streams, then the manifest customization application 170 selects the matching timed text stream associated with the partial timed text subtype. Otherwise, the manifest customization application 170 selects the matching timed text stream associated with the full timed text subtype.
[0078] If the manifest customization application 170 does not find any matching timed text stream sets, then the manifest customization application 170 performs any number and / or types of heuristics to select a timed text stream set based on the user metadata set 138 and the package metadata set 132, optionally based on the client device metadata set 136, and optionally based on the selected video stream set.
[0079] The manifest customization application 170 selects at least one video stream from at least one video stream included in the selected video stream set for inclusion in the video track. Thus, the video track is a subset of the video streams included in the selected video stream set. As used herein, a "subset" may be a proper subset or an improper subset. The manifest customization application 170 selects at least one audio stream from the audio streams included in the selected audio stream set for inclusion in the audio track. Thus, the audio track is a subset of the audio streams included in the selected audio stream set. If a timed text stream set is selected, then the manifest customization application 170 selects at least one timed text stream from the timed text streams included in the selected timed text stream set for inclusion in the timed text track. The manifest customization application 170 may select one or more streams for inclusion in the track in any technically feasible manner.
[0080] In some embodiments, manifest customization application 170 selects at least one stream from the streams included in the selected stream set for inclusion in the track based on the corresponding stream metadata set, the associated media metadata set, the client device metadata set 136, the CDN metadata set 134, the user metadata set 138, any amount and / or type of other data, any number of selection criteria, or any combination thereof. Some examples of selection criteria that manifest customization application 170 may use to select streams for inclusion in the track are network throughput, stream bit rate, user preferences, compatibility with the client device, rights restrictions associated with the location of the client device, and rights restrictions associated with the user.
[0081] If the manifest customization application 170 selects at least one timed text stream for inclusion in the timed text track, then the manifest customization application 170 further combines the video track, the audio track, and the timed text track to generate the recommended presentation. Otherwise, the manifest customization application 170 combines the video track and the audio track to generate the recommended presentation.
[0082] Manifest customization application 170 defines zero or more other recommended presentations, each specifying a different combination of video tracks, audio tracks, and optional timed text tracks. Manifest customization application 170 may define the recommended presentations and any number of other audio tracks and / or other timed text tracks specified in the recommended presentations, in any technically feasible manner, based on any number and / or types of heuristics and any number and / or types of criteria.
[0083] The manifest customization application 170 generates a video track description for the video track, a different audio track description for each audio track, and a different timed text track description for each timed text track. Herein, the video tracks, audio tracks, and timed text tracks are also referred to as "tracks." Each track description identifies a track identifier, an associated subset of media metadata sets, the streams contained in the corresponding track, a subset of stream metadata sets for each stream, and one or more locators for each stream.
[0084] A locator for a stream is a reference that can be used, along with any amount (including none) of other data in the manifest file, to request individual segments of the stream from the recommended server device. In some embodiments, the track description may identify a different Universal Record Locator (URL) for each segment of each stream in the corresponding track. In some other embodiments, the track description may identify a reference Universal Record Locator (URL) associated with the first segment of the first stream in the corresponding track and with segment offsets for each other stream segment in the corresponding track.
[0085] In particular, manifest customization application 170 may generate a different alternative video track summary for each of zero or more "alternate" video tracks for inclusion in the manifest file. Each alternative video track summary indicates that an associated alternative video track may be available. More specifically, each alternative video track summary identifies different alternative video tracks that may originate from different non-selected video stream sets included in comprehensive media package 130, specifies at least a subset of the associated video metadata set, and optionally specifies any amount and / or type of other associated metadata. For example, in some embodiments, each alternative video track summary specifies metadata including an alternative video track identifier, language, and video type, and optionally aspect ratio. In the same or other embodiments, the alternative video track identifier embeds an identifier for the corresponding video stream set.
[0086] Manifest customization application 170 generates manifest file 174 that identifies, without limitation, each recommended server device, title description set, each recommended presentation, each track description, and each alternative video track summary. The title description set identifies a media title and further specifies any amount and / or type of metadata associated with the playback of the media title. In some embodiments, the title description set is at least a subset of the title metadata set included in comprehensive media package 130. Referring now to FIG. 3, an example manifest file 174 is described in more detail.
[0087] As indicated by circled number 3, manifest customization application 170 sends manifest file 174 to endpoint application 180, thereby fulfilling manifest request 172. As described in more detail below, the manifest file enables endpoint application 180 to play the media title via any recommended presentations.
[0088] To play the media title, endpoint application 180 selects a track identified in one of the recommended presentations. For each segment of the media title, endpoint application 180 selects a corresponding segment from one stream in each selected track. As indicated by circled number 4, endpoint application 180 sends segment request 182 for the selected stream segment to CDN edge server device 150. In response, CDN edge server device 150 sends stream segment 184 to endpoint application 180 (indicated by circled number 5). Endpoint application 180 decodes each stream segment to generate decoded media content and plays the media content.
[0089] Advantageously, playing a media title via a recommended presentation automatically tailors (e.g., personalizes) the media viewing experience for the user to, at least in part, the user's preferences as identified in user metadata set 138. Additionally, as described below, each alternative video track description in manifest file 174 may provide endpoint application 180 with the opportunity to perform video track switching to further tailor the media viewing experience to the user's preferences during playback of the media title.
[0090] For example, in some embodiments, endpoint application 180 allows a user, via a graphical user interface, to select any alternative video track based on the language and / or aspect ratio specified in the corresponding alternative video track summary. In some other embodiments, including the embodiment shown in FIG. 1, endpoint application 180 automatically selects an alternative video track based on any type of user information.
[0091] As indicated by circled number 6, during playback of the media title, endpoint application 180 receives an audio language change request 198 from the user via input device 196. In response, endpoint application 180 determines that the language associated with one of the alternative video tracks matches the newly requested audio language. As indicated by circled number 7, while playback of the media title continues, endpoint application 180 sends an alternative manifest request 176 to manifest customization application 170. Alternate manifest request 176 identifies the client device, the user, the media title, and an alternative video track that matches the requested audio language.
[0092] Generating alternative manifest files for video track switching Manifest customization application 170 generates alternate manifest files individually tailored to different users in response to alternate manifest requests received from client devices. Each alternate manifest request identifies the user, the media title, and the requested alternate video track. As used herein, an "alternate manifest file" refers to a manifest file for a media title generated based on a request for a manifest file that items the alternate video track. In some embodiments, alternate manifest request 176 is a request for a manifest file that identifies an alternate video track and / or a video stream set from which the alternate video track is derived.
[0093] In response to alternate manifest request 176, manifest customization application 170 again retrieves package metadata set 132, CDN metadata set 134, client device metadata set 136, and user metadata set 138 from media service database 162 (indicated by circled number 8). Manifest customization application 170 selects a set of video streams from the set of video streams included in comprehensive media package 130 based on package metadata set 132 and the metadata associated with the requested alternate video tracks and specified in alternate manifest request 176. Manifest customization application 170 selects one or more video streams from the video streams included in the selected set of video streams for inclusion in the alternate video tracks.
[0094] In some embodiments, manifest customization application 170 preferentially selects a set of audio streams from the set of audio streams included in comprehensive media package 130 based on associated metadata identified in package metadata set 132 and languages associated with alternative video tracks. Manifest customization application 170 selects one or more audio streams from the audio streams included in the selected set of audio streams for inclusion in the audio track.
[0095] In some embodiments, manifest customization application 170 optionally selects one or more timed text streams from comprehensive media package 130 for inclusion in any timed text track, using techniques similar to those described herein with respect to generating manifest file 174. If manifest customization application 170 selects at least one timed text stream for inclusion in a timed text track, then manifest customization application 170 further assembles alternative video tracks, audio tracks, and timed text tracks to generate the recommended presentation. Otherwise, manifest customization application 170 assembles alternative video tracks and audio tracks to generate the recommended presentation.
[0096] Manifest customization application 170 defines zero or more other recommended presentations, each specifying a different combination of alternative video tracks and audio tracks and optional timed text tracks. Manifest customization application 170 generates video track descriptions for the alternative video tracks, different audio track descriptions for each audio track, and different timed text track descriptions for each timed text track. Optionally, manifest customization application 170 may generate any number of different alternative video track summaries for inclusion in alternative manifest files 174.
[0097] Manifest customization application 170 generates an alternate manifest file 178 that identifies, without limitation, one or more recommended server devices, a title description set, one or more recommended presentations, video track descriptions for requested alternate video tracks, an audio track description for each audio track included in at least one of the recommended presentations, a timed text track description for each timed text track included in at least one of the recommended presentations, and zero or more alternate video track summaries. An example alternate manifest file 178 will now be described in more detail with reference to FIG. 4.
[0098] As indicated by circled number 9, manifest customization application 170 sends alternate manifest file 178 to endpoint application 180, thereby fulfilling alternate manifest request 176. In response, endpoint application 180 deselects the currently selected track and selects the track identified in one of the recommended presentations identified in alternate manifest request 176.
[0099] For each segment of the media title, endpoint application 180 selects a corresponding stream segment from one stream in each newly selected track. As indicated by circled number 10, endpoint application 180 sends an alternative segment request 186 for the selected stream segment to CDN edge server device 150. In response, CDN edge server device 150 sends alternative stream segment 188 to endpoint application 180 (indicated by circled number 11). Endpoint application 180 decodes each alternative stream segment to generate decoded media content and plays the media content. In this aspect, manifest customization application 170 enables video switching during playback of the media title.
[0100] It should be noted that the techniques described herein are illustrative rather than limiting, and may be modified without departing from the broader spirit and scope of the present embodiments. Those skilled in the art will appreciate that numerous modifications and variations will be apparent to those of ordinary skill in the art regarding the functionality of the media processing pipeline 120, ingest application 122, encoding application 124, deployment application 126, cloud-based media service 160, manifest customization application 170, endpoint application 180, origin server device 140, CDN, and CDN edge server device 150 described herein without departing from the scope and spirit of the described embodiments.
[0101] Similarly, the storage units, configurations, amounts and / or types of data described herein are exemplary and not limiting, and may be modified without departing from the broader spirit and scope of the present embodiments. In that regard, those skilled in the art will recognize numerous modifications and variations of the media source files 102, comprehensive media packages 130, package metadata sets 132, CDN metadata sets 134, media service databases 162, client device metadata sets 136, user metadata sets 138, manifest requests 172, manifest files 174, alternate manifest requests 176, alternate manifest files 178, segment requests 182, stream segments 184, alternate segment requests 186, alternate stream segments 188, and audio language change requests 198 described herein without departing from the scope and spirit of the described embodiments.
[0102] For example, in some embodiments, ingest application 122 may generate a video metadata set that specifies values for any number and / or type of characteristics that collectively distinguish the corresponding video stream sets, i.e., different video tracks. Manifest customization application 170 may use the distinguishing characteristics to select the video stream sets and then generate a manifest file based on the video streams. Endpoint application 180 may use the distinguishing characteristics (specified in the manifest file) to request an alternative manifest file based on one of the alternative video tracks specified in the manifest file.
[0103] In some alternative embodiments, the media source files 102 include one or more trick play source files. Each trick play source file specifies a subset of frames of an associated video source file. For each trick play source file included in the media source files 102, the ingest application 122 generates an associated trick play metadata set and adds the trick play metadata set to the comprehensive media package 130. Each trick play metadata set specifies a matching video ID, a trick play type, and zero or more other characteristics of the trick play associated with streaming the media title. The matching video ID identifies the video source file and / or associated video metadata set with which the trick play source file is intended to be used in any technically feasible manner. Some examples of trick play types are fast forward, skip mode, rewind, and search. For each trick play metadata set, the encoding application 124 generates a trick play stream set based on the associated trick play source file and / or associated video source file and adds the trick play stream set to the comprehensive media package 130. Each trick-play stream identifies a subset of encoded frames in the corresponding video stream. Deployment application 126 adds trick-play metadata sets to package metadata set 132. Manifest customization application 170 adds zero or more trick-play track descriptions to manifest file 174 and alternate manifest files 178. Endpoint application 180 provides visual feedback using one or more trick-play tracks specified in a manifest file (e.g., manifest file 174, alternate manifest file 178) during non-normal playback modes, such as fast-forward mode, skip mode, rewind mode, and search mode.
[0104] It will be appreciated that the system 100 illustrated herein is illustrative only, and that modifications and variations are possible. For example, the functionality provided by the ingest application 122, encoding application 124, and deployment application 126 as described herein may be incorporated into or distributed across any number of software applications (including one) and any number of components of the system 100. Furthermore, the connection topology between the various units of Figure 1 may be modified as desired.
[0105] Exemplary Comprehensive Media Package Figure 2 is a more detailed diagram of the comprehensive media package 130 of Figure 1 in accordance with various embodiments. More specifically, the comprehensive media package 130 shown in Figure 2 is a typical comprehensive media package. As shown, in some embodiments, the comprehensive media package 130 includes, but is not limited to, title metadata set 210, video stream sets 220(1) through 220(3), video metadata sets 230(1) through 230(3), audio stream sets 240(1) through 240(6), audio metadata sets 250(1) through 250(6), timed text stream sets 260(1) through 260(6), and timed text metadata sets 270(1) through 270(6).
[0106] As previously described in this specification with reference to FIG. 1, the title metadata set 210 identifies the title ID (titleID) associated with the media title, the runtime of the media title, the segment definitions (segmentDefinitions) for the media title, the baseline video ID (baselineVideoID), and the fallback video ID (fallbackVideoID).
[0107] As shown, video metadata set 230(1) identifies a language of English (en) and a video type (videoType) of original. Video metadata set 230(1) is associated with video stream set 220(1). Thus, video stream set 220(1) identifies one or more video streams and a different stream metadata set for each video stream, each video stream being a different encoded version of the English original video source file.
[0108] As shown, video metadata set 230(2) identifies a language of French (fr) and a video type of localized. Video metadata set 230(2) is associated with video stream set 220(2). Thus, video stream set 220(2) identifies one or more video streams and a different stream metadata set for each video stream, each video stream being a different encoded version of a French-localized video source file.
[0109] As shown, video metadata set 230(3) identifies a language of Spanish (es) and a video type of localized. Video metadata set 230(3) is associated with video stream set 220(3). Thus, video stream set 220(3) identifies one or more video streams and a different stream metadata set for each video stream, each video stream being a different encoded version of a Spanish-localized video source file.
[0110] As shown, audio metadata set 250(1) identifies a language of English and an audio type (audioType) of Original. Audio metadata set 250(1) is associated with audio stream set 240(1). Audio stream set 240(1) therefore identifies one or more audio streams and a different stream metadata set for each audio stream, each audio stream being a different encoded version of the English original audio source file.
[0111] As shown, audio metadata set 250(2) identifies the language French and the audio type dubbed. Audio metadata set 250(2) is associated with audio stream set 240(2). Audio stream set 240(2) therefore identifies one or more audio streams and a different stream metadata set for each audio stream, each audio stream being a different encoded version of the French dubbed audio source file.
[0112] As shown, audio metadata set 250(3) identifies a language of Spanish, a timed text type of subtitle, and a timed text subtype of full. Audio metadata set 250(3) is associated with audio stream set 240(3). Thus, audio stream set 240(3) identifies one or more audio streams and a different stream metadata set for each audio stream, each audio stream being a different encoded version of a Spanish-dubbed audio source file.
[0113] As shown, audio metadata set 250(4) identifies a language of English and an audio type of description. Audio metadata set 250(4) is associated with audio stream set 240(4). Thus, audio stream set 240(4) identifies one or more audio streams and a different stream metadata set for each audio stream, each audio stream being a different encoded version of the English audio description source file.
[0114] As shown, audio metadata set 250(5) identifies a language of French and an audio type of description. Audio metadata set 250(5) is associated with audio stream set 240(5). Thus, audio stream set 240(5) identifies one or more audio streams and a different stream metadata set for each audio stream, each audio stream being a different encoded version of a French audio description source file.
[0115] As shown, audio metadata set 250(6) identifies a language of Spanish and an audio type of description. Audio metadata set 250(6) is associated with audio stream set 240(6). Thus, audio stream set 240(6) identifies one or more audio streams and a different stream metadata set for each audio stream, each audio stream being a different encoded version of a Spanish audio description source file.
[0116] As shown, timed text metadata set 270(1) identifies a language of English, a timed text type (timedTextType) of subtitle, and a timed text subtype (timedTextSubtype) of full. Timed text metadata set 270(1) is associated with timed text stream set 260(1). Timed text stream set 260(1) therefore identifies one or more timed text streams and a different stream metadata set for each timed text stream, each timed text stream being a different encoded version of the English full subtitle source file.
[0117] As shown, timed text metadata set 270(2) identifies the language French, the timed text type Subtitle, and the timed text subtype Full. Timed text metadata set 270(2) is associated with timed text stream set 260(2). Timed text stream set 260(2) therefore identifies one or more timed text streams and a different stream metadata set for each timed text stream, where each timed text stream is a different encoded version of a French full subtitle source file.
[0118] As shown, timed text metadata set 270(3) identifies a language of Spanish, a timed text type of subtitle, and a timed text subtype of full. Timed text metadata set 270(3) is associated with timed text stream set 260(3). Timed text stream set 260(3) therefore identifies one or more timed text streams and a different stream metadata set for each timed text stream, each timed text stream being a different encoded version of a Spanish full subtitle source file.
[0119] As shown, timed text metadata set 270(4) identifies a language of English, a timed text type of subtitle, and a partial timed text subtype. Timed text metadata set 270(4) is associated with timed text stream set 260(4). Timed text stream set 260(4) therefore identifies one or more timed text streams and a different stream metadata set for each timed text stream, each timed text stream being a different encoded version of an English partial subtitle source file.
[0120] As shown, timed text metadata set 270(5) identifies the language French, the timed text type subtitle, and the partial timed text subtype. Timed text metadata set 270(5) is associated with timed text stream set 260(5). Timed text stream set 260(5) therefore identifies one or more timed text streams and a different stream metadata set for each timed text stream, each timed text stream being a different encoded version of a French partial subtitle source file.
[0121] As shown, timed text metadata set 270(6) identifies the language Spanish, the timed text type subtitle, and the partial timed text subtype. Timed text metadata set 270(6) is associated with timed text stream set 260(6). Timed text stream set 260(6) therefore identifies one or more timed text streams and a different stream metadata set for each timed text stream, each timed text stream being a different encoded version of a Spanish partial subtitle source file.
[0122] Example Manifest File Figure 3 is a more detailed diagram of the manifest file 174 of Figure 1 in accordance with various embodiments. More particularly, the manifest file 174 shown in Figure 3 is a typical manifest file. For illustrative purposes, the manifest customization application 170 generates the manifest file 174 based on the example comprehensive media package 130 shown in Figure 2, a location of Spain, a preferred language (e.g., profile language) of English, a preferred audio language of English, a preferred timed text language of English, and a preferred timed text type of subtitles.
[0123] As shown, manifest file 174 identifies, without limitation, recommended server device 310, media title description 320, recommended presentations 330(1) through 330(6), video track description 340, audio track description 350(1), audio track description 350(2), timed text track description 360(1), timed text track description 360(2), alternative video track summary 370(1), and alternative video track summary 370(2). As shown, recommended server device 310 is CDN edge server device 150 of FIG. 1. Media title description 320 is equivalent to title metadata set 210 identified in comprehensive media package 130.
[0124] Recommended presentation 330(1) identifies an English original video track associated with video track description 340, an English original audio track associated with audio track description 350(1), and an English partial subtitle track associated with timed text track description 360(1). Recommended presentation 330(2) identifies an English original video track, an English original audio track, and a Spanish full subtitle track associated with timed text track description 360(2). Recommended presentation 330(3) identifies an English original video track and an English original audio track.
[0125] Recommended presentation 330(4) identifies the English original video track, the English audio description track associated with audio track description 350(2), and the English partial subtitle track. Recommended presentation 330(5) identifies the English original video track, the English audio description track, and the Spanish full subtitle track. Recommended presentation 330(6) identifies the English original video track and the English audio description track.
[0126] Video track description 340 describes a video track that includes a subset of video streams in video stream set 220(1). Referring again to FIG. 2, video stream set 220(1) is an English original video stream set. In some embodiments, video track description 340 identifies a video track identifier, a language of English, a video type of original, the constituent video streams, and stream-specific metadata for each constituent video stream. The stream-specific metadata for a given constituent video stream identifies the size in bytes, bit rate, resolution, encoding profile, frame rate, quality score, and one or more URLs that CDN edge server device 150 can use to retrieve each segment of the video stream.
[0127] Audio track description 350(1) describes an audio track that includes a subset of the audio streams in audio stream set 240(1). Referring again to FIG. 2, audio stream set 240(1) is an English original audio stream set. In some embodiments, audio track description 350(1) identifies an audio track identifier, a language of English, an audio type of original, one or more constituent audio streams, and stream-specific metadata for each constituent audio stream. Similarly, audio track description 350(2) describes an audio track that includes a subset of the audio streams in audio stream set 240(4). Referring again to FIG. 2, audio stream set 240(4) is an English audio description stream set.
[0128] Timed text track description 360(1) describes a timed text track that includes a subset of timed text streams in timed text stream set 260(4). Referring again to Figure 2, timed text stream set 260(4) is a set of English partial subtitle streams. Timed text track description 360(2) describes a timed text track that includes a subset of timed text streams in timed text stream set 260(3). Referring again to Figure 2, timed text stream set 260(3) is a set of Spanish full subtitle streams.
[0129] Alternate video track summary 370(1) identifies alternate video tracks that may be derived from video stream set 220(3) and specifies any amount of video metadata set 230(3) that corresponds to video stream set 220(3). Referring again to Figure 2, video stream set 220(3) is a Spanish localized video stream set. In some embodiments, alternate video track summary 370(1) therefore specifies alternate video track identifiers that embed identifiers for video stream set 220(3), a language of Spanish, and a video type of localized.
[0130] Alternate video track summary 370(2) identifies alternate video tracks that may be derived from video stream set 220(2) and specifies any amount of video metadata set 230(2) that corresponds to video stream set 220(2). Referring again to Figure 2, video stream set 220(2) is a French-localized video stream set. Thus, in some embodiments, alternate video track summary 370(2) specifies alternate video track identifiers that embed identifiers for video stream set 220(2), a language of French, and a video type of localized.
[0131] Advantageously, alternate video track summary 370(1) and alternate video track summary 370(2) enable endpoint application 180 to coordinate a video switching process during which endpoint application 180 switches from playing a media title via a recommended presentation that includes an English original video track to playing a media title via a recommended presentation that includes either a Spanish localized video track or a French localized video track.
[0132] Example Alternate Manifest File Figure 4 is a more detailed diagram of the alternative manifest file 178 of Figure 1 in accordance with various embodiments. More particularly, the alternative manifest file 178 shown in Figure 4 is an exemplary alternative manifest file. For illustrative purposes, the manifest customization application 170 generates the alternative manifest file 178 based on an identifier for the alternative video track summary 370(2) of Figure 3, the comprehensive media package 130 shown in Figure 2, a preferred timed text language of English, and a preferred timed text type of subtitles.
[0133] More specifically, manifest customization application 170 selects video stream set 220(2) based on the identifier for alternative video track summary 370(2). Then, manifest customization application 170 generates alternative manifest file 178 based on video stream set 220(2), comprehensive media package 130, and a timed text preference of English subtitles. Referring again to Figure 2, video stream set 220(2) is a French-localized video stream set, and therefore, manifest customization application 170 generates alternative manifest file 178 based on the French-localized video tracks.
[0134] As shown, alternate manifest file 178 identifies, without limitation, recommended server device 410, media title description 320, recommended presentations 430(1) through 430(6), video track description 440, audio track description 450(1), audio track description 450(2), timed text track description 460(1), timed text track description 460(2), alternate video track summary 470(1), and alternate video track summary 470(2). As shown, recommended server device 410 is CDN edge server device 150 of FIG. 1. Media title description 320 is equivalent to title metadata set 210 identified in comprehensive media package 130.
[0135] Recommended presentation 430(1) identifies a French localized video track associated with video track description 440, a French dubbed audio track associated with audio track description 450(1), and a full English subtitle track associated with timed text track description 460(1). Recommended presentation 430(2) identifies a French localized video track, a French dubbed audio track, and a French partial subtitle track associated with timed text track description 460(2). Recommended presentation 430(3) identifies a French localized video track and a French dubbed audio track.
[0136] Recommended presentation 430(4) identifies a French localized video track, an English original audio track associated with video track description 440, a French audio description track associated with audio track description 450(2), and an English full subtitle track. Recommended presentation 430(5) identifies a French localized video track, a French audio description track, and a French partial subtitle track. Recommended presentation 430(3) identifies a French localized video track and a French audio description track.
[0137] Video track description 440 describes a French localized video track that includes a subset of the video streams in video stream set 220(2). In some embodiments, video track description 440 identifies a video track identifier, a language of French, a video type of localized, the constituent video streams, and stream-specific metadata for each constituent video stream.
[0138] Audio track description 450(1) describes a French dubbed audio track that includes a subset of the audio streams in audio stream set 240(2). Audio track description 450(2) describes a French audio description track that includes a subset of the audio streams in audio stream set 240(5). Timed text track description 460(1) describes an English full subtitle track that includes a subset of the timed text streams in timed text stream set 260(1). Timed text track description 460(2) describes a French partial subtitle track that includes a subset of the timed text streams in timed text stream set 260(5).
[0139] Alternate video track summary 470(1) identifies alternative English original video tracks that may be derived from video stream set 220(1) and specifies any amount of video metadata set 230(1) that corresponds to video stream set 220(1). Alternate video track summary 470(2) identifies alternative Spanish localized video tracks that may be derived from video stream set 220(3) and specifies any amount of video metadata set 230(3) that corresponds to video stream set 220(3).
[0140] Technique for providing a playback stream associated with a media title 5 is a flow diagram of method steps for generating and deploying streams associated with media titles to a content delivery network according to various embodiments. The method steps are described with reference to the systems of FIGS. 1-4, but one skilled in the art will recognize that any system configured to perform the method steps in any order is within the scope of the present embodiments.
[0141] As shown, method 500 begins at step 502, where an ingest application 122 generates an initial version of a comprehensive media package including a title metadata set based on baseline video source files associated with a media title. At step 504, for each video source file associated with the media title, the ingest application 122 adds a video metadata set to the comprehensive media package that identifies the video type and language.
[0142] In step 506, for each audio source file associated with the media title, the ingest application 122 adds an audio metadata set identifying the audio type and language to the comprehensive media package, and in step 508, for each of zero or more timed text source files associated with the media title, the ingest application 122 adds a timed text metadata set identifying the language, timed text type, and any timed text subtype to the comprehensive media package.
[0143] At step 510, the encoding application 124 generates one or more streams for each media source file, and each segment of each stream can be decoded independently. At step 512, the encoding application 124 propagates each media metadata set to the corresponding stream, generating a stream metadata set for each stream. At step 514, for each media source file, the encoding application 124 adds a stream set including the corresponding stream and associated stream metadata set to the comprehensive media package.
[0144] At step 516, the deployment application 126 generates a package metadata set for the media title based on the title metadata set, the media metadata set, and the stream metadata set. At step 518, the deployment application 126 sends the package metadata set to a software application that generates a customized manifest file for the media title. At step 520, the deployment application 126 stores the comprehensive media package on the origin server device and deploys the stored comprehensive media package from the origin server device to the CDN. Method 500 then ends.
[0145] 6 is a flow diagram of method steps for generating a manifest file that itemizes streams associated with a media title according to various embodiments. The method steps are described with reference to the systems of FIGS. 1-4, but one skilled in the art will recognize that any system configured to perform the method steps in any order is within the scope of the present embodiments.
[0146] As shown, method 600 begins at step 602, when manifest customization application 170 receives a request for a manifest file from a client device. At step 604, manifest customization application 170 retrieves a package metadata set, a user metadata set, a client device metadata set, and a CDN metadata set 134 corresponding to a comprehensive media package associated with the media title based on the request.
[0147] At step 606, the manifest customization application 170 sets a recommended server device equivalent to an edge server device in the CDN based on the CDN metadata set 134 and the client device metadata set. At step 608, the manifest customization application 170 selects one of a set of video streams, one of a set of audio streams, and optionally one of a set of timed text streams in the comprehensive media package based on the package metadata set and the user metadata set. At step 610, the manifest customization application 170 defines a first recommended presentation that identifies a video track derived from the selected set of video streams and a different track derived from each of the other selected sets of streams.
[0148] In step 612, manifest customization application 170 defines zero or more other recommended presentations, each identifying a different combination of video tracks, audio tracks, and optional timed text tracks derived from the comprehensive media package. In step 614, for each identified track, manifest customization application 170 generates a track description that identifies a subset of streams in the corresponding stream set, associated metadata, and locators for segments of each identified stream.
[0149] At step 616, the manifest customization application 170 generates different alternative video track summaries that identify the language and video type for each of any number (including none) of the unselected video stream sets in the comprehensive media package. At step 618, the manifest customization application 170 generates a manifest file that identifies recommended server devices, recommended presentations, track descriptions, and one or more alternative video track summaries. At step 620, the manifest customization application 170 transmits the manifest file to the client device. Method 600 then ends.
[0150] 7 is a flow diagram of method steps for generating a manifest file to enable video track switching during playback of a media title according to various embodiments. The method steps are described with reference to the systems of FIGS. 1-4, but one skilled in the art will recognize that any system configured to perform the method steps in any order is within the scope of the present embodiments.
[0151] As shown, method 700 begins at step 702, where manifest customization application 170 sends a manifest file to a client device that enables streaming of one video track and identifies one or more alternate video tracks. At step 704, manifest customization application 170 receives a request from the client device to generate an alternate manifest file that enables streaming of the alternate video tracks.
[0152] At step 706, the manifest customization application 170 retrieves the package metadata set, the user metadata set, the client device metadata set, and the CDN metadata set 134 corresponding to the comprehensive media package associated with the media title. At step 708, the manifest customization application 170 selects a video stream set corresponding to the requested alternative video track and further selects an associated language based on the package metadata set.
[0153] At step 710, the manifest customization process selects a set of audio streams in the comprehensive media package based on the selected language and the package metadata set. At step 712, the manifest customization application 170 optionally selects a set of timed text streams in the comprehensive media package based on the user metadata set, the package metadata set, and the selected language.
[0154] At step 714, the manifest customization application 170 defines a first recommended presentation that identifies the recommended video tracks and different tracks derived from each of the other selected stream sets. At step 716, the manifest customization application 170 defines zero or more other recommended presentations, each identifying a different combination of recommended video tracks and audio and optional timed text tracks derived from the comprehensive media package.
[0155] At step 718, the manifest customization application 170 generates an alternative manifest file that enables the client device to stream the recommended presentation and therefore the corresponding track from an edge server of the CDN. At step 720, the manifest customization application 170 transmits the alternative manifest file to the client device. Method 700 then ends.
[0156] That is, the techniques of this disclosure may be used to generate and deploy a comprehensive media package for a media title that enables streaming playback of the media title via suggested presentations flexibly tailored to different users. In some embodiments, the media title is associated with an original language video file, one or more other video files, one or more audio files, zero or more timed text files, and zero or more other media files. An ingest application included in the media processing pipeline generates a title metadata set that identifies and describes any number of characteristics of the media title based on the original language video file associated with the media title. The ingest application adds the title metadata set to an initially empty comprehensive media package for the media title. For each media file (including the original language video file), the ingest application generates a corresponding media metadata set and adds it to the comprehensive media package. In particular, the media metadata set describes relevant characteristics of the associated stream sets, identifies those stream sets, and defines compatibility between them. In particular, each media metadata set identifies a language. Additionally, each timed text metadata set may identify any timed text type that indicates a language-based compatibility with the set of video streams.
[0157] An encoding application included in the media processing pipeline encodes each media file one or more times based on different encoding profiles to generate one or more corresponding streams. The encoding application propagates each media metadata set to the corresponding streams and generates a stream metadata set for each stream. The stream metadata set included in a given stream set describes relevant characteristics of and identifies the associated streams. For each media metadata set, the encoding application generates a corresponding stream set that includes the corresponding stream and stream metadata set, and adds the corresponding stream set to a comprehensive media package.
[0158] A deployment application included in the media processing pipeline generates a package metadata set including a title metadata set, a media metadata set, and a stream set. The deployment application stores the package metadata set in a media service database accessible to a manifest customization application and any number of other applications included in the cloud-based media service. The deployment application stores the comprehensive media package on an origin server device. A software application running on the origin server device provides the comprehensive media package on request to software applications running on server devices included in the CDN.
[0159] In some embodiments, the manifest customization application 170 receives a request for a manifest file for a media title from an endpoint application running on a client device associated with a user. In response, the manifest customization application 170 retrieves the package metadata set, CDN metadata, client device metadata, and user metadata associated with the media title from the media service database. The user may specify any number and / or types of preferences associated with the user (e.g., timed text language, timed text type). The manifest customization application 170 selects one of the video stream sets in the comprehensive media package based on the package metadata set and the preferences associated with the user specified in the user metadata (e.g., preferred language, preferred aspect ratio). The manifest customization application 170 selects at least one of the video streams in the selected video stream set for inclusion in a video track and generates one or more recommended presentations based on the video track, the package metadata set, and the user metadata. Each recommended presentation includes a different combination of a video track, an audio track, and an optional timed text track. The manifest customization application 170 generates a different alternative video track summary that identifies the alternative video track identifier, language, and video type for each of any number (including none) of the non-selected video stream sets in the comprehensive media package.
[0160] The manifest customization application 170 then generates a manifest file identifying recommended server devices included in the CDN, recommended presentations, track descriptions for each track identified in at least one recommended presentation, and zero or more alternative video track summaries. Each track description identifies a track identifier, language, media type, constituent streams, and stream-specific metadata for each constituent stream. The stream-specific metadata for a given constituent stream enables the endpoint application to request each segment of the constituent stream from the recommended server device. Each alternative video track summary provides an opportunity to switch from streaming a video track identified in the recommended presentation to streaming a video track associated with a different language. The manifest customization application sends the manifest file to the endpoint application.
[0161] In some embodiments, while streaming and playing a media title through one of the recommended presentations, the endpoint application sends a request to the manifest customization application 170 for an alternative manifest file that identifies alternative video track identifiers. In response, the manifest customization application 170 selects a set of video streams included in the comprehensive media package based on the alternative video track identifiers and the package metadata set. The manifest customization application 170 designates one or more of the video streams identified in the selected set of video streams as alternative video tracks and generates one or more new recommended presentations, each including the alternative video tracks, based on the package metadata set and the user metadata. The manifest customization application 170 then generates an alternative manifest file that identifies the recommended server device, the new recommended presentations, track descriptions for each track identified in the at least one new recommended presentation, and zero or more alternative video track summaries. The manifest customization application 170 then sends the alternative manifest file to the endpoint application. Upon receiving the alternate manifest file, the endpoint application switches from streaming the recommended presentations identified in the manifest file to streaming the recommended presentations identified in the alternate manifest file.
[0162] At least one technical advantage of the disclosed technology over the prior art is that, using the disclosed technology, when a single, all-encompassing media package that more comprehensively represents a media title is generated and deployed to a content delivery network, source files are not processed redundantly by a media processing pipeline, as compared to a conventional video-only package. Therefore, the amount of storage resources, processing resources, and time consumed to generate and deploy a media title may be significantly reduced compared to prior art approaches. Another technical advantage of the disclosed technology is that, unlike prior art approaches, tracks identified in a manifest file generated using a single all-encompassing media package may be instantly selected based on predetermined user preferences from the entire set of streams derived from all source files associated with the corresponding media title. Therefore, the disclosed technology may generate tracks for playback that are more closely tailored to a user's preferences than is possible using prior art approaches in which a manifest file is generated based on a subset of the streams contained in a single video-only package. These technical advantages provide one or more technical advances over prior art approaches. 1. In some embodiments, a computer-implemented method for switching video tracks during playback of a media title includes selecting one or more video streams for inclusion in a first video track, the one or more video streams derived from a first video source file associated with the media title and included in a media package generated for the media title; generating a manifest file based on the first video track, the manifest file itemizing the first video track and indicating that a first alternative video track may be made available; transmitting the manifest file to a client device; receiving a request from the client device to generate an alternative manifest file that itemizes the first alternative video track; selecting one or more video streams for inclusion in the first alternative video track, the one or more video streams derived from a second video source file associated with the media title and included in the media package; generating the alternative manifest file based on the first alternative video track; and transmitting the alternative manifest file to the client device. 2. The computer-implemented method of claim 1, wherein generating the manifest file includes generating a first alternative video track summary associated with the first alternative video track based on metadata associated with a first set of video streams derived from a second video source file and included in the media package. 3. The computer-implemented method of claim 1 or 2, wherein the first alternative video track summary identifies at least one of an identifier, a language, or an aspect ratio. 4. The computer-implemented method of any one of paragraphs 1 to 3, wherein the step of selecting one or more video streams for inclusion in the first alternative video track includes the steps of selecting a first set of video streams from a plurality of sets of video streams included in the media package based on metadata associated with the first alternative video track and specified in the request, and selecting one or more video streams from the first set of video streams. 5. The computer-implemented method of any one of paragraphs 1 to 4, wherein generating the alternative manifest file includes selecting a first set of audio streams from a plurality of sets of audio streams included in the media package based on a first language associated with a first alternative video track, selecting at least one audio stream from the first set of audio streams for inclusion in the first audio track, and combining the first alternative video track and the first audio track to generate a first recommended presentation. 6. The computer-implemented method of any one of clauses 1 to 5, wherein generating the alternative manifest file includes selecting a first set of timed text streams from a plurality of sets of timed text streams included in the media package based on at least one of a first preference associated with a first user or a first language associated with a first alternative video track; selecting at least one timed text stream from the first set of timed text streams for inclusion in the first timed text track; and combining the first alternative video track, the first audio track, and the first timed text track to generate a first recommended presentation. 7. The computer-implemented method of any one of clauses 1 to 6, wherein the first preference includes at least one of a timed text language or a timed text type. 8. The computer-implemented method of any one of paragraphs 1 to 7, wherein generating the alternative manifest file includes generating a description of the first alternative video track that identifies at least one of a different bit rate, a different resolution, a different quality score, or a different locator for each video stream included in the first alternative video track. 9. The computer-implemented method of any one of clauses 1 to 8, wherein the first video source file is associated with a first language and the second video source file is associated with a second language different from the first language. 10. The computer-implemented method of any one of clauses 1 to 9, wherein generating the manifest file includes selecting a first set of audio streams from a plurality of sets of audio streams included in the media package based on an audio language associated with the first user, selecting at least one audio stream from the first set of audio streams for inclusion in a first audio track, and assembling the first video track and the first audio track to generate a first recommended presentation. 11. In some embodiments, one or more non-transitory computer-readable media include instructions that, when executed by one or more processors, cause the one or more processors to: select one or more video streams for inclusion in a first video track, the one or more video streams derived from a first video source file associated with the media title and included in a media package generated for the media title; generate a manifest file based on the first video track, the manifest file itemizing the first video track and indicating that a first alternate video track may be made available; and a media package for selecting one or more video streams for inclusion in the first alternative video track, the one or more video streams being derived from a second video source file associated with the media title and included in the media package; generating an alternative manifest file based on the first alternative video track; and transmitting the alternative manifest file to the client device, thereby switching video tracks during playback of the media title. 12. One or more non-transitory computer-readable media described in paragraph 11, wherein generating the manifest file includes generating multiple alternative video track summaries based on metadata associated with multiple sets of video streams included in the media package, each alternative video track summary indicating that a different alternative video track may be made available. 13. One or more non-transitory computer-readable media described in paragraph 11 or 12, wherein each alternative video track summary included in the plurality of alternative video track summaries specifies at least one of a different identifier, a different language, or a different aspect ratio. 14. One or more non-transitory computer-readable media described in any one of clauses 11 to 13, wherein the step of selecting one or more video streams for inclusion in the first alternative video track includes the steps of selecting a first set of video streams from a plurality of sets of video streams included in the media package based on metadata associated with the first alternative video track and specified in the request, and selecting one or more video streams from the first set of video streams. 15. One or more non-transitory computer-readable media described in any one of clauses 11 to 14, wherein the step of generating the alternative manifest file includes the steps of selecting a first set of audio streams from a plurality of sets of audio streams included in the media package based on a first language associated with a first alternative video track, selecting at least one audio stream from the first set of audio streams for inclusion in the first audio track, and assembling the first alternative video track and the first audio track to generate a first recommended presentation. 16. One or more non-transitory computer-readable media described in any one of clauses 11 to 15, wherein generating the alternative manifest file includes selecting a first set of timed text streams from a plurality of sets of timed text streams included in the media package based on at least one of a first preference associated with a first user or a first language associated with a first alternative video track; selecting at least one timed text stream from the first set of timed text streams for inclusion in the first timed text track; and combining the first alternative video track, the first audio track, and the first timed text track to generate a first recommended presentation. 17. One or more non-transitory computer-readable media described in any one of paragraphs 11 to 16, wherein the first preference includes at least one of a timed text language or a timed text type. 18. One or more non-transitory computer-readable media described in any one of paragraphs 11 to 17, wherein generating the alternative manifest file includes generating a description of the first alternative video track that identifies at least one of a different bit rate, a different resolution, a different quality score, or a different locator for each video stream included in the first alternative video track. 19. One or more non-transitory computer-readable media described in any one of clauses 11 to 18, wherein the first video source file is associated with a first aspect ratio and the second video source file is associated with a second aspect ratio that is different from the first aspect ratio. 20. In some embodiments, a system includes one or more memories that store instructions and one or more processors coupled to the one or more memories, the instructions being configured, when executed, to perform the following steps: select one or more video streams for inclusion in a first video track, the one or more video streams derived from a first video source file associated with the media title and included in a media package generated for the media title; generate a manifest file based on the first video track, the manifest file itemizing the first video track and indicating that a first alternative video track may be made available; transmit the manifest file to a client device; receive a request from the client device to generate an alternate manifest file that itemizes the first alternative video track; select one or more video streams for inclusion in the first alternative video track, the one or more video streams derived from a second video source file associated with the media title and included in the media package; generate the alternate manifest file based on the first alternative video track; and transmit the alternate manifest file to the client device.
[0163] Any and all combinations of any claim element recited in any claim and / or any element described herein, in any form, are within the intended scope of the invention and protection.
[0164] The description of various embodiments is for purposes of illustration and is not intended to be exhaustive or to limit the disclosed embodiments. Numerous modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the disclosed embodiments.
[0165] Aspects of the present disclosure may be embodied as a system, method, or computer program product. Accordingly, aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be referred to herein as a "module" or "system." Furthermore, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable medium(s) having computer-readable program code embedded therein.
[0166] Any combination of one or more computer-readable media may be used. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media may include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory, a flash memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination thereof. In the context of this specification, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0167] Aspects of the present disclosure have been described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine. The instructions, when executed by a processor of the computer or other programmable data processing apparatus, cause the processor to perform the function(s) / act(s) identified in the block or blocks of the flowcharts and / or block diagrams. Such a processor may be, but is not limited to, a general-purpose processor, a special-purpose processor, an application-specific processor, or a field-programmable gate array.
[0168] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams represents a module, segment, or portion of code, which includes one or more executable instructions for implementing a particular logical function. It should also be noted that in some alternative embodiments, the functions shown in the blocks may be performed in a different order than that shown in the figures. For example, two blocks shown in succession may actually be performed substantially simultaneously, depending on the functionality involved, or in some cases, the blocks may be performed in the reverse order. Furthermore, it will be understood that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs a particular function or operation, or by a combination of dedicated hardware and computer instructions.
[0169] While embodiments of the present disclosure have been described above, other and further embodiments of the present disclosure may be contemplated without departing from the basic scope of the disclosure, which scope is determined by the following claims. [Explanation of symbols]
[0170] 102 media source files 130 Comprehensive Media Package 134 CDN Metadata Sets 136 Client Device Metadata Set 138 User Metadata Set 140 Origin Server Device 150 Content transmission network edge server device 174 Manifest File 178 Alternate Manifest File 192 Display device 194 Audio Equipment 196 Input Devices
Claims
1. 1. A computer-implemented method for switching video tracks during playback of a media title, comprising: selecting one or more video streams for inclusion in a first video track, the one or more video streams derived from a first video source file associated with the media title and included in a media package generated for the media title; generating a manifest file based on the first video track, the manifest file itemizing the first video track and indicating that a first alternative video track may be made available; transmitting the manifest file to a client device; receiving a request from the client device to generate an alternate manifest file that itemizes the first alternate video track; selecting one or more video streams for inclusion in the first alternative video track, the one or more video streams derived from a second video source file associated with the media title and included in the media package; generating the alternative manifest file based on the first alternative video track; transmitting the alternate manifest file to the client device; 10. A computer-implemented method comprising:
2. 2. The computer-implemented method of claim 1, wherein generating the manifest file includes generating a first alternative video track summary associated with the first alternative video track based on metadata associated with a first set of video streams derived from the second video source file and included in the media package.
3. The computer-implemented method of claim 2 , wherein the first alternative video track summary identifies at least one of an identifier, a language, or an aspect ratio.
4. selecting the one or more video streams for inclusion in the first alternative video track comprises: selecting a first set of video streams from a plurality of sets of video streams included in the media package based on metadata associated with the first alternative video track and specified in the request; selecting the one or more video streams from the first set of video streams; The computer-implemented method of claim 1 , comprising:
5. The step of generating the alternative manifest file includes: selecting a first set of audio streams from a plurality of sets of audio streams included in the media package based on a first language associated with the first alternative video track; selecting at least one audio stream from the first set of audio streams for inclusion in a first audio track; combining the first alternative video track and the first audio track to generate a first recommended presentation; The computer-implemented method of claim 1 , comprising:
6. The step of generating the alternative manifest file includes: selecting a first set of timed text streams from a plurality of sets of timed text streams included in the media package based on at least one of a first preference associated with a first user or a first language associated with the first alternative video track; selecting at least one timed text stream from the first set of timed text streams for inclusion in a first timed text track; combining the first alternative video track, the first audio track, and the first timed text track to generate a first recommended presentation; The computer-implemented method of claim 1 , comprising:
7. The computer-implemented method of claim 6 , wherein the first preference includes at least one of a timed text language or a timed text type.
8. 2. The computer-implemented method of claim 1, wherein generating the alternative manifest file includes generating a description of the first alternative video track that identifies at least one of a different bit rate, a different resolution, a different quality score, or a different locator for each video stream included in the first alternative video track.
9. 10. The computer-implemented method of claim 1, wherein the first video source file is associated with a first language and the second video source file is associated with a second language different from the first language.
10. The step of generating the manifest file includes: selecting a first set of audio streams from a plurality of sets of audio streams included in the media package based on an audio language associated with a first user; selecting at least one audio stream from the first set of audio streams for inclusion in a first audio track; combining the first video track and the first audio track to generate a first recommended presentation; The computer-implemented method of claim 1 , comprising:
11. One or more non-transitory computer-readable media containing instructions that, when executed by one or more processors, cause the one or more processors to: selecting one or more video streams for inclusion in a first video track, the one or more video streams derived from a first video source file associated with the media title and included in a media package generated for the media title; generating a manifest file based on the first video track, the manifest file itemizing the first video track and indicating that a first alternative video track may be made available; transmitting the manifest file to a client device; receiving a request from the client device to generate an alternate manifest file that itemizes the first alternate video track; selecting one or more video streams for inclusion in the first alternative video track, the one or more video streams derived from a second video source file associated with the media title and included in the media package; generating the alternative manifest file based on the first alternative video track; transmitting the alternate manifest file to the client device; and one or more non-transitory computer-readable media for causing a video track to be switched during playback of a media title.
12. generating the manifest file includes generating a plurality of alternative video track summaries based on metadata associated with a plurality of sets of video streams included in the media package; 12. The one or more non-transitory computer-readable media of claim 11, wherein each alternative video track summary indicates that a different alternative video track may be enabled.
13. 13. The one or more non-transitory computer-readable media of claim 12, wherein each alternative video track summary included in the plurality of alternative video track summaries specifies at least one of a different identifier, a different language, or a different aspect ratio.
14. selecting the one or more video streams for inclusion in the first alternative video track comprises: selecting a first set of video streams from a plurality of sets of video streams included in the media package based on metadata associated with the first alternative video track and specified in the request; selecting the one or more video streams from the first set of video streams; 12. The one or more non-transitory computer-readable media of claim 11, comprising:
15. The step of generating the alternative manifest file includes: selecting a first set of audio streams from a plurality of sets of audio streams included in the media package based on a first language associated with the first alternative video track; selecting at least one audio stream from the first set of audio streams for inclusion in a first audio track; combining the first alternative video track and the first audio track to generate a first recommended presentation; 12. The one or more non-transitory computer-readable media of claim 11, comprising:
16. The step of generating the alternative manifest file includes: selecting a first set of timed text streams from a plurality of sets of timed text streams included in the media package based on at least one of a first preference associated with a first user or a first language associated with the first alternative video track; selecting at least one timed text stream from the first set of timed text streams for inclusion in a first timed text track; combining the first alternative video track, the first audio track, and the first timed text track to generate a first recommended presentation; 12. The one or more non-transitory computer-readable media of claim 11, comprising:
17. 17. The one or more non-transitory computer-readable media of claim 16, wherein the first preference includes at least one of a timed text language or a timed text type.
18. 12. The one or more non-transitory computer-readable media of claim 11, wherein generating the alternative manifest file includes generating a description of the first alternative video track that identifies at least one of a different bit rate, a different resolution, a different quality score, or a different locator for each video stream included in the first alternative video track.
19. 12. The one or more non-transitory computer-readable media of claim 11, wherein the first video source file is associated with a first aspect ratio and the second video source file is associated with a second aspect ratio that is different from the first aspect ratio.
20. In the system, one or more memories for storing instructions; one or more processing units coupled to the one or more memories; the one or more processing units, when the instructions are executed, selecting one or more video streams for inclusion in a first video track, the one or more video streams derived from a first video source file associated with the media title and included in a media package generated for the media title; generating a manifest file based on the first video track, the manifest file itemizing the first video track and indicating that a first alternative video track may be made available; transmitting the manifest file to a client device; receiving a request from the client device to generate an alternate manifest file that itemizes the first alternate video track; selecting one or more video streams for inclusion in the first alternative video track, the one or more video streams derived from a second video source file associated with the media title and included in the media package; generating the alternative manifest file based on the first alternative video track; transmitting the alternate manifest file to the client device; A system that performs the following: