Methods, apparatuses, and computer readable media for signaling chained auxiliary media content on a DASH media stream
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-15
- Publication Date
- 2026-08-11
AI Technical Summary
[0006]然而,当DASH播放器使用W3C媒体源扩展时,由于使用单个MSE源缓冲区来解决这种非线性播放问题非常具有挑战性,因而即使是MPD链接和前贴片广告插入也会失败
[0008] This disclosure addresses one or more technical problems. It includes methods, processes, apparatuses, and non-transitory computer-readable media for implementing the novel concept—the DASH standard—of secondary presentation and secondary MPD, which allow secondary or standalone media presentations to be described in accordance with the primary media presentation. Embodiments of this disclosure also relate to secondary presentations including secondary media content, which may be presented as pre-roll, mid-roll, or end-roll media content in other secondary presentations. Embodiments also involve stacking multiple secondary presentations.
Smart Images

Figure CN116965008B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to U.S. Provisional Application No. 63 / 298,919, filed January 12, 2022, and U.S. Application No. 18 / 065,154, filed December 13, 2022, the entire contents of which are expressly incorporated herein by reference. Technical Field
[0003] Embodiments of this disclosure relate to streaming media content, and more specifically, to streaming media, advertising, and live content based on Dynamic Adaptive Streaming over Hypertext Transfer Protocol (DASH) according to the Moving Picture Experts Group (MPEG). Background Technology
[0004] MPEG DASH provides a standard for streaming media content over IP networks. In MPEG DASH, Media Representation Descriptions (MPDs) and events are used to deliver media timeline-related events to clients. The ISO / IEC 23009-1 DASH standard allows for streaming of multi-rate content. The DASH standard provides a single linear timeline, where each segment is a continuation of the previous one within a single timeline. ISO / IEC 23009-1 also provides tools for MPD linking, i.e., signaling the URL of the next MPD for playback within an MPD that can be used for pre-roll ad insertion.
[0005] MPEG DASH provides a standard for streaming multimedia content over IP networks. While this standard solves the problem of linear playback of media content, it does not address non-linear scenarios, such as media segments associated with different timelines that are independent of each other. These shortcomings can be overcome using MPEG DASH links and pre-roll ad insertion.
[0006] However, when the DASH player uses the W3C Media Source Extension, it is very challenging to solve this non-linear playback problem using a single MSE source buffer, so even MPD links and pre-roll ad insertions will fail.
[0007] Therefore, a method is needed for combining supplementary or independent content that differs from the main media content. Specifically, a method and apparatus are needed for combining supplementary content with the main media content as a pre-roll, mid-roll, or end-roll playback. A method for stacking supplementary content is needed. Furthermore, a method is needed for carrying information associated with the supplementary content and stacking information. Summary of the Invention
[0008] This disclosure addresses one or more technical problems. It includes methods, processes, apparatuses, and non-transitory computer-readable media for implementing the novel concept—the DASH standard—of secondary presentation and secondary MPD, which allow secondary or standalone media presentations to be described in accordance with the primary media presentation. Embodiments of this disclosure also relate to secondary presentations including secondary media content, which may be presented as pre-roll, mid-roll, or end-roll media content in other secondary presentations. Embodiments also involve stacking multiple secondary presentations.
[0009] Embodiments of this disclosure may provide a method for signaling chained auxiliary media content, including pre-roll media content, mid-roll media content, and end-roll media content, in an HTTP-based Dynamic Adaptive Streaming (DASH) main media stream. The method may include: receiving one or more auxiliary descriptors, wherein each of the one or more auxiliary descriptors includes a Uniform Resource Locator (URL) referencing one or more auxiliary Media Presentation Descriptions (MPDs) and a stacking mode value indicating a stacking operation supported by the main DASH media stream; retrieving one or more auxiliary media segments based on the URLs referenced in the respective auxiliary descriptors, wherein the one or more auxiliary media segments are independent of one or more main DASH media segments; and playing the one or more auxiliary media segments and one or more main DASH media segments at least once from a Media Source Extension (MSE) source buffer in at least one order based on the one or more auxiliary descriptors and the stacking mode value.
[0010] Embodiments of this disclosure may provide an apparatus for signaling a chain of auxiliary media content, including pre-roll media content, middle-roll media content, and end-roll media content, in an HTTP-based Dynamic Adaptive Streaming (DASH) main media stream. The apparatus may include: at least one memory configured to store computer program code; and at least one processor configured to access the computer program code and operate according to the instructions in the computer program code. The program code may include: receiving code configured to cause the at least one processor to receive one or more auxiliary descriptors, wherein each of the one or more auxiliary descriptors includes a Uniform Resource Locator (URL) referencing one or more auxiliary Media Presentation Descriptions (MPDs) and a stacking mode value indicating a stacking operation supported by a primary DASH media stream; retrieval code configured to cause the at least one processor to retrieve one or more auxiliary media segments based on the URLs referenced in the respective auxiliary descriptors of the one or more auxiliary descriptors, wherein the one or more auxiliary media segments are independent of one or more primary DASH media segments; and playback code configured to cause the at least one processor to play one or more auxiliary media segments and one or more primary DASH media segments at least once from a Media Source Extension (MSE) source buffer based on the one or more auxiliary descriptors and the stacking mode value in at least one order.
[0011] Embodiments of this disclosure may provide a non-transitory computer-readable medium storing instructions. These instructions may include one or more instructions that, when executed by one or more processors of a device for signaling chained auxiliary media content including pre-roll media content, middle-roll media content, and end-roll media content in an HTTP-based Dynamic Adaptive Streaming (DASH) main media stream, cause the one or more processors to: receive one or more auxiliary descriptors, wherein each of the one or more auxiliary descriptors includes a Uniform Resource Locator (URL) referencing one or more auxiliary Media Presentation Descriptions (MPDs) and a stacking mode value indicating stacking operations supported by the main DASH media stream; retrieve one or more auxiliary media segments based on the URLs referenced in each of the one or more auxiliary descriptors, wherein the one or more auxiliary media segments are independent of one or more main DASH media segments; and play the one or more auxiliary media segments and one or more main DASH media segments at least once from a Media Source Extension (MSE) source buffer in at least one order based on the one or more auxiliary descriptors and the stacking mode value. Attached Figure Description
[0012] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:
[0013] Figure 1This is a simplified illustration of a communication system according to an embodiment.
[0014] Figure 2 This is an example illustration of component placement in a streaming environment according to an embodiment.
[0015] Figure 3 This is a simplified block diagram of the DASH processing model according to an embodiment.
[0016] Figure 4 This is an exemplary flowchart, according to an embodiment, for signaling a chain of auxiliary media content, including pre-roll media content, middle-roll media content, and end-roll media content, in an HTTP-based Dynamic Adaptive Streaming (DASH) main media stream.
[0017] Figure 5 A simplified diagram of a computer system according to an embodiment. Detailed Implementation
[0018] The proposed features discussed below can be used individually or in any combination in any order. Furthermore, embodiments can be implemented using processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.
[0019] Figure 1 A simplified block diagram of a communication system 100 according to an embodiment of the present disclosure is shown. The communication system 100 may include at least two terminals 102 and 103 interconnected via a network 105. For one-way data transmission, the first terminal 103 may encode video data from its local location for transmission to the other terminal 102 via the network 105. The second terminal 102 may receive the encoded video data from the other terminal from the network 105, decode the encoded data, and display the recovered video data. One-way data transmission is common in media service applications, etc.
[0020] Figure 1 A second pair of terminals 101 and 104 is shown, providing bidirectional transmission of encoded video, for example, during video conferencing. For bidirectional data transmission, each terminal 101 and 104 can encode video data captured at a local location for transmission to the other terminal via network 105. Each terminal 101 and 104 can also receive encoded video data transmitted by the other terminal, decode the encoded data, and display the recovered video data on a local display device.
[0021] exist Figure 1In this disclosure, terminals 101, 102, 103, and 104 may be shown as servers, personal computers, and smartphones, but the principles of this disclosure are not limited thereto. Embodiments of this disclosure are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network 105 refers to any number of networks, including, for example, wired and / or wireless communication networks, that transmit encoded video data between terminals 101, 102, 103, and 104. Communication network 105 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks (LANs), wide area networks (WANs), and / or the Internet. For the purposes of this discussion, unless explained below, the architecture and topology of network 105 may be irrelevant to the operation of this disclosure.
[0022] As an example, Figure 2 The placement of the video encoder and video decoder in a streaming environment is illustrated. This embodiment can be applied to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0023] The streaming system may include an acquisition subsystem 203, which may include a video source 201, such as a digital camera, that creates, for example, an uncompressed video sample stream 213. This sample stream 213 may be emphasized as having a high data volume compared to an encoded video bitstream and may be processed by an encoder 202 coupled to the video source 201. The encoder 202 may include hardware, software, or a combination of hardware and software to implement or carry out aspects of the embodiments described in more detail below. The encoded video bitstream 204 may be emphasized as having a lower data volume compared to the sample stream and may be stored on a streaming server 205 for future use. One or more streaming clients 212 and 207 may access the streaming server 205 to retrieve encoded video bitstreams 208 and 206, which may be copies of the encoded video bitstream 204. Client 212 may include a video decoder 211 that decodes the incoming copy 208 of the encoded video bitstream and produces an output video sample stream 210 that may be displayed on a display 209 or other presentation device. In some streaming systems, encoded video bitstreams can be encoded in 204, 206, and 208 octaves according to certain video coding / compression standards. Examples of these standards have been mentioned above and are described further in this document.
[0024] Figure 3An example DASH processing model 300 is shown, for example, as an example client architecture for handling DASH and CMAF events. In DASH processing model 300, client requests for media segments (e.g., advertising media segments and live media segments) can be based on the addresses described in Listing 303. Listing 303 also describes metadata paths from which clients can access segments, parse these segments, and send them to application 301.
[0025] Listing 303 includes MPD events or events, in-band events, and a "moof" parser 306 that can parse MPD event segments or event segments and append these event segments to the event and metadata buffer 330. The in-band event and "moof" parser 306 can also retrieve media segments and append them to the media buffer 340. The event and metadata buffer 330 can send event and metadata information to the event and metadata synchronizer and scheduler 335. The event and metadata synchronizer and scheduler 335 can schedule specific events to DASH player control, selection, and heuristic logic 302, and schedule application-related event and metadata paths to the application 301.
[0026] According to some embodiments, the MSE may include a pipeline including a file format parser 350, a media buffer 340, and a media decoder 345. The MSE 320 is one or more logical buffers for media segments, in which media segments can be recorded and ordered based on their presentation time. Media segments may include, but are not limited to, advertising media segments associated with advertising MPDs and live media segments associated with live MPDs. Each media segment can be added to or appended to the media buffer 340 based on its timestamp offset, and the timestamp offset can be used to order the media segments in the media buffer 340.
[0027] Since embodiments of this application can construct a Linear Media Source Extension (MSE) buffer from two or more non-linear media sources using MPD links, and the non-linear media sources can be advertising MPDs and live MPDs, the file format parser 350 can be used to process the different media and / or codecs used by the live media segments included in the live MPD. In some embodiments, the file format parser can publish a change type based on the codec, profile, and / or level of the live media segment.
[0028] As long as a media segment exists in the media buffer 340, the event and metadata buffer 330 retains the corresponding event segment and metadata. The example DASH processing model 300 may include a timed metadata path resolver 325 to record metadata associated with in-band and MPD events. Figure 3The MSE 320 includes only a file format parser 350, a media buffer 340, and a media decoder 345. The event and metadata buffer 330, as well as the event and metadata synchronizer and scheduler 335, are not native devices of the MSE 320, thus preventing the MSE 320 from processing events locally and sending them to the application.
[0029] Auxiliary presentation
[0030] Embodiments of this disclosure define auxiliary media presentation as a media presentation independent of the main media presentation of the MPD. For example, an advertising media segment or a live media segment, independent of the main media segment, can be an auxiliary presentation. Updates to any auxiliary media presentation or auxiliary media segment do not affect the main media segment. Similarly, updates to the main media segment do not affect the auxiliary media segment. Therefore, an auxiliary media segment (also referred to as an auxiliary media presentation or auxiliary presentation) can be completely independent of the main media segment (also referred to as the main media presentation and media presentation in this disclosure).
[0031] Assisted MPD
[0032] An MPD is a media presentation description that may include media presentations in a hierarchical structure. An MPD may include one or more time-segment sequences, where each time-segment may include one or more adaptive sets. Each adaptive set in an MPD may include one or more presentations, and each presentation includes one or more media segments. These one or more media segments carry the actual media data that is encoded, decoded, and / or played, along with associated metadata. Auxiliary MPDs may include one or more auxiliary media segments.
[0033] As described above, embodiments of this disclosure define auxiliary MPDs representing supplementary content independent of the main media content. According to one aspect, the main MPD may include references to at least one auxiliary MPD using auxiliary descriptors, or in some embodiments, may include references to each auxiliary MPD using auxiliary descriptors. The auxiliary descriptors may have a specific syntax. As an example, an auxiliary descriptor may include a descriptor called a basic descriptor, or may include descriptors called supplementary descriptors that can describe or identify the auxiliary MPD.
[0034] According to one aspect of this disclosure, the main MPD may include URL links to one or more auxiliary MPDs, each of which has references to one or more auxiliary media contents. A departure point can be configured during playback of the main MPD. The departure point can be a point in time at which the auxiliary media segment begins playback by leaving the main media segment. In some embodiments, the departure point can be located before the start of the main media segment or the current auxiliary media segment. This may be referred to as pre-roll playback. In some embodiments, the departure point can be located at the end of the current auxiliary media segment or the main media segment. This may be referred to as end-roll playback. In some embodiments, the departure point can be any point in time during playback of the main media segment or the current media segment. This may be referred to as mid-roll playback. In some embodiments, an offset can be used to indicate mid-roll playback, the offset indicating the departure point from the currently available start time of the main media segment.
[0035] The rejoin point during playback can also be configured. In some embodiments, the rejoin point may be located at the end of playback of one or more secondary media segments. In some embodiments, the rejoin point may be located at the live edge of the main media segment. In some embodiments, the rejoin point may be located at the exit point when the main media segment stops. In some embodiments, the rejoin point may be located after a specific time period starting from the exit point when the main media segment stops.
[0036] In embodiments that stack one or more auxiliary MPDs (i.e., to play one or more MPDs sequentially), the main MPD can support multiple stacking modes. These stacking modes can execute or process MPDs in a specific order or method and can be referred to as stacking operations. The first stacking mode can be a "one-way" mode. In this stacking mode, after the MPD of the last URL has been played, the MPD of the first URL in the stack (the main MPD) is played. In some embodiments, multiple MPDs in an MPD stack including the main MPD and auxiliary MPDs can be played in the order in which they will be presented. As an example of a one-way mode, MPD1→MPD2→…→MPDn→MPD1, where MPDn is the nth MPD, the main MPD starts from n=0 and the auxiliary MPDs start from n>0.
[0037] The second stacking mode can be a "play-once" mode. In play-once mode, each URL's MPD in the stack is played only once. When returning to the stack, if the URL has already been played, the links and / or stacking are no longer considered. As an example of play-once mode, MPD1→MPD2→MPD3→MPD2→MPD1, where MPDn is the nth MPD, the main MPD starts from n=0 and the auxiliary MPDs start from n>0. The third stacking mode can be a "play-everytime" mode. In play-everytime mode, each auxiliary descriptor (also called a link descriptor) can be re-evaluated at each stack level, regardless of the stack's playback. As an example of play-everytime mode, MPD1→MPD2→MPD3→MPD2→MPD3→MPD2→MPD3, where MPDn is the nth MPD, the main MPD starts from n=0 and the auxiliary MPDs start from n>0.
[0038] According to one aspect of this disclosure, auxiliary MPD support for the primary MPD can be implemented using signaling and employing basic or supplementary descriptors. Signaling descriptors can be used at the MPD level.
[0039] Table 1 - Descriptor Semantics
[0040]
[0041] According to one aspect of this disclosure, auxiliary MPD support can be implemented using MPD events. In this embodiment, event stream semantics can be used.
[0042] Table 2 - Event Flow Semantics
[0043]
[0044] Table 3 - Event Semantics
[0045]
[0046]
[0047] According to one aspect, since alternative MPDs need to be downloaded before the event's presentation time, the event scheme can be an on_receive scheduling mode. In some embodiments, event instances can be repeated across different time periods. In particular, it is desirable to have a pre-roll overlay during playback in any time period. If only one pre-roll overlay is needed even if the player plays multiple time periods (i.e., the pre-roll overlay at the start of playback in the first time period), then an equivalent rule can be applied to all event instances across time periods representing that pre-roll overlay.
[0048] Embodiments of this disclosure relate to a method for signaling auxiliary media presentations from a main media presentation defined in an MPD, for inserting pre-roll, mid-roll, and end-roll auxiliary media content into the media presentation, wherein auxiliary MPD URLs, departure and re-entry times, and stacking operations between different levels of auxiliary MPDs are signaled. In some embodiments, the main content may depart at the beginning, middle, or end before its playback begins. In some embodiments, after playing auxiliary content or a specific duration thereof, the player may be instructed to resume playback of the main content from its missed point, from the current moment, or at any point in between. When an auxiliary MPD sequence exists, various supported stacking operation modes may also be signaled.
[0049] In some embodiments, auxiliary MPD support can be signaled using a basic or supplementary descriptor at the MPD level. In some embodiments, the basic or supplementary descriptor includes information required for leaving and rejoining the main media content playback, as well as the auxiliary MPD URL.
[0050] In some embodiments, auxiliary MPD support can be signaled using MPD events. These MPD events include all the information required to leave and rejoin the main media content playback, as well as the auxiliary MPD URL. Furthermore, in some embodiments, based on the playback of the auxiliary media content, equivalent and non-equivalent events can be repeated at various time intervals.
[0051] Figure 4 An exemplary flowchart is shown for a process 400 of signaling a chain of auxiliary media content, including pre-roll media content, middle-roll media content and end-roll media content, in an HTTP-based Dynamic Adaptive Streaming (DASH) main media stream.
[0052] In operation 410, one or more auxiliary descriptors may be received. In an embodiment, each of the one or more auxiliary descriptors may include a Uniform Resource Locator (URL) referencing one or more auxiliary MPDs and a stacking mode value indicating a stacking operation supported by the primary DASH media stream.
[0053] In some embodiments, the stacking mode value may include a first stacking mode value, which may indicate that one or more auxiliary media segments in the stack are played in a loop or sequentially. A second stacking mode value may indicate that one or more auxiliary media segments in the stack are played only once. A third stacking mode value may indicate that one or more auxiliary descriptors are evaluated at each level of the stack. As an example, the first stacking mode value may be "oneWay", the second stacking mode value may be "playOnce", and the third stacking mode value may be "playEverytime".
[0054] In some embodiments, one or more auxiliary descriptors further include departure and rejoin information. The departure information may include a first value for playing one or more auxiliary media segments. The first value may be relative to the MPD's Availability Start Time (AST). In some embodiments, the departure information may include a first departure value indicating that one or more auxiliary media segments should be played immediately upon retrieval. For example, the first departure value may be 0. A second departure value may indicate that one or more auxiliary media segments should be played at the end of the current MPD, where the current MPD may be the main MPD or one of one or more auxiliary MPDs. For example, the second departure value may be end. A third departure value may indicate that one or more auxiliary media segments should be played at a specific offset from the MPD's Availability Start Time. As an example, the third departure value may be an offset time.
[0055] Rejoin information may include a second value for rejoining the primary MPD. A first rejoin value may indicate returning to the primary MPD at the departure time from the primary MPD to one or more auxiliary MPDs. The first rejoin value can be 0. A second rejoin value may indicate returning to the primary MPD at the end of one or more auxiliary MPDs. The second rejoin value can be "end". A third rejoin value may indicate returning to the live edge of the primary MPD, and can be "live". A fourth rejoin value may indicate returning to the primary MPD at an offset from the departure time from the primary MPD to one or more auxiliary MPDs. The fourth rejoin value can be a specific offset time relative to the MPD AST.
[0056] In operation 415, one or more secondary media segments can be retrieved based on URLs referenced in each of the one or more secondary descriptors. These one or more secondary media segments can be independent of one or more primary DASH media segments.
[0057] In operation 420, one or more secondary media segments and one or more primary DASH media segments can be played from the Media Source Extension (MSE) source buffer based on one or more secondary descriptors and stacking mode values.
[0058] In some embodiments, one or more auxiliary descriptors may be signaled in the base descriptor or the supplementary descriptor at the MPD level. In some embodiments, one or more auxiliary descriptors may be signaled as MPD events. MPD events may have an event scheme with an on_receive scheduling mode. In some embodiments, MPD events may have equivalent rules applicable to all instances of the MPD event. In some embodiments, MPD events may have equivalent rules applicable to a specific instance of the MPD event. MPD events may include leave information, rejoin information, and stack mode values.
[0059] although Figure 4 An example block of process 400 is shown, but in embodiments, process 400 may include... Figure 4 The blocks shown are compared to additional blocks, fewer blocks, different blocks, or blocks with different arrangements. In an embodiment, any block of process 400 can be combined or arranged in any number or order as needed. In an embodiment, two or more blocks of process 400 can be executed in parallel.
[0060] The above-described techniques can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media, or implemented by one or more hardware processors with a specific configuration. For example, Figure 5 A computer system 500 suitable for implementing various embodiments is shown.
[0061] Computer software can be coded using any suitable machine code or computer language. Any suitable machine code or computer language can be assembled, compiled, linked, or similarly processed to create code containing instructions that can be executed directly by the computer's central processing unit (CPU), graphics processing unit (GPU), or through interpretation, microcode execution, etc.
[0062] The instructions can be executed on various types of computers or their components, including personal computers, tablets, servers, smartphones, gaming devices, and Internet of Things devices.
[0063] Figure 5The components of the computer system 500 shown are exemplary in nature and are not intended to impose any limitation on the scope or functionality of computer software implementing embodiments of this disclosure. The configuration of the components should also not be construed as having any dependencies or requirements relating to any one or a combination of components shown in the exemplary embodiments of the computer system 500.
[0064] Computer system 500 may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movement), audio input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input. The human-machine interface device may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images acquired from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0065] The input human-machine interface device may include one or more of the following (only one of each is shown): keyboard 501, mouse 502, touchpad 503, touch screen 510, joystick 505, microphone 506, scanner 508, camera 507.
[0066] Computer system 500 may also include certain human-machine interface (HMI) output devices. Such HMI output devices may, for example, stimulate the senses of one or more human users through tactile output, sound, light, and smell / taste. Such HMI output devices may include tactile output devices (e.g., haptic feedback of touchscreen 510, or joystick 505, but may also be haptic feedback devices that are not input devices), audio output devices (e.g., speakers 509, headphones), visual output devices (e.g., screen 510 including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input functionality, each with or without haptic feedback functionality—some of these screens are capable of outputting two-dimensional or more three-dimensional visual outputs via devices such as stereoscopic image output, virtual reality glasses, holographic displays, and smoke boxes), and printers.
[0067] The computer system 500 may also include human-machine-accessible storage devices and associated media, such as optical media including CD / DVD ROM / RW 520 with media such as CD / DVD 511, finger drives 522, removable hard disk drives or solid-state drives 523, conventional magnetic media such as magnetic tapes and floppy disks, and devices based on dedicated ROM / ASIC / PLD such as security dongles.
[0068] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.
[0069] Computer system 500 may also include an interface 599 for connecting one or more communication networks 598. Network 598 may be, for example, wireless, wired, or optical. Network 598 may further be a local area network, wide area network, metropolitan area network, vehicle and industrial network, real-time network, latency-tolerant network, etc. Examples of network 598 include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANbus, etc. Some networks 598 typically require an external network interface adapter (e.g., a USB port of computer system 500) to connect to certain general-purpose data ports or peripheral buses (550 and 551); other network interfaces are typically integrated into the core of computer system 500 by connecting to a system bus (e.g., an Ethernet interface in a PC computer system or a cellular network interface in a smartphone computer system). Computer system 500 can use any of these networks 598 to communicate with other entities. Such communication can be one-way receiving (e.g., broadcast television), one-way transmitting (e.g., CANbus connected to certain CANbus devices), or bidirectional, such as connecting to other computer systems using a local area or wide area digital network. As mentioned above, certain protocols and protocol stacks can be used on each of these networks and network interfaces.
[0070] The aforementioned human-machine interface device, human-machine accessible storage device, and network interface can be attached to the kernel 540 of the computer system 500.
[0071] The core 540 may include one or more central processing units (CPUs) 541, graphics processing units (GPUs) 542, image adapters 517, dedicated programmable processing units 543 in the form of field-programmable gate areas (FPGAs), hardware accelerators 544 for certain tasks, etc. These devices, along with read-only memory (ROM) 545, random access memory 546, and internal mass storage 547 such as internal non-user-accessible hard disk drives (SDs), etc., may be connected via a system bus 548. In some computer systems, the system bus 548 may be accessed via one or more physical connectors to allow for expansion with additional CPUs, GPUs, etc. Peripheral devices may be directly connected to the core's system bus 548 or connected via a peripheral bus 551. Peripheral bus architectures include PCI, USB, etc.
[0072] CPU 541, GPU 542, FPGA 543, and accelerator 544 can execute certain instructions, which can be combined to form the aforementioned computer code. This computer code can be stored in ROM 545 or RAM 546. Transient data can also be stored in RAM 546, while permanent data can be stored, for example, in internal mass storage 547. Fast storage and retrieval to any storage device can be achieved using a cache, which can be closely associated with one or more CPUs 541, GPU 542, mass storage 547, ROM 545, RAM 546, etc.
[0073] Computer-readable media may have computer code thereon for performing various computer-implemented operations. The media and computer code may be media and computer code specifically designed and constructed for the purposes of this disclosure, or the media and computer code may be of a type known and available to those skilled in the art of computer software.
[0074] By way of example and not limitation, software contained in one or more tangible computer-readable media can be executed by one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) to enable a computer system 500 having an architecture, particularly a kernel 540, to provide functionality. Such computer-readable media may be media associated with user-accessible mass storage as described above, and some non-transitory memory of the kernel 540, such as internal kernel mass memory 547 or ROM 545. Software implementing the various embodiments of this disclosure may be stored in such means and executed by the kernel 540. Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software may cause the kernel 540, particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to perform specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM 546 and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system may be made functional by logic hard-wired or otherwise embodied in the circuitry (e.g., accelerator 544), which may replace or operate with the software to perform the specific process or a specific portion of the specific process described herein. Where appropriate, references to software may include logic, and vice versa. Where appropriate, references to computer-readable media may include circuitry storing software for execution (e.g., integrated circuits (ICs)), circuitry embodying logic for execution, or both. This disclosure includes any suitable combination of hardware and software.
[0075] While several exemplary embodiments have been described in this disclosure, there are variations, substitutions, and various equivalent embodiments that fall within the scope of this disclosure. Therefore, it should be understood that those skilled in the art will be able to design numerous systems and methods that, although not expressly shown or described herein, embody the principles of this disclosure and are therefore within its spirit and scope.
Claims
1. A method for signaling chained auxiliary media content on a dynamically adaptive streaming DASH media stream based on the Hypertext Transfer Protocol, the method being executed by at least one processor, characterized in that, The method includes: Receive one or more auxiliary descriptors, wherein each of the one or more auxiliary descriptors includes a Uniform Resource Locator URL and a stacking mode value, the URL referencing one or more auxiliary Media Presentation Descriptions (MPDs), and the stacking mode value indicating a stacking operation supported by the main DASH media stream; The stacking mode value includes one of the following: The first stacking mode value indicates whether one or more auxiliary media segments in the stack will play in a loop or sequentially. A second stacking mode value indicates that the auxiliary media segment in the one or more auxiliary media segments in the stack is played only once; and The third stacking mode value indicates the evaluation of the auxiliary descriptors in the one or more auxiliary descriptors at each level of the stack; One or more secondary media segments are retrieved based on the URLs referenced in the various secondary descriptors, wherein the one or more secondary media segments are independent of one or more primary DASH media segments; and Based on the one or more auxiliary descriptors and the stacking mode value, the MSE source buffer is expanded from the media source to play the one or more auxiliary media segments and the one or more main DASH media segments at least once in at least one order.
2. The method according to claim 1, characterized in that, The one or more auxiliary descriptors also include: Departure information, wherein the departure information includes a first value for playing the one or more auxiliary media segments, wherein the first value is relative to the MPD availability start time AST of the main MPD; and Re-joining information, wherein the re-joining information includes a second value for re-joining the master MPD.
3. The method according to claim 2, characterized in that, The departure information includes one of the following: The first departure value indicates that the one or more auxiliary media segments should be played immediately upon retrieval; The second departure value indicates that the one or more auxiliary media segments should be played at the end of the current MPD, wherein the current MPD is the main MPD or one of the one or more auxiliary MPDs; and The third departure value indicates the playback of the one or more auxiliary media segments at a specific offset from the start time of the MPD availability.
4. The method according to claim 2, characterized in that, The rejoining information includes one of the following: The first rejoin value indicates the return to the main MPD at the departure time from the main MPD to the one or more auxiliary MPDs; The second rejoin value indicates a return to the main MPD at the end of the one or more auxiliary MPDs; The third rejoin value indicates a return to the live edge of the main MPD; as well as The fourth rejoin value indicates the return to the main MPD at the offset of the departure time from the main MPD to the one or more auxiliary MPDs.
5. The method according to any one of claims 1 to 4, characterized in that, The one or more auxiliary descriptors are signaled in the basic descriptor at the MPD level or in the supplementary descriptor at the MPD level.
6. The method according to any one of claims 1 to 4, characterized in that, Send the one or more auxiliary descriptors as MPD events using signals.
7. The method according to claim 6, characterized in that, The MPD event has an event scheme with an on_receive scheduling mode.
8. The method according to claim 6, characterized in that, The MPD event has equivalent rules that apply to all instances of the MPD event.
9. The method according to claim 6, characterized in that, The MPD event has equivalent rules that apply to a specific instance of the MPD event.
10. The method according to claim 6, characterized in that, The MPD event includes departure information, rejoin information, and the stacking mode value.
11. An apparatus for signaling chained auxiliary media content on a dynamically adaptive streaming DASH media stream based on the Hypertext Transfer Protocol, characterized in that, The device includes: At least one memory configured to store computer program code; At least one processor is configured to access and operate according to the instructions of the computer program code, the computer program code comprising: A receiving code is configured to cause the at least one processor to receive one or more auxiliary descriptors, wherein each of the one or more auxiliary descriptors includes a Uniform Resource Locator URL and a stacking mode value, the URL referencing one or more auxiliary Media Presentation Descriptions (MPDs) and the stacking mode value indicating a stacking operation supported by the main DASH media stream. The stacking mode value includes one of the following: The first stacking mode value indicates whether one or more auxiliary media segments in the stack will play in a loop or sequentially. A second stacking mode value indicates that the auxiliary media segment in the one or more auxiliary media segments in the stack is played only once; and The third stacking mode value indicates the evaluation of the auxiliary descriptors in the one or more auxiliary descriptors at each level of the stack; Retrieval code, configured to cause the at least one processor to retrieve one or more secondary media segments based on the URLs referenced in the respective secondary descriptors of the one or more secondary descriptors, wherein the one or more secondary media segments are independent of one or more primary DASH media segments; and Playback code configured to cause the at least one processor to expand the MSE source buffer from the media source based on the one or more auxiliary descriptors and the one or more main DASH media segments at least once in at least one order.
12. The apparatus according to claim 11, characterized in that, The one or more auxiliary descriptors also include: Departure information, wherein the departure information includes a first value for playing the one or more auxiliary media segments, wherein the first value is relative to the MPD availability start time AST of the main MPD; and Re-joining information, wherein the re-joining information includes a second value for re-joining the master MPD.
13. The apparatus according to any one of claims 11 to 12, characterized in that, The one or more auxiliary descriptors are signaled in the basic descriptor at the MPD level or in the supplementary descriptor at the MPD level.
14. The apparatus according to any one of claims 11 to 12, characterized in that, Send the one or more auxiliary descriptors as MPD events using signals.
15. The apparatus according to claim 14, characterized in that, The MPD event has equivalent rules that apply to all instances of the MPD event.
16. A non-transitory computer-readable medium storing instructions, characterized in that, The instructions include one or more instructions that, when executed by one or more processors of a device for signaling chained auxiliary media content comprising pre-roll media content, middle-roll media content and end-roll media content in an HTTP-based dynamically adaptive streaming DASH main media stream, cause the one or more processors to implement the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Content-Specific Identification and Timing Behavior in Dynamic Adaptive Streaming over Hypertext Transfer Protocol
US20140013003A1