Method and apparatus for dynamic DASH picture-in-picture streaming

The proposed method addresses the lack of explicit picture-in-picture signaling in DASH by merging main and picture-in-picture streams and dynamically updating their characteristics within the DASH framework, resulting in enhanced streaming adaptability and user experience.

JP7697040B2Active Publication Date: 2025-06-23TENCENT AMERICA LLC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023561377
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-09-21
Filing Date
2022-09-23
Publication Date
2025-06-23
Estimated Expiration
2042-09-23

AI Technical Summary

Technical Problem

Existing Dynamic Adaptive Streaming over HTTP (DASH) technologies lack explicit and interoperable methods for signaling picture-in-picture media content, limiting the ability to dynamically update the position, size, and resolution of picture-in-picture streams during streaming sessions.

Method used

The method involves determining the presence of main and picture-in-picture video streams based on role values, merging these streams using a pre-selection descriptor, and dynamically updating the media presentation descriptor (MPD) to adjust the position, size, and resolution of the picture-in-picture stream.

Benefits of technology

This solution enables efficient and dynamic signaling of picture-in-picture content within DASH streaming, allowing for real-time adjustments to the position, size, and resolution of picture-in-picture streams, thereby enhancing user experience and adaptability to changing media presentations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007697040000001
    Figure 0007697040000001
  • Figure 0007697040000002
    Figure 0007697040000002
  • Figure 0007697040000003
    Figure 0007697040000003
Patent Text Reader

Abstract

A method and apparatus may be provided for dynamically signaling picture-in-picture video during media streaming, which may include determining whether video data includes a first main video stream and a second picture-in-picture video stream based on a first role value associated with the first main video stream and a second role value associated with the second picture-in-picture video stream, determining a pre-selection descriptor indicating that the second picture-in-picture video stream is selected to be signaled with the first main video stream in the DASH media streaming, merging the first main video stream and the second picture-in-picture video stream as a combined video stream using the pre-selection descriptor, and updating the pre-selection descriptor with an updated media presentation descriptor (MPD).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority based on U.S. Provisional Patent Application No. 63 / 252,398, filed on October 5, 2021, and U.S. Patent Application No. 17 / 949,528, filed on September 21, 2022, the contents of which are hereby incorporated by reference in their entirety.

[0002] Embodiments of the present disclosure relate to streaming media content, and more particularly, to streaming picture - in - picture content via Dynamic Adaptive Streaming over hypertext transfer protocol (DASH) according to the Moving Picture Experts Group (MPEG) hypertext transfer protocol.

Background Art

[0003] MPEG DASH provides a standard for streaming media content over an IP network. In MPEG DASH, various contents can be described by a Media Presentation Descriptor (MPD), which is a DASH manifest, but explicit picture - in - picture signaling in DASH cannot be provided. Furthermore, the implicit methods in the prior art cannot provide an interoperable method or solution for picture - in - picture signaling either.

[0004] Therefore, there is a need for a method for delivering picture - in - picture media streaming using explicit extensions and existing DASH standards.

Summary of the Invention

Means for Solving the Problems

[0005] The present disclosure addresses one or more technical problems. The present disclosure includes a method, a process, an apparatus, and a non-transitory computer-readable medium for implementing picture-in-picture media content using DASH streaming. Further, embodiments of the present disclosure are also related to dynamically updating the position, size, resolution, etc. of picture-in-picture media content during a streaming session.

[0006] Embodiments of the present disclosure may provide a method for dynamically signaling a picture-in-picture video during dynamic adaptive HTTP streaming (DASH) media streaming. The method may be executed by a processor and may include determining whether video data includes a first main video stream and a second picture-in-picture video stream based on a first role value associated with the first main video stream and a second role value associated with the second picture-in-picture video stream; determining a pre-selection descriptor indicating that a second picture-in-picture video stream is selected to be signaled with the first main video stream in DASH media streaming; merging the first main video stream and the second picture-in-picture video stream as a combined video stream using the pre-selection descriptor; and updating the pre-selection descriptor using an updated media presentation descriptor (MPD).

[0007] Embodiments of the present disclosure may provide an apparatus for dynamically signaling picture-in-picture video during dynamic adaptive HTTP streaming (DASH) media streaming. The apparatus may include at least one memory configured to store computer program code and at least one processor configured to access the computer program code and operate as directed by the computer program code. The program code may include first determination code for causing at least one processor to determine whether video data includes a first main video stream and a second picture-in-picture video stream based on a first role value associated with the first main video stream and a second role value associated with the second picture-in-picture video stream, second determination code for causing at least one processor to determine a pre-selection descriptor indicating that a second picture-in-picture video stream is selected to be signaled with a first main video stream in DASH media streaming, first grouping code for causing at least one processor to merge a first main video stream and a second picture-in-picture video stream as a combined video stream using the pre-selection descriptor, and first update code for causing at least one processor to update the pre-selection descriptor using an updated media presentation descriptor (MPD).

[0008] Embodiments of the present disclosure may provide a non-transitory computer-readable medium storing instructions. When the instructions are executed by one or more processors of a device for dynamically signaling a picture-in-picture video during dynamic adaptive HTTP streaming (DASH) media streaming, the instructions cause the one or more processors to determine whether video data includes a first main video stream and a second picture-in-picture video stream based on a first role value associated with the first main video stream and a second role value associated with the second picture-in-picture video stream, cause the one or more processors to determine a pre-selection descriptor indicating that the second picture-in-picture video stream is selected to be signaled with the first main video stream in DASH media streaming, cause the one or more processors to merge the first main video stream and the second picture-in-picture video stream as a combined video stream using the pre-selection descriptor, and cause the one or more processors to update the pre-selection descriptor using an updated media presentation descriptor (MPD). The instructions may include one or more instructions. [1] Further features, properties, and various advantages of the subject matter of the present disclosure will become more apparent from the following detailed description and the accompanying drawings.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Mode for Carrying Out the Invention

[0010] The proposed features described below may be used separately or combined in any order. Further, the embodiments may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.

[0011] FIG. 1 illustrates a simplified block diagram of a communication system 100 according to an embodiment of the present disclosure. The communication system 100 may include at least two terminals 102 and 103 interconnected via a network 105. For unidirectional data transmission, the first terminal 103 may code video data at a local location for transmission to the other terminal 102 via the network 105. The second terminal 102 may receive the coded video data of the other terminal from the network 105, decode the coded data, and display the restored video data. Unidirectional data transmission may be common in media providing applications and the like.

[0012] FIG. 1 illustrates a second pair of terminals 101 and 104 provided to support two-way transmission of coded video that may occur, for example, during a video conference. For two-way data transmission, each terminal 101 and 104 may code video data captured at a local location for transmission to the other terminal via the network 105. Each terminal 101 and 104 may also receive the coded video data transmitted by the other terminal, decode the coded data, and display the restored video data on a local display device.

[0013] In FIG. 1, terminals 101, 102, 103, and 104 can be exemplified as a server, a personal computer, and a smartphone, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure can be applied to uses involving a laptop computer, a tablet computer, a media player, and / or dedicated video conferencing equipment. Network 105 represents any number of networks that carry encoded video data among terminals 101, 102, 103, and 104, including, for example, wired and / or wireless communication networks. The communication network 105 can exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of network 105 may not be important for the operation of the present disclosure, unless otherwise described herein below.

[0014] FIG. 2 shows, as an example, the arrangement of video encoders and decoders in a streaming environment. Embodiments can be applicable to other video-related uses, including, for example, video conferencing, digital TV, and further including storage of compressed video on digital media such as CDs, DVDs, and memory sticks.

[0015] A streaming system may include a capture subsystem 203 that can include a video source 201, such as a digital camera, for creating, for example, an uncompressed video sample stream 213. The sample stream 213 may be emphasized as a high data volume when compared to an encoded video bitstream and can be processed by an encoder 202 coupled to the video source 201. The encoder 202 may include hardware, software, or a combination thereof to enable or implement aspects of the embodiments, as will be described in detail below. The encoded video bitstream 204 may be emphasized as a lower data volume compared to the sample stream and can be stored in a streaming server 205 for future use. One or more streaming clients 212 and 207 may access the streaming server 205 to obtain encoded video bitstreams 208 and 206, which may be copies of the encoded video bitstream 204. Client 212 may include a video decoder 211 that decodes an incoming copy of the encoded video bitstream 208 and creates an outgoing video sample stream 210 that can be rendered on a display 209 or other rendering device. In some streaming systems, the encoded video bitstreams 204, 206, and 208 may be encoded according to a particular video coding / compression standard. Examples of these standards have been mentioned above and will be further described herein.

[0016] Figure 3 shows a sample DASH processing model 300, such as a sample client architecture for processing DASH events and CMAF events. In the DASH processing model 300, requests for client media segments (e.g., advertisement media segments and live media segments) can be based on the addresses described in the manifest 303. The manifest 303 also describes metadata tracks that the client can access the segments of, parse them, and further send them to the application 301.

[0017] The manifest 303 includes MPD events or events, and the in-band event and "moof" parser 306 can parse MPD event segments or event segments and append the event segments to the event and metadata buffer 330. Also, the in-band event and "moof" parser 306 can fetch media segments and append them to the media buffer 340. The event and metadata buffer 330 can send event and metadata information to the event and metadata synchronizer and dispatcher 335. The event and metadata synchronizer and dispatcher 335 can dispatch specific events to the DASH player control, selection, and heuristic logic 302 and dispatch application-related events and metadata tracks to the application 301.

[0018] According to some embodiments, the MSE may include a pipeline that includes a file format parser 350, a media buffer 340, and a media decoder 345. The MSE 320 is a logical buffer for media segments, and the media segments can be tracked and ordered based on the presentation time of the media segments. The media segments can include, but are not limited to, media segments associated with an advertisement MPD and live media segments associated with a live MPD. Each media segment may be added or appended to the media buffer 340 based on the time stamp offset of the media segment, and the time stamp offset may be used to order the media segments within the media buffer 340.

[0019] Embodiments of the present application can be directed to constructing a media source extension (MSE) buffer from two or more non-linear media sources using an MPD chain, where the non-linear media sources can be an advertisement MPD and a live MPD, and the file format parser 350 can be used to process different media and / or codecs used by the live media segments included in the live MPD. In some embodiments, the file format parser can issue a change type based on the codec, profile, and / or level of the live media segment.

[0020] As long as the media segment exists within the media buffer 340, the event and metadata buffer 330 maintains the corresponding event segment and metadata. The sample DASH processing model 300 may include a time-stamped metadata tracking parser 325 for maintaining tracking of metadata associated with in-band events and MPD events. According to FIG. 3, the MSE 320 includes only a file format parser 350, a media buffer 340, and a media decoder 345. The event and metadata buffer 330 and the event and metadata synchronizer and dispatcher 335 are not native to the MSE 320, preventing the MSE 320 from natively processing events and sending them to the application.

[0021] According to one aspect, the MPD is a media presentation description that may include a media presentation in a hierarchical structure. The MPD can include one or more periodic sequences, and each period can include one or more adaptation sets. Each adaptation set within the MPD can include one or more representations, and each representation can include one or more media segments. These one or more media segments carry the actual media data and related metadata that are encoded, decoded, and / or played back. According to embodiments of the present disclosure, an overlay video can refer to a single or combined experience stream or can refer to a pip video stream.

[0022] FIG. 4 is an exemplary diagram showing a picture-in-picture media presentation 400.

[0023] As shown in FIG. 4, the main picture 405 captures the entire screen, and the overlay picture (picture-in-picture 410) captures a small area of the screen that covers the corresponding area of the main picture. The coordinates of the picture-in-picture (pip) are indicated by x, y, height, and width, and these parameters define the position and size of the pip relative to the main picture coordinates.

[0024] In connection with streaming, the main video (or main picture 405) and the pip video (or picture-in-picture 410) can be delivered as two separate streams. If there are independent streams, they can be decoded by separate decoders and then combined together for rendering. As another embodiment, if the video codec used for the main video supports stream merging, the pip video stream is combined with the main video stream. In some embodiments, the pip video stream can replace the main video streamed in the cover area of the main video with the pip video. Next, the single and / or combined streams are sent to the decoder for decoding and rendering.

[0025] According to one aspect, when the main video and the pip video are related (such as a sign video at the corner of the screen indicating what is being said on the main screen), it may be necessary to change the position of the pip video within the main video. As an example, when the area of the pip picture changes from the background to the foreground, it may be necessary to change the position of the pip picture. Similarly, it may be necessary to change the resolution of the pip picture. Also, during a certain duration of the media presentation, the pip picture may not be required. These dynamic changes related to the streaming of these pictures-in-picture are not handled by DASH. Furthermore, DASH cannot handle this dynamic change in position and resolution without changing the main video stream and the pip video stream.

[0026] Embodiments of the present disclosure relate to solving the above technical problems. According to one embodiment, picture-in-picture is delivered in DASH when the main video and the pip video can be independently decoded and then combined. According to other embodiments, picture-in-picture is delivered in DASH when the main video and the pip video can be combined into a single stream before decoding and decoded together as a single stream (also referred to herein as a "combined stream").

[0027] Delivery of picture-in-picture in DASH when the main video and the pip video are decoded independently The Pip video stream and the main video stream can be identified using the DASH role scheme. According to one aspect, the main video stream can use the "main" role value, while the pip video stream can use a new value "pip" or "picture-in-picture" to identify the corresponding adaptation set. In some embodiments, since the overlaid picture is not necessarily a "signaled" video, the value of the role attribute (e.g., main or pip) is independent of the value "signaled". If so, both values can be used to signal the characteristics of the overlaid video.

[0028] The main video stream and the pip video stream may be grouped as a single experience using a preselection descriptor. In some embodiments, the pip video stream adaptation set may include a preselection descriptor that references the adaptation set of the main video stream.

[0029] The position of the PIP video can be signaled using DASH specifications, such as the 23009-1 Annex H Spatial Relationship Descriptor (SRD). This descriptor can be used to signal the position and size of the PIP video. The SRD descriptor enables the definition of x, y, width, and height with respect to a common coordinate system. This descriptor can be used in both the main and overlay video adaptation sets to define the relationship of the components to each other. Updates to the position and size can be achieved using one of the following mechanisms.

[0030] (1) When the position, size, and / or resolution are changed, the MPD introduces a new period and updates the value of the SRD to update the position, size, resolution, etc. of the picture-in-picture.

[0031] (1) Use of a metadata track containing coordinate information according to 23009-1 Annex H.

[0032] Delivery of picture-in-picture in DASH when the main video segment and the PIP video segment are merged with the video stream before decoding According to an embodiment, the main video stream and the pip video stream can be grouped as a single experience using a preselection element. In some embodiments, grouping the main video stream and the pip video stream may include identifying the pip video stream and the main video stream based on the DASH role scheme within the preselection element. The main video stream can use the value "main" for the role, while the pip video stream can use a new value "picture-in-picture" or "pip" to identify the corresponding adaptation set. It can be understood that the role value of "pip" or "picture-in-picture" may be used interchangeably or may have special instructions. As an example, the use of "pip" may indicate independent decoding followed by an overlay. As another example, the user of "picture-in-picture" can indicate grouping into a single experience followed by decoding.

[0033] According to the same or other embodiments, grouping the main video stream and the pip video stream may further include that a new value "pip" for the preselection@order can be defined to replace a part of the main video stream with an overlaid video stream. Further, in some embodiments, a new attribute, preselection@replacementRules, may be added to define the replacement rules. As an example, when the codec used is VVC, the @replacementRule may include sub-picture OD. The semantics of the @replacementRules attribute are codec-dependent.

[0034] An MPD update may be defined to insert a new Period and update the value of a Preselected element that includes @replacementRules. In one embodiment, the MPD inserts a new Period and updates the value of a Preselected element where a merge of the pip video stream with the main video stream may be defined.

[0035] The advantages of the present disclosure are only sophisticated necessary extensions to the DASH standard for dynamically and efficiently performing picture-in-picture signaling. In one embodiment, a new role value for "picture-in-picture" (or an appropriate version thereof) may be added to the DASH standard to indicate the presence of the main video stream and the pip video stream. In the same or other embodiments, a new @order value for "replacement" (or an appropriate version thereof) may be added to the DASH standard to indicate that the pip video stream may need to replace part of the main video stream. According to the same or other embodiments, a new attribute called @replacementRules may be added to the DASH standard to define one or more replacement rules based on the codecs of the main video stream and the pip video stream.

[0036] Embodiments of the present disclosure can relate to a method, system, and process for dynamically signaling the relationship of picture-in-picture video and its relationship to the main video in DASH streaming. When the picture-in-picture video and the main video are decoded independently, a role attribute with a special value can be used to signal the picture-in-picture video stream. In some embodiments, the main video can have a role value of "main". In some embodiments, a preselection descriptor can be used to couple the picture-in-picture video adaptation set to the main video adaptation set. In some embodiments, the position and size of the picture-in-picture video on the main video can be defined by SRD descriptors of both the main and picture-in-picture video adaptation sets. The position and size of the picture-in-picture video can be updated by using an MPD update and inserting a new period with a new SRD value. In other embodiments, a metadata track can be used to dynamically convey and / or update the position and size information.

[0037] Embodiments of the present disclosure can relate to a method, system, and process for dynamically signaling the relationship of a picture-in-picture video and a main video in DASH streaming when the picture-in-picture video and the main video can be merged before decoding. In some embodiments, a preselection element can be used to signal a group of main and picture-in-picture adaptation sets. A role attribute with a special value can be used to signal a picture-in-picture video stream while the main video has a role value of "main". In some embodiments, a new value of the attribute order can be used to signal a picture-in-picture application, and new attributes can be used to define how the two streams are merged before transmitting to a decoder. The position and size of the picture-in-picture can be updated by an MPD update that can insert a new period. In some embodiments, the attribute that defines the merge rule can be updated to reflect that a new region of the main video stream should be replaced with the picture-in-picture stream.

[0038] FIG. 5 is an exemplary flowchart of a process 500 for dynamically signaling a picture-in-picture video during media streaming.

[0039] In operation 510, it can be determined whether video data includes a first main video stream and a second picture-in-picture video stream based on role values associated with the first main video stream and the second picture-in-picture video stream.

[0040] In project 515, a first main video stream with a second picture-in-picture video stream may be merged as a single video stream based on a pre-selection descriptor, which may be associated with the second picture-in-picture video stream.

[0041] In some embodiments, whether the first main video stream and the second picture-in-picture video stream are independently decoded before grouping can be determined based on that a first role value associated with the first main video stream is a main value and a second role value associated with the second picture-in-picture video stream is a picture-in-picture value. Then, based on a pre-selection descriptor associated with the second picture-in-picture video stream that refers to an adaptation set in the first main video stream, the first main video content associated with the adaptation set is grouped with the second picture-in-picture video content as a single video stream. The grouping may further include signaling the position of the second picture-in-picture video content, the size of the second picture-in-picture video content, or the resolution of the second picture-in-picture video content using a spatial relationship descriptor.

[0042] In some embodiments, the first main video stream and the second picture-in-picture video stream can be identified based on that, before grouping, the first role value associated with the first main video stream is the main value and the second role value associated with the second picture-in-picture video stream is the picture-in-picture value. Then, an order value in a preselection descriptor for replacing a part of the first main video content with the second picture-in-picture video content may be defined, and one or more replacement rules in the preselection descriptor for replacing a part of the first main video content with the second picture-in-picture video content may be defined. The first main video content with the second picture-in-picture video content can be merged based on at least one or more replacement rules in the preselection descriptor, and the merge is performed before decoding a single video stream.

[0043] In step 520, the preselection descriptor can be updated using an updated media presentation descriptor. Based on the fact that the first main video stream and the second picture-in-picture video stream are independently decoded before grouping, the position or size of the second picture-in-picture video content can be updated in the spatial relationship descriptor based on the updated MPD. In some embodiments, the position or size of the second picture-in-picture video content can be updated in the spatial relationship descriptor based on a metadata track including coordinate information.

[0044] FIG. 5 shows exemplary blocks of process 500, but in embodiments, process 500 may include additional blocks, fewer blocks, different blocks, or blocks in a different arrangement than those shown in FIG. 5. In embodiments, any blocks of process 500 may be combined or arranged in any amount or order as needed. In embodiments, two or more of the blocks of process 500 may be executed in parallel.

[0045] The techniques described above may be implemented using computer-readable instructions and as computer software physically stored on one or more computer-readable media, or by one or more hardware processors specifically configured. For example, FIG. 6 shows a computer system 600 suitable for implementation of various embodiments.

[0046] Computer software can be coded using any suitable machine code or computer language that can be subject to mechanisms such as assembly, compilation, linking, etc. to create code including instructions that can be executed directly by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or through interpretation, execution of microcode, etc.

[0047] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, game devices, Internet of Things devices, etc.

[0048] The components shown in FIG. 6 with respect to computer system 600 are exemplary in nature and are not intended to suggest any limitation regarding the use or functionality scope of the computer software implementing embodiments of the present disclosure. Also, the configuration of the components should not be construed as having any dependencies or requirements related to any one or combination of the components shown in the exemplary embodiments of computer system 600.

[0049] The computer system 600 may include a specific human interface input device. Such a human interface input device may respond to input by one or more human users via, for example, tactile input (such as keystrokes, swipes, movement of a data glove), voice input (such as voice, clapping), visual input (such as gestures), or olfactory input. Using the human interface device, it is possible to capture specific media that is not necessarily directly related to conscious human input, such as voice (speech, music, ambient sound, etc.), images (such as scanned images, photographic images obtained from a still image camera), video (such as two-dimensional video, three-dimensional video including stereoscopic video, etc.).

[0050] The input human interface device may include one or more of a keyboard 601, a mouse 602, a trackpad 603, a touch screen 610, a joystick 605, a microphone 606, a scanner 608, and a camera 607 (only one of each is shown in the figure).

[0051] In addition, computer system 600 may include a specific human interface output device. Such a human interface output device can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by touch screen 610 or joystick 605, although there may also be tactile feedback devices that do not function as input devices), audio output devices (e.g., speaker 609, headphones, etc.), visual output devices (e.g., screens 610 including CRT screens, LCD screens, plasma screens, OLED screens, some of which have a touch screen input function and some do not, and some of which have a tactile feedback function and some do not, and some of these can output two-dimensional visual output or three-dimensional output through means such as stereographic output, virtual reality glasses, holographic displays, and smoke tanks), and may also include printers.

[0052] In addition, computer system 600 may also include a human-accessible memory device and optical media such as CD / DVD ROM / RW 620 with media such as CD / DVD 611, thumb drive 622, removable hard drive or solid state drive 623, legacy magnetic media such as tapes and floppy disks, and their associated media such as dedicated ROM / ASIC / PLD-based devices such as security dongles.

[0053] Also, as will be understood by those skilled in the art, the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not include a transmission medium, a carrier wave, or other transient signals.

[0054] In addition, computer system 600 may include an interface 699 to one or more communication networks 698. Network 698 can be, for example, wireless, wired, or optical. Network 698 can further be local, wide area, metropolitan, vehicular and industrial, real-time, delay-tolerant, etc. Examples of network 698 include, for example, local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial television, vehicular and industrial networks including CANBus, etc. A particular network 698 generally requires an external network interface adapter attached to a particular general-purpose data port or peripheral bus (650 and 651) (such as a USB port of computer system 600), and other networks are generally incorporated into the core of computer system 600 by attachment to the system bus as described later (such as an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks 698, computer system 600 can communicate with other entities. Such communication can be one-way reception only (such as broadcast TV), one-way transmission only (such as CANbus to a particular CANbus device), or two-way, such as communication to other computer systems using local area or wide area digital networks. Particular protocols and protocol stacks can be used in each of those networks and network interfaces as described above.

[0055] The aforementioned human interface devices, human-accessible storage devices, and network interfaces can be attached to the core 640 of computer system 600.

[0056] The core 640 may include one or more central processing units (CPUs) 641, a graphics processing unit (GPU) 642, a graphics adapter 617, a special programmable processing device in the form of a field programmable gate array (FPGA) 643, a hardware accelerator 644 for specific tasks, and the like. These devices may be connected via a system bus 648 together with a read-only memory (ROM) 645, a random access memory 646, an internal hard drive that is not accessible to the user, an internal mass storage such as an SSD 647. In some computer systems, it is possible to access the system bus 648 in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be directly attached to the system bus 648 of the core or can be attached via a peripheral bus 651. The architecture of the peripheral bus includes PCI, USB, and the like.

[0057] The CPU 641, GPU 642, FPGA 643, and accelerator 644 may execute specific instructions that can together constitute the aforementioned computer code. The computer code can be stored in the ROM 645 or the RAM 646. Migration data can be stored in the RAM 646, while persistent data can be stored, for example, in the internal mass storage 647. By using a cache memory that can be closely associated with one or more CPUs 641, GPUs 642, mass storage 647, ROM 645, RAM 646, etc., it is possible to enable fast storage and fast retrieval to any of the memory devices.

[0058] A computer-readable medium may have computer code for performing various computer-implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of the present disclosure or they may be of the kind well-known and available to persons having skill in the computer software arts.

[0059] As an example and not by way of limitation, a computer system 600 having the illustrated architecture, specifically a core 640, can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied within one or more tangible computer-readable media. Such computer-readable media can be associated with the user-accessible mass storage introduced above, as well as with specific storage of the core 640 of a non-transitory nature such as the core internal mass storage 647 or ROM 645. The software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core 640. The computer-readable media can include one or more memory devices or chips, depending on the specific requirements. The software can cause the core 640 and specifically the processors therein (including a CPU, GPU, FPGA, etc.) to define data structures stored in the RAM 646 and modify such data structures according to processes defined by the software, to execute the specific processes described herein, or specific portions of specific processes. In addition to or instead of this, the computer system can provide functionality as a result of logic embodied in a circuit (e.g., an accelerator 644) in a hard-wired or other manner, which can operate instead of or in conjunction with software to execute the specific processes described herein or specific portions of specific processes. Optionally, references to software can include logic and vice versa. Optionally, references to computer-readable media can include circuits (such as integrated circuits (ICs)) storing software for execution, circuits embodying logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0060] Although several exemplary embodiments of the present disclosure have been described, there are modifications, substitutions, and various alternative equivalents within the scope of the present disclosure. Therefore, it can be understood by those skilled in the art that although not explicitly illustrated or described herein, many systems and methods that embody the principles of the present disclosure and are thus within its spirit and scope can be devised.

Description of Reference Numerals

[0061] 100 Communication system 101 Terminal 102 Second terminal 103 First terminal 104 Terminal 105 Network 201 Video source 202 Encoder 203 Capture subsystem 204 Video bitstream 205 Streaming server 206 Encoded video bitstream 207 Streaming client 208 Encoded video bitstream 209 Display 210 Transmitted video sample stream 211 Video decoder 212 Streaming client, client 213 Sample stream 300 DASH processing model 301 Application 302 DASH player control, selection, and heuristic logic 303 Manifest 306 In-band event and "moof" parser 325 Time-specification metadata tracking parser 330 Event and metadata buffer 335 Event and metadata synchronizer and dispatcher 340 Media buffer 345 Media Decoder 350 File Format Parser 400 Picture-in-Picture Media Presentation 405 Main Picture 410 Picture-in-Picture 600 Computer System 601 Keyboard 602 Mouse 603 Track Pad 605 Joystick 606 Microphone 607 Camera 608 Scanner 609 Speaker 610 Touch Screen, Screen 617 Graphics Adapter 620 CD / DVD ROM / RW 622 Thumb Drive 623 Removable Hard Drive or Solid State Drive 640 Core 641 Central Processing Unit (CPU) 642 Graphics Processing Unit (GPU) 643 Field-Programmable Gate Array (FPGA) 644 Hardware Accelerator, Accelerator 645 Read Only Memory (ROM) 646 Random Access Memory 647 Internal Mass Storage, Mass Storage 648 System Bus 651 Peripheral Bus 698 Network 699 Interface 714 Network

Claims

1. A method for dynamically signaling picture-in-picture video during dynamic adaptive HTTP streaming (DASH) media streaming, the method being executed by one or more processors, the method comprising: determining whether the video data includes a first main video stream and a second picture-in-picture video stream based on a first role value associated with the first main video stream indicating that the first main video stream is the main video stream and a second role value different from the first role value associated with the second picture-in-picture video stream indicating that the second picture-in-picture video stream is the picture-in-picture video stream, the first role value and the second role value being included in a pre-selection descriptor; determining the pre-selection descriptor indicating that the second picture-in-picture video stream is selected to be signaled with the first main video stream in the DASH media streaming; merging the first main video stream and the second picture-in-picture video stream as a combined video stream using the pre-selection descriptor; and updating the pre-selection descriptor by inserting a new period having a spatial relationship descriptor (SRD descriptor) indicating the position and size of the second picture-in-picture video stream.

2. The step of merging the first main video stream and the second picture-in-picture video stream as the combined video stream comprises: The step of determining that the first main video stream and the second picture-in-picture video stream are independently decoded before the step of merging, based on the fact that the first role value associated with the first main video stream is the main value and the second role value associated with the second picture-in-picture video stream is the picture-in-picture value or a special value; The step of merging the first main video content and the second picture-in-picture video content associated with the adaptation set as the combined video stream, based on the preselection descriptor that refers to the adaptation set in the first main video stream; The method according to claim 1, comprising the step of signaling the position, the size, or the resolution of the second picture-in-picture video content in the second picture-in-picture video content using a spatial relationship descriptor.

3. The step of updating the preselection descriptor using the updated MPD comprises The method according to claim 2, comprising the step of updating the position or the size of the second picture-in-picture video content in the spatial relationship descriptor based on the updated MPD.

4. The step of updating the preselection descriptor using the updated MPD comprises The method according to claim 2, comprising the step of updating the position or the size of the second picture-in-picture video content in the spatial relationship descriptor based on a metadata track including coordinate information.

5. The step of merging the first main video stream and the second picture-in-picture video stream as the combined video stream comprises The step of identifying the first main video stream and the second picture-in-picture video stream before the step of merging, based on the fact that the first roll value associated with the first main video stream is the main value and the second roll value associated with the second picture-in-picture video stream is the picture-in-picture value or a special value; The step of defining an order value in the pre-selection descriptor to replace a part of the first main video content with the second picture-in-picture video content; The method according to claim 1, further comprising the step of defining one or more replacement rules in the pre-selection descriptor to replace a part of the first main video content with the second picture-in-picture video content.

6. The step of merging the first main video stream and the second picture-in-picture video stream as the combined video stream is The step of merging the first main video content and the second picture-in-picture video content based on at least the one or more replacement rules in the pre-selection descriptor, the step of merging being performed before decoding the combined video stream, the method according to claim 5.

7. The step of updating the pre-selection descriptor using the updated MPD is The method according to claim 5, further comprising the step of updating the one or more replacement rules in the pre-selection descriptor based on the updated MPD.

8. The method according to claim 7, wherein the one or more replacement rules are codec-dependent.

9. An apparatus for signaling picture-in-picture video during dynamic adaptive HTTP streaming (DASH) media streaming, configured to perform the method according to any one of claims 1 to 7.

10. A computer program for causing a computer to execute the method according to any one of claims 1 to 3, 5, and 6.

Citation Information

Patent Citations

  • Picture synthesis method and apparatus

    CN106649794A

  • Transmitting method and receiving method

    JP2020039174A

  • Signalling of substitution of video data unit in picture-in-picture region

    JP2023008947A

  • Method, device and medium for video processing - Patents.com

    JP2024534616A

  • Method and apparatus for video coding and decoding

    US20150304665A1