Picture-in-picture signaling method using preselected signaling with external code stream operation instructions

By introducing @interleaving and @order attributes into the DASH standard, allowing the decoder to define code stream operations and multiplexing instructions, the scalability and interoperability problems of the DASH standard in multimedia stream multiplexing and picture-in-picture applications are solved, and flexible processing and multiplexing of multimedia streams are achieved.

CN120077628APending Publication Date: 2025-05-30TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004457.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-07-09
Filing Date
2024-07-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing DASH standards lack scalability and interoperability when dealing with multimedia stream multiplexing, especially in picture-in-picture applications, where there is no effective solution for annotation and reuse of VVC sub-pictures.

Method used

By introducing new attributes of preselected elements in the DASH code stream, the decoder allows the decoder to define code stream operations and multiplexing instructions, and realizes flexible multiplexing and processing of multimedia segments.

Benefits of technology

Improves the scalability and interoperability of the DASH standard, providing a flexible solution to handle and multiplex multimedia streams, especially in picture-in-picture applications, supporting effective annotation and multiplexing of VVC sub-pictures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077628A_ABST
    Figure CN120077628A_ABST
Patent Text Reader

Abstract

A method includes: receiving a dynamic adaptive streaming (DASH) code stream based on a hypertext transfer protocol (HTTP); determining that the DASH code stream comprises a preselected element, wherein the preselected element is used for multiplexing a plurality of media segments included in the DASH code stream; analyzing the plurality of media segments from the code stream; multiplexing, by the DASH application, the plurality of segments using the preselected element and at least one policy associated with a decoder to generate a multiplexed code stream; and outputting the multiplexed code stream.
Need to check novelty before this filing date? Find Prior Art

Description

Cross - Reference to Related Applications

[0001] This application claims priority to U.S. Provisional Application No. 63 / 526,145, filed on July 11, 2023, U.S. Provisional Application No. 63 / 526,140, filed on July 11, 2023, and U.S. Application No. 18 / 767,475, filed on July 9, 2024, the disclosures of each of which are incorporated herein by reference. Technical Field

[0002] The present disclosure generally relates to signaling multiplexing instructions using pre - selected elements with external bitstream operation instructions. Background Art

[0003] Pre - selected elements are defined in Hyper - Text Transfer Protocol (HTTP) - based Dynamic Adaptive Streaming over HTTP (DASH) for providing a media experience by combining at least two streams. Current pre - selection designs define a specific set of methods for multiplexing received streams before providing them to at least one decoder.

[0004] Moving Picture Experts Group (MPEG) DASH provides a standard for streaming multimedia content over IP networks. The DASH standard provides a way to describe various contents and their relationships using pre - selected elements. However, current pre - selection designs define a specific and well - defined set of methods for multiplexing received streams before providing them to at least one decoder. Since each codec specification may require a different way to operate and multiplex at least two streams, with the introduction of each new codec, one or more of its methods need to be included in the DASH standard, which limits the scalability and usability of the DASH specification.

[0005] The DASH CDAM2 document is developing the use of pre - selected picture - in - picture signaling. However, it includes explicit signaling for sprite replacement.

[0006] Although the DASH standard provides a way to describe various contents and their relationships, it does not provide an interoperable solution for annotating VVC sprites for picture - in - picture applications. Picture - in - picture has many applications, from watching an alternative channel while watching the main channel, to adding a sign - language video for hearing - impaired viewers, i.e., a small video in the corner of the main video showing a person using sign language to convey audio information. Summary of the Invention

[0007] According to one aspect of the present disclosure, a method executed by at least one processor of a decoder includes: receiving a Dynamic Adaptive Streaming over HTTP (DASH) bitstream; determining that the DASH bitstream includes a preselected element for multiplexing a plurality of media segments included in the DASH bitstream; parsing the plurality of media segments from the bitstream; multiplexing, by a DASH application, the plurality of segments using the preselected element and at least one policy associated with the decoder to generate a multiplexed bitstream; and outputting the multiplexed bitstream.

[0008] According to one aspect of the present disclosure, a decoder includes: at least one memory configured to store program code; and at least one processor configured to read the program code and operate in accordance with the instructions of the program code, the program code including: receiving code configured to cause the at least one processor to receive a Dynamic Adaptive Streaming over HTTP (DASH) bitstream; first determination code configured to cause the at least one processor to determine that the DASH bitstream includes a preselected element for multiplexing a plurality of media segments included in the DASH bitstream; parsing code configured to cause the at least one processor to parse the plurality of media segments from the bitstream; multiplexing code configured to cause the at least one processor to multiplex, by a DASH application, the plurality of segments using the preselected element and at least one policy associated with the decoder to generate a multiplexed bitstream; and output code configured to cause the at least one processor to output the multiplexed bitstream.

[0009] According to one aspect of the present disclosure, a method executed by at least one processor of a decoder includes: processing a Dynamic Adaptive Streaming over HTTP (DASH) bitstream; wherein the DASH bitstream includes a preselected element for multiplexing a plurality of media segments included in the DASH bitstream, the plurality of media segments being parsed from the bitstream, and a DASH application multiplexes the plurality of segments using the preselected element and at least one policy associated with the decoder to generate a multiplexed bitstream.

[0010] Additional embodiments will be set forth in the following description, and in part will be obvious from the description, and / or may be learned by practice of the embodiments presented in the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0012] Figure 1It is a schematic diagram of an environment that can implement the methods, apparatuses, and systems described herein according to an embodiment.

[0013] Figure 2 is Figure 1 a block diagram of an example component of at least one device of.

[0014] Figure 3 It is a schematic diagram of an example client architecture for processing DASH and CMAF events according to an embodiment.

[0015] Figure 4 It is an example picture-in-picture schematic diagram according to an embodiment.

[0016] Figure 5 It is a flowchart of an example process for signaling multiplexing information according to an embodiment. Detailed Description of Specific Embodiments

[0017] The following detailed description of example embodiments refers to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements.

[0018] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the embodiments to the precise forms disclosed. Modifications and variations are possible in light of the foregoing disclosure, or may be obtained from practice of the embodiments. Additionally, at least one feature or component of one embodiment may be incorporated into (or combined with at least one feature of) another embodiment. Further, in the flowcharts and descriptions of operations provided below, it should be understood that at least one operation may be omitted, at least one operation may be added, at least one operation may be performed simultaneously (at least in part), and the order of at least one operation may be switched.

[0019] It is evident that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual specific control hardware or software code for implementing these systems and / or methods does not limit the embodiments. Thus, the operations and performance of the systems and / or methods are described herein without reference to specific software code—it should be understood that software and hardware may be designed based on the description herein to implement the systems and / or methods.

[0020] Although specific combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features can be combined in ways not specifically recited in the claims and / or not disclosed in the specification. Although each dependent claim listed below may directly depend on only one claim, the disclosure of possible implementations includes each dependent claim combined with every other claim in the set of claims.

[0021] Elements, acts, or instructions used herein should not be construed as critical or essential unless explicitly described as such. Additionally, as used herein, the articles "a" and "an" are intended to include at least one item and may be used interchangeably with "at least one." Where only one item is intended, the term "one" or similar language is used. Additionally, as used herein, the terms "has," "have," "having," "include," "including," etc. are intended to be open-ended terms. Additionally, the phrase "based on" is intended to mean "at least partially based on" unless otherwise explicitly stated. Additionally, expressions such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood to include only A, only B, or both A and B.

[0022] References to "an embodiment," "embodiments," or similar language throughout this specification mean that a particular feature, structure, or characteristic described in connection with the indicated embodiment is included in at least one embodiment of the present solution. Thus, the phrases "in an embodiment," "in embodiments," and similar language throughout this specification may, but do not necessarily, all refer to the same embodiment.

[0023] Furthermore, the described features, advantages, and characteristics of the present disclosure may be combined in any suitable manner in at least one embodiment. Based on the description herein, those skilled in the relevant art will recognize that the present disclosure may be practiced without at least one specific feature or advantage of a particular embodiment. In other instances, additional features and advantages may be recognized in certain embodiments that may not exist in all embodiments of the present disclosure.

[0024] Embodiments of the present disclosure provide a scalable method for operating and multiplexing signaling flows in DASH preselection. Embodiments of the present disclosure further provide a method for signaling a picture-in-picture experience using preselection, but allowing the codec to define stream operation and multiplexing instructions.

[0025] Figure 1is a schematic diagram of an environment 100 in accordance with an embodiment, in which the methods, apparatuses, and systems described herein can be implemented. As Figure 1 shown, the environment 100 can include a user device 110, a platform 120, and a network 130. The devices of the environment 100 can be interconnected by a wired connection, a wireless connection, or a combination of wired and wireless connections.

[0026] The user device 110 includes at least one device capable of receiving, generating, storing, processing, and / or providing information related to the platform 120. For example, the user device 110 can include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smartphone, a wireless phone, etc.), a wearable device (e.g., smart glasses or a smart watch), or a similar device. In some embodiments, the user device 110 can receive information from and / or send information to the platform 120.

[0027] The platform 120 includes at least one device as described elsewhere herein. In some embodiments, the platform 120 can include a cloud server or a group of cloud servers. In some embodiments, the platform 120 can be designed to be modular such that software components can be swapped in or out according to specific needs. In this way, the platform 120 can be easily and / or quickly reconfigured to have different uses.

[0028] In some embodiments, as shown, the platform 120 can be hosted in a cloud computing environment 122. It should be noted that although the embodiments described herein describe the platform 120 as being hosted in the cloud computing environment 122, in some embodiments, the platform 120 is not cloud-based (i.e., can be implemented outside of a cloud computing environment) or can be partially cloud-based.

[0029] The cloud computing environment 122 includes an environment that hosts the platform 120. The cloud computing environment 122 can provide services such as computing, software, data access, storage, etc., which do not require an end user (e.g., the user device 110) to know the physical location and configuration of the systems and / or devices of the hosted platform 120. As shown, the cloud computing environment 122 can include a set of computing resources 124 (collectively referred to as "computing resources 124" and individually referred to as "computing resource 124").

[0030] The computing resource 124 includes at least one personal computer, workstation computer, server device, or other type of computing and / or communication device. In some embodiments, the computing resource 124 may host the platform 120. Cloud resources may include computing instances executed in the computing resource 124, storage devices provided in the computing resource 124, data transmission devices provided by the computing resource 124, and the like. In some embodiments, the computing resource 124 may communicate with other computing resources 124 via a wired connection, wireless connection, or a combination of wired and wireless connections.

[0031] Further as Figure 1 shown, the computing resource 124 includes a set of cloud resources, such as at least one application program (APP) 124-1, at least one virtual machine (VM) 124-2, virtualized storage (VS) 124-3, at least one hypervisor (HYP) 124-4, and the like.

[0032] The application program 124-1 includes at least one software application program, which may be provided to the user device 110 and / or the platform 120, or accessed by the user device 110 and / or the platform 120. The application program 124-1 does not require the installation and execution of the software application program on the user device 110. For example, the application program 124-1 may include software related to the platform 120, and / or any other software that can be provided through the cloud computing environment 122. In some embodiments, an application program 124-1 may send / receive information to / from at least one other application program 124-1 via the virtual machine 124-2.

[0033] The virtual machine 124-2 includes a software implementation of a machine (e.g., a computer) that executes programs, similar to a physical machine. The virtual machine 124-2 may be a system virtual machine or a process virtual machine, depending on the usage and correspondence of the virtual machine 124-2 to any real machine. The system virtual machine may provide a complete system platform that supports the execution of a complete operating system (OS). The process virtual machine may execute a single program and may support a single process. In some embodiments, the virtual machine 124-2 may execute on behalf of a user (e.g., the user device 110) and may manage the infrastructure of the cloud computing environment 122, such as data management, synchronization, or long-term data transmission.

[0034] The virtualized storage 124-3 includes at least one storage system and / or at least one device that uses virtualization technology within the storage system or device of the computing resources 124. In some embodiments, within the context of a storage system, the types of virtualization can include block virtualization and file virtualization. Block virtualization can refer to the abstraction (or separation) of logical storage from physical storage so that the storage system can be accessed without regard to the physical storage or heterogeneous architecture. The separation can allow the administrator of the storage system to flexibly manage the storage of end users. File virtualization can eliminate the dependence between the data accessed at the file level and the location of the physical storage file. This can optimize the performance of storage utilization, server consolidation, and / or uninterrupted file migration.

[0035] The hypervisor 124-4 can provide hardware virtualization technology that allows at least two operating systems (e.g., "guest operating systems") to execute simultaneously on a host computer such as the computing resources 124. The hypervisor 124-4 can provide a virtual operating platform to the guest operating systems and can manage the execution of the guest operating systems. At least two instances of various operating systems can share the virtualized hardware resources.

[0036] The network 130 includes at least one wired and / or wireless network. For example, the network 130 can include a cellular network (e.g., a fifth generation (5G) network, a Long-Term Evolution (LTE) network, a third generation (3G) network, a Code Division Multiple Access (CDMA) network, etc.), a Public Land Mobile Network (PLMN), a Local Area Network (LAN), a Wide Area Network (WAN), a Metropolitan Area Network (MAN), a telephone network (e.g., a Public Switched Telephone Network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-based network, etc., and / or a combination of these or other types of networks.

[0037] Figure 1 The number and arrangement of the devices and networks shown are provided as examples. In fact, compared with Figure 1 the devices and / or networks shown, there can be more devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks with a different arrangement. Additionally, Figure 1The at least two devices shown can be implemented within a single device, or Figure 1 the single device shown can be implemented as at least two distributed devices. Additionally or alternatively, a set of devices (e.g., at least one device) of environment 100 can perform at least one function described as being performed by another set of devices of environment 100.

[0038] Figure 2 is Figure 1 a block diagram of example components of at least one device. Device 200 can correspond to user device 110 and / or platform 120. As Figure 2 shown, device 200 can include bus 210, processor 220, memory 230, storage component 240, input component 250, output component 260, and communication interface 270.

[0039] Bus 210 includes components that permit communication between the components of device 200. Processor 220 is implemented in hardware, firmware, or a combination of hardware and software. Processor 220 is a central processing unit (CPU), graphics processing unit (GPU), accelerated processing unit (APU), microprocessor, microcontroller, digital signal processor (DSP), field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), or another type of processing component. In some implementations, processor 220 includes at least one processor capable of being programmed to perform functions. Memory 230 includes random access memory (RAM), read-only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, and / or optical memory) that stores information and / or instructions for use by processor 220.

[0040] Storage component 240 stores information and / or software related to the operation and use of device 200. For example, storage component 240 can include a hard disk (e.g., a magnetic disk, optical disk, magneto-optical disk, and / or solid state disk), compact disc (CD), digital versatile disc (DVD), floppy disk, cassette tape, magnetic tape, and / or another type of non-transitory computer-readable medium, as well as corresponding drives.

[0041] Input component 250 includes components that permit device 200 to receive information, such as via user input, e.g., a touch screen display, keyboard, keypad, mouse, button, switch, and / or microphone. Additionally or alternatively, input component 250 can include sensors for sensing information (e.g., a global positioning system (GPS) component, accelerometer, gyroscope, and / or actuator). Output component 260 includes components that provide output information from device 200, such as a display, speaker, and / or at least one light-emitting diode (LED).

[0042] The communication interface 270 includes transceiver-like components (e.g., a transceiver and / or separate receiver and transmitter) that enable the device 200 to communicate with other devices, for example, via a wired connection, a wireless connection, or a combination of wired and wireless connections. The communication interface 270 may allow the device 200 to receive information from another device and / or provide information to another device. For example, the communication interface 270 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, etc.

[0043] The device 200 may perform at least one of the processes described herein. The device 200 may perform these processes in response to the processor 220 executing software instructions stored by a non-transitory computer-readable medium (e.g., the memory 230 and / or the storage component 240). A computer-readable medium is defined herein as a non-transitory memory device. A memory device includes storage space within a single physical storage device or storage space distributed across at least two physical storage devices.

[0044] The software instructions may be read into the memory 230 and / or the storage component 240 from another computer-readable medium or from another device via the communication interface 270. When executed, the software instructions stored in the memory 230 and / or the storage component 240 may cause the processor 220 to perform at least one of the processes described herein. Additionally or alternatively, hardware wired circuitry may be used in place of or in combination with the software instructions to perform at least one of the processes described herein. Accordingly, the embodiments described herein are not limited to any particular combination of hardware circuitry and software.

[0045] Figure 2 The number and arrangement of the components shown are provided as an example. In fact, compared to the components shown, the device 200 may include more components, fewer components, different components, or components arranged differently. Additionally or alternatively, a set of components (e.g., at least one component) of the device 200 may perform at least one function described as being performed by another set of components of the device 200. Figure 2 Shown is a sample DASH processing model 300 such as a sample client architecture for handling DASH and CMAF events. In the DASH processing model 300, requests by the client for media segments (e.g., advertisement media segments and live media segments) may be based on the addresses described in the manifest 303. The manifest 303 also describes metadata tracks from which the client may access segments of the metadata tracks, parse them, and send them to the application 301.

[0046] Figure 3

[0047] ​Manifest 303 includes at least two MPD events or events, and the in-band event and "moof" parser 306 can parse at least two MPD event fragments or event fragments and append the at least two event fragments to the event and metadata buffer 330. The in-band event and "moof" parser 306 can also obtain media fragments and append them to the media buffer 340. The event and metadata buffer 330 can send event and metadata information to the event and metadata synchronizer and scheduler 335. The event and metadata synchronizer and scheduler 335 can schedule specific events to the DASH player control, selection, and heuristic logic 302, and schedule application-related event and metadata tracks to the application 301.

[0048] According to some embodiments, the MSE may include a pipeline that includes a file format parser 350, a media buffer 340, and a media decoder 345. The MSE 320 is at least one logical buffer for media fragments, where media fragments can be detected and sorted based on the presentation time of the media fragments. Media fragments can include, but are not limited to, advertisement media fragments associated with an advertisement MPD and live media fragments associated with a live MPD. Each media fragment can be added or appended to the media buffer 340 based on the timestamp offset of the media fragment, and the timestamp offset can be used to sort the media fragments in the media buffer 340.

[0049] Since embodiments of the present application may involve constructing a linear media source extension (MSE) buffer from at least two non-linear media sources using an MPD chain, and the non-linear media sources can be an advertisement MPD and a live MPD, the file format parser 350 can be used to handle different media and / or codecs used for live media fragments included in the live MPD. In some embodiments, the file format parser can issue change types based on the codec, profile, and / or level of the live media fragment.

[0050] As long as media fragments exist in the media buffer 340, the event and metadata buffer 330 maintains the corresponding event fragments and metadata. The sample DASH processing model 300 can include a timing metadata detection parser 325 to detect metadata associated with in-band and MPD events. According to Figure 3 , the MSE 320 only includes a file format parser 350, a media buffer 340, and a media decoder 345. The event and metadata buffer 330 and the event and metadata synchronizer and scheduler 335 are not local to the MSE 320, thereby prohibiting the MSE 320 from locally processing events and sending them to the application.

[0051] The semantics of DASH preselected elements are shown in Table 1. Table 1

[0052] As shown in the above table, the attribute @order defines a very specific multiplexing scheme. Each time a new scheme is needed, a new value and the consistency rules for stream multiplexing for that value need to be added to the specification.

[0053] According to at least one embodiment, a new attribute for preselection, called interleaving, is provided, as shown in Table 2. Table 2

[0054] The new attribute in Table 2 is shown underlined. As shown in Table 2, the @interleaving attribute is an opaque attribute. For example, the content of this attribute is not defined by the DASH specification. The syntax and semantics of its content are defined by the decoder specification or the relevant specification used in the preselected adaptation set or content component element. Therefore, the job of the DASH client is to provide the @interleaving information to the DASH application, and the DASH application uses this information to operate and multiplex the received segments / sub-segments.

[0055] Since the syntax and semantics of interleaving are defined by an external specification, this attribute is extensible in the future, which means that any new decoder specification can define at least one interleaving scheme, syntax, and semantics. And since the application streaming those streams knows the decoder, it should know the interleaving instructions and consistency rules. Therefore, based on the embodiments of the present disclosure, the DASH specification is not bound to any specific codec.

[0056] Table 3 shows another way of signaling interleaving information, where the interleaving information is included as an option in @order. Table 3

[0057] As shown in Table 3, @order has a new value. According to at least one embodiment, if the @order value starts with the substring "opaque", the rest of the @order value provides information about the multiplexing and operation of the received (sub-) segments. In at least one example, the syntax and semantics are defined by at least one decoder specification and are outside the scope of the DASH specification.

[0058] Embodiments of the present disclosure also provide an alternative method that allows flexible bitstream operation independent of the DASH specification for picture-in-picture.

[0059] Figure 4 Shows an example picture-in-picture use case. As shown, the main picture 400 can occupy the entire screen, while the overlay picture 402 can occupy a small area of the screen, covering the corresponding area of the main picture. The coordinates of the picture-in-picture (PiP) can be represented by x, y, height, and width, where these parameters define the position and size of the PiP relative to the coordinates of the main picture accordingly.

[0060] In the case of streaming, the main video and the PiP video can be transmitted as two separate streams. If there are independent streams, the main video and the PiP video are decoded by separate decoders and then combined for rendering. If the video codec used supports merged streams, the PiP video stream is combined with the main video stream, possibly replacing the stream representing the overlay area of the main video with the PiP video, and then the single stream is sent to the decoder for decoding and then rendering.

[0061] DASH CDAM2 provides the following solution for picture-in-picture signaling.

[0062] In at least one example, the SupplementalProperty element whose @schemeIdUri attribute is equal to urn:mpeg:dash:pinp:2022 is called a picture-in-picture (PiP) descriptor. In at least one example, a PiP descriptor can exist at a preselection level. The presence of a PiP descriptor in the preselection indicates that the purpose of the preselection is to provide a PiP experience.

[0063] In at least one example, the PiP service provides the ability to include a video with a smaller spatial resolution in a video with a larger spatial resolution. In this case, different code streams / representations of the main video are included in the preselected main adaptation set, and different code streams / representations of the supplementary video (also called the PiP video) are included in the preselected partial adaptation set.

[0064] When a PiP descriptor exists in the preselection and the picInPicInfo@dataUnitsReplacable attribute exists and is equal to "true", the client can choose to replace the encoded video data units representing the target PiP area in the main video with the corresponding encoded video data units of the PiP video before sending them to the video decoder. In this way, separate decoding of the main video and the PiP video can be avoided. For a specific picture in the main video, the corresponding video data units of the PiP video are all the encoded video data units in the decoded time-synchronized samples of the supplementary video representation.

[0065] The @value attribute of the PiP descriptor should not be present. The PiP descriptor may include a picInPicInfo element, whose attributes are shown in Table 4. Table 4

[0066] In at least one example, the XML syntax of the PicInpicInfo element can be specified as follows:

[0067] In at least one example, another solution for PiP support in DASH is to use the "pip" value of Role, and use ContentComponent together with Role and @tag to signal the subpicture ID or any other ID required.

[0068] In at least one example, the operations of streaming and composition can be specified by the decoder specification rather than the specification of the dash client. In at least one example, the dash client provides content attributes and metadata to the DASH application, and it is the job of the DASH application to perform any streaming operations, PiP positioning, and rendering.

[0069] As shown above, preselected elements can be used, but a new element PicInpicInfo inside the preselection is required to signal the streaming operation by replacing at least one subpicture stream of the main video with the encoded stream of the PiP video. This operation is very specific to the video decoder and cannot be generalized to other streaming operations and multiplexing.

[0070] According to at least one embodiment, the preselected interleaving attribute or order attribute can be used to provide opaque information about multiplexing to the application instead of explicitly providing multiplexing information. In at least one example, the opaque information is transparent to the DASH client and the decoder. In this case, the DASH client does not need to parse and understand the content of the interleaving attribute. In at least one example, the DASH application is responsible for streaming operations and multiplexing, and the instructions, syntax, and semantics are defined by at least one decoder specification. Since the DASH application knows the decoder it is using, the DASH application is able to understand the opaque instructions provided by the DASH client.

[0071] According to at least one embodiment, a new value is added to the DASH role scheme. The value "pip" can indicate that preselected elements are used for the PiP experience. Table 5 shows an example semantics of the "pip" value. Table 5

[0072] The underlined lines indicate new values.

[0073] In at least one example, a normal audio / video program labels both the main audio and video as "main". However, when the two media component types are not equally important, e.g., (a) a video that provides a pleasant visual experience to accompany a music track as the main content, or (b) an ambient audio that accompanies a video as the main content, which shows live scenes such as a sports event, the accompanying media can be assigned a "supplementary" role.

[0074] In at least one example, especially when at least two alternative media content components including at least two supplementary media content components are available, it can be expected that the alternative media content components carry additional descriptors to indicate their differences from the main media content components (e.g., viewpoint descriptors or role descriptors).

[0075] In at least one example, open ("burned") captions or translated captions can be labeled only as the media type component "video", but with a descriptor indicating "captions" or "translated captions".

[0076] In at least one example, role descriptors with values such as "translated captions", "captions", "description", "sign language", or "metadata" can be used to assign a "kind" value to the tracks exposed in an HTML 5 application for a DASH MPD.

[0077] In at least one example, preselection elements can be used to describe a set of adaptation sets in the MPD that are suitable for a picture-in-picture (PiP) experience. The PiP experience provides the ability to include a video with a smaller spatial resolution within a video with a larger spatial resolution. In this case, different bitstreams / representations of the main video can be included in a preselected main adaptation set, and different bitstreams / representations of the supplementary video (also known as the PiP video) are included in a preselected partial adaptation set.

[0078] In at least one example, the preselection element indicating a PiP presentation can exactly include one Role element according to a role scheme and the value "pip". The preexistence of this descriptor with this value can indicate that the purpose of the preselection is to provide a PiP experience.

[0079] In at least one example, Preselection@interleaving, if present in a DASH bitstream, provides instructions to a DASH application to process presentation segments / sub-segments before providing them to at least one decoder. In at least one example, the syntax and semantics used in this attribute can be defined by the specification and / or related documentation that defines the decoder used in the adaptation set. In at least one example, content components belong to this preselection via Preselection@preselectionComponents.

[0080] In at least one example, in the case of VVC, sub-picture ids can be utilized to identify sub-pictures. In at least one example, the following syntax can be used for Preselection@interleaving: subpic1 subpic2…, where subpic1, subpic2, and … are spatially separated sub-picture ids of a VVC bitstream, each sub-picture id defining a sub-picture, and the set defining the entire region available for picture-in-picture overlay.

[0081] Figure 5 An example process for signaling and processing multiplexing information in a DASH bitstream is shown.

[0082] The process can start at operation S502: Receive a DASH bitstream. The DASH bitstream can include multiple media segments.

[0083] The process proceeds to operation S504: Determine that the DASH bitstream includes a preselection element for multiplexing media segments. The preselection element can be Preselection@interleaving or Preselection@order with a predefined string value (e.g., "opaque").

[0084] The process proceeds to operation S506: Parse media segments from the bitstream. For example, the multiplexing information for multiplexing media segments can specify at least two media segments to be multiplexed or interleaved, where these specified media segments are parsed from the bitstream.

[0085] The process proceeds to operation S508: Multiplex the media segments using the preselection element and at least one strategy of the decoder.

[0086] According to at least one embodiment, a method includes: using a preselection for picture-in-picture experience signaling, where attributes are used to indicate multiplexing information and instructions to an application, the multiplexing information and instructions being invisible to a DASH client, the DASH application being capable of decoding the multiplexing information and instructions and applying corresponding multiplexing and operations to received media segments, thereby providing scalability of the multiplexing scheme without being explicitly defined by the DASH standard specification.

[0087] According to at least one embodiment, a method includes: making a DASH preselection element scalable for bitstream operation and multiplexing, where a new attribute is added to the DASH preselection element, the new attribute carrying bitstream operation and multiplexing information and instructions, the corresponding syntax and semantics of the bitstream operation and multiplexing information being defined by an external specification of a decoder standard, the specification of the decoder standard defining a set of syntax and semantics independent of the DASH specification, where the DASH specification includes instructions associated with the set of syntax and semantics to achieve scalability of preselected bitstream operation and multiplexing.

[0088] The above techniques can be implemented as computer software using computer-readable instructions and physically stored in at least one computer-readable medium.

[0089] Embodiments of the present disclosure can be used alone or in any combination. In addition, each embodiment (and its method) can be implemented by a processing circuit (e.g., at least one processor or at least one integrated circuit). In one example, at least one processor executes a program stored in a non-volatile computer-readable medium.

[0090] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the embodiments to the precise forms disclosed. Modifications and variations can be made in light of the above disclosure, or can be obtained from practice of the embodiments.

[0091] As used herein, the term "component" is intended to be broadly construed as hardware, firmware, or a combination of hardware and software.

[0092] Although combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features can be combined in ways not specifically recited in the claims and / or not disclosed in the specification. Although each dependent claim listed below may directly depend on only one claim, the disclosure of possible implementations includes each dependent claim combined with every other claim in the claim set.

[0093] Unless expressly stated otherwise, elements, acts, or instructions used herein should not be construed as critical or essential. Additionally, as used herein, the articles "a" and "an" are intended to include at least one item and may be used interchangeably with "one or more." Further, as used herein, the term "set" is intended to include at least one item (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and may be used interchangeably with "one or more." Where only one item is intended, the term "one" or similar language is used. Additionally, as used herein, the terms "has," "have," "having," etc. are intended to be open-ended terms. Further, the phrase "based on" is intended to mean "at least partially based on" unless expressly stated otherwise.

Claims

1. A method performed by at least one processor of a decoder, characterized in that The method comprises: Receiving DASH stream based on Hypertext Transfer Protocol HTTP; Determining that the DASH code stream includes a pre-selected element, where the pre-selected element is used to multiplex a plurality of media segments included in the DASH code stream; Parsing the plurality of media segments from the bitstream; multiplexing, by a DASH application, the plurality of segments using the preselected elements and at least one policy associated with the decoder to generate a multiplexed codestream; and The multiplexed code stream is output.

2. The method according to claim 1, characterized in that The preselection element is preselection@interleaving.

3. The method according to claim 2, characterized in that The preselection@interleaving comprises at least one instruction for interleaving at least two media segments of the plurality of media segments.

4. The method according to claim 2, characterized in that: The preselection@intervleaving replaces the preselectionon@order element included in the DASH code stream.

5. The method according to claim 1, characterized in that The preselection element is preselection@order, which specifies a predetermined string value.

6. The method according to claim 5, characterized in that The predetermined string value is "opaque".

7. The method according to claim 6, characterized in that The preselection element preselection@order specifying the "opaque" string value further includes at least one instruction for interleaving at least two media segments of the plurality of media segments.

8. The method according to claim 1, characterized in that Further including: It is determined that the DASH code stream includes a role element, the role element has a value, and the value indicates that the multiplexed code stream is a picture-in-picture.

9. The method according to claim 8, characterized in that The role element is Role@pip.

10. The method according to claim 8, characterized in that The preselection element is preselection@interleaving, The preselection@interleaving includes at least one instruction for interleaving at least two of the multiple media segments to provide a picture-in-picture, wherein the first of the at least two media segments is a main picture, and the second media segment is an overlay picture to be displayed within the main picture.

11. A decoder, characterized in that: include: at least one memory configured to store program code; as well as At least one processor is configured to read the program code and operate according to instructions of the program code, wherein the program code includes: A receiving code is configured to enable the at least one processor to receive a dynamic adaptive streaming DASH code stream based on the hypertext transfer protocol HTTP; A first determination code is configured to cause the at least one processor to determine that the DASH code stream includes a pre-selected element, where the pre-selected element is used to multiplex a plurality of media segments included in the DASH code stream; A parsing code configured to enable the at least one processor to parse the plurality of media segments from the bitstream; A multiplexing code configured to cause the at least one processor to multiplex the plurality of segments by a DASH application using the preselected elements and at least one policy associated with the decoder to generate a multiplexed code stream; and An output code is configured to cause the at least one processor to output the multiplexed code stream.

12. The decoder according to claim 11, characterized in that The preselection element is preselection@interleaving.

13. The decoder according to claim 12, characterized in that The preselection@interleaving comprises at least one instruction for interleaving at least two media segments of the plurality of media segments.

14. The decoder according to claim 12, characterized in that The preselection@intervleaving replaces the preselectionon@order element included in the DASH code stream.

15. The decoder according to claim 11, characterized in that The preselection element is preselection@order, which specifies a predetermined string value.

16. The decoder according to claim 15, characterized in that The predetermined string value is "opaque".

17. The decoder according to claim 16, characterized in that The preselection element preselection@order specifying the "opaque" string value further includes at least one instruction for interleaving at least two media segments of the plurality of media segments.

18. The decoder according to claim 11, characterized in that The second determination code is configured to enable the at least one processor to determine that the DASH code stream includes a role element, and the role element has a value, and the value indicates that the multiplexed code stream is a picture-in-picture.

19. The method according to claim 8, characterized in that The role element is Role@pip.

20. A method performed by at least one processor of a decoder, the method comprising: Processing of dynamic adaptive streaming DASH stream based on Hypertext Transfer Protocol HTTP; The DASH code stream includes a pre-selected element, and the pre-selected element is used to multiplex a plurality of media segments included in the DASH code stream. The plurality of media segments are parsed from the bitstream, and The DASH application multiplexes the plurality of segments using the preselected elements and at least one policy associated with the decoder to generate a multiplexed codestream.