Method of signaling haptic information for DASH selection process by using initialized segmentation

By realizing the reception, parsing and adaptation set selection of haptic media files in DASH applications, the shortcomings of the MPEG DASH standard for haptic streaming are solved and a high-quality haptic experience is achieved.

CN120077627APending Publication Date: 2025-05-30TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004372.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-07-09
Filing Date
2024-07-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The MPEG DASH standard fails to effectively support haptic streaming, especially when multiple haptic tracks are present, it is difficult to select the right adaptation set for a high-quality haptic experience.

Method used

A method and apparatus are provided to receive, parse and select adapter sets in a haptic media file through a DASH application to realize streaming and rendering based on the haptic adapter set. The device includes a processor and memory, enabling haptic processing and adapter set selection through computer program code.

Benefits of technology

Through this method, it is possible to effectively select and stream tactile adaptation sets, improve the quality of tactile experience in multimedia presentation, and solve the shortcomings of the DASH standard for tactile streaming.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077627A_ABST
    Figure CN120077627A_ABST
Patent Text Reader

Abstract

Comprising a method and apparatus including computer code configured to cause one or more processors to be implemented by an HTTP-based DASH application implemented by the one or more processors: receive a media file, the media file comprising haptics indicated by an adaptation set of a DASH MPD file; parsing any one of the haptic virtual object, perception, channel, and band identifier; selecting a subset of a haptic adaptation set including the adaptation set based on parsing any one of the virtual object, perception, channel, and band identifier; and streaming and rendering, in the DASH, haptic feedback of the media file based on the selected subset of the haptic adaptation set.
Need to check novelty before this filing date? Find Prior Art

Description

Cross - Reference to Related Applications

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 526,146, filed on July 11, 2023; U.S. Provisional Patent Application No. 63 / 526,147, filed on July 11, 2023; U.S. Provisional Patent Application No. 63 / 526,156, filed on July 11, 2023; U.S. Provisional Patent Application No. 63 / 526,154, filed on July 11, 2023; U.S. Provisional Patent Application No. 63 / 543,261, filed on October 9, 2023; and U.S. Patent Application No. 18 / 767,053, filed on July 9, 2024. The entire contents of all prior applications are hereby expressly incorporated by reference into this application. Technical Field

[0002] This application provides for signaling a haptic experience configuration in a Dynamic Adaptive Streaming over Hypertext Transfer Protocol (DASH) file, at least at a high level or a middle level, for selecting the correct adaptation set. Background Art

[0003] Haptic experience has become part of multimedia presentation. In such applications, haptic signals are delivered to a device or a wearable device, and the user feels haptic sensations during the use of the application. Recently, the Moving Picture Experts Group (MPEG) has started researching compression standards for haptics.

[0004] MPEG has also released the DASH standard for Internet media streaming. Recently, MPEG is researching the use of DASH for streaming haptic experiences, either as part of or the entire multimedia presentation.

[0005] MPEG DASH provides a standard for streaming multimedia content over IP networks. Although the DASH standard provides a way to describe various contents and their relationships, the standard does not provide explicit support for haptic streaming, especially when there are multiple haptic tracks to choose from.

[0006] For any of the above reasons, it is thus desirable to provide a technical solution to these problems that arise in computer audio technology. Summary of the Invention

[0007] To solve these technical problems, a method and an apparatus are provided. The apparatus includes a memory configured to store computer program code and one or more processors configured to access the computer program code and operate according to the instructions of the computer program code. The computer program is configured to cause the one or more processors to implement a DASH application for tactile processing based on the Hypertext Transfer Protocol (HTTP): receive code configured to cause at least one processor to receive a media file, the media file including a haptic indicated by an adaptation set of a DASH Media Presentation Description (MPD) file; parse code configured to cause at least one processor to parse any one of a virtual object, a perception, a channel, and a band identifier of the haptic; select code configured to cause at least one processor to select a subset of a haptic adaptation set including the adaptation set based on parsing any one of the virtual object, the perception, the channel, and the band identifier; and stream and render code configured to cause at least one processor to stream and render haptic feedback of the media file in DASH based on the selected subset in the haptic adaptation set.

[0008] The DASH application is configured to command the DASH client to stream from the selected subset in the haptic adaptation set.

[0009] The DASH application can receive a media file from the DASH client.

[0010] The media file can also indicate the haptic through a Haptics Interchange Format (HJIF) profile.

[0011] The media file can also indicate the haptic by referring to a link pointing to an HJIF file.

[0012] The media file can also indicate the haptic through an HJIF configuration object.

[0013] The media file can also indicate the haptic through an Extensible Markup Language (XML) element. Brief Description of the Drawings

[0014] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, wherein:

[0015] Figure 1 is a simplified schematic diagram of a computer environment according to an embodiment.

[0016] Figure 2 is a simplified schematic diagram of media processing according to an embodiment.

[0017] Figure 3 is a simplified block diagram of a decoder according to an embodiment.

[0018] Figure 4 is a simplified block diagram of an encoder according to an embodiment.

[0019] Figure 5 is a simplified block diagram of CMAF features according to an embodiment.

[0020] Figure 6 is a simplified block diagram of an HTTP DASH environment according to an embodiment.

[0021] Figure 7 is a simplified diagram of event message features according to an embodiment.

[0022] Figure 8 is a simplified flowchart according to an embodiment.

[0023] Figure 9 is a simplified flowchart according to an embodiment.

[0024] Figure 10 is a simplified flowchart according to an embodiment.

[0025] Figure 11 is a simplified flowchart according to an embodiment.

[0026] Figure 12 is a simplified flowchart according to an embodiment.

[0027] Figure 13 is a simplified block diagram of a computer environment according to an embodiment. Detailed implementation

[0028] The features presented in the following description can be used alone or in any combination. In addition, the embodiments of the present application can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-volatile computer-readable medium.

[0029] Figure 1FIG. 0 shows a simplified block diagram of a communication system 100 according to an embodiment of the present disclosure. The communication system 100 includes at least two terminals 102 and 103 interconnected via a network 105. For unidirectional data transmission, a third terminal 103 may encode video data at a local location for transmission via the network 105 to another terminal 102. The second terminal 102 may receive the encoded video data of another terminal from the network 105, decode the encoded data, and display the recovered video data. Unidirectional data transmission is common in media service applications and the like.

[0030] Figure 1 FIG. 4 shows a first terminal 101 and a fourth terminal 104 provided to support bidirectional transmission of encoded video, which may occur, for example, during a video conference. For bidirectional data transmission, each of a terminal 101 and a fourth terminal 104 may encode video data captured at a local location for transmission via the network 105 to the other of the first terminal 101 and the fourth terminal 104. Each of the first terminal 101 and the fourth terminal 104 may also receive the encoded video data transmitted by the other of the first terminal 101 and the fourth terminal 104, may decode the encoded data, and may display the recovered video data on a local display device.

[0031] In Figure 1 FIG. 9, the first terminal 101, the second terminal 102, the third terminal 103, and the fourth terminal 104 may be shown as servers, personal computers, and smart phones, but the principles of the present application are not limited thereto. Embodiments of the present application are applied to laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network 105 represents any number of networks that transmit encoded video data between the first terminal 101, the second terminal 102, the third terminal 103, and the fourth terminal 104, including, for example, wired (wired) and / or wireless communication networks. The communication network 105 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks (LANs), wide area networks, and / or the Internet. For the purposes of this discussion, unless otherwise explained below, the architecture and topology of the network 105 may be immaterial to the operation of the present application.

[0032] As an example of an application of the disclosed subject matter, Figure 2 FIG. 14 shows how a video encoder and a video decoder are placed in a streaming environment. The subject matter disclosed in the present application is equally applicable to other video-supported applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including compact discs (CDs), digital video discs (DVDs), memory sticks, and the like.

[0033] A streaming system may include an acquisition subsystem 203, which includes a video source 201 such as a digital camera. The video source creates a stream of video samples 213, such as an uncompressed stream of video samples. The sample stream 213 is emphasized as a high-data-volume sample stream compared to an encoded video bitstream and may be processed by an encoder 202 coupled to the camera 201. The encoder 202 includes hardware, software, or a combination of both to implement or enforce aspects of the disclosed subject matter described in more detail below. The encoded video bitstream 204 is emphasized as a lower-data-volume encoded video bitstream compared to the sample stream and may be stored on a streaming server 205 for future use. One or more streaming clients 212 and 207 may access the streaming server 205 to retrieve copies 208 and 206 of the encoded video bitstream 204. The client 212 may include a video decoder 211 that decodes an incoming copy of the encoded video bitstream 208 and produces an output video sample stream 210 that can be displayed on a display 209 or other display device (not depicted). In some streaming systems, the video bitstreams 204, 206, and 208 may be encoded according to certain video codec / compression standards. Examples of these standards were mentioned above and will be further described herein.

[0034] Figure 3 It may be a functional block diagram of a video decoder 300 according to an embodiment of the present invention.

[0035] A receiver 302 may receive one or more encoded video sequences to be decoded by the decoder 300. In the same or another embodiment, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequences may be received from a channel 301, which may be hardware / software leading to a storage device storing the encoded video data. The receiver 302 may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective using entities (not labeled). The receiver 302 may separate the encoded video sequences from the other data. To prevent network jitter, a buffer memory 303 may be coupled between the receiver 302 and an entropy decoder / parser 304 (hereinafter referred to as "parser"). When the receiver 302 receives data from a storage / forwarding device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory 303 may not be needed, or it may be made smaller. For use on a service packet network such as the Internet, the buffer memory 303 may also be needed, which may be relatively large and may advantageously have an adaptive size.

[0036] Video decoder 300 may include a parser 304 to reconstruct symbols 313 from an entropy-coded video sequence. The categories of these symbols include information for managing the operation of decoder 300, and potentially information for controlling a display device such as display 312, which is not part of the decoder but may be coupled to the decoder. The control information for one or more display devices may be in the form of Supplementary Enhancement Information (SEI) or a Video Usability Information (VUI) parameter set segment (not labeled). The parser 304 may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence may be according to a video coding technology or standard, and may follow various guidelines known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 304 may extract subgroup parameter sets for at least one subgroup among subgroups of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the group. Subgroups may include Groups of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CU), blocks, Transform Units (TU), Prediction Units (PU), etc. The entropy decoder / parser may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, Motion Vectors (MV), etc.

[0037] The parser 304 may perform entropy decoding / parsing operations on the video sequence received from buffer memory 303 to create symbols 313. The parser 304 may receive the encoded data and selectively decode specific symbols 313. Additionally, the parser 304 may determine whether a specific symbol 313 is to be provided to the motion compensation prediction unit 306, the scaler / inverse transform unit 305, the intra prediction unit 307, or the loop filter 311.

[0038] Depending on the type of the encoded video picture or a portion of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of symbols 313 may involve multiple different units. Which units are involved and the manner of involvement may be controlled by subgroup control information parsed by the parser 304 from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser 304 and the multiple units below are not described.

[0039] In addition to the functional blocks already mentioned, decoder 300 can conceptually be subdivided into several functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the functional units below.

[0040] The first unit is the scaler / inverse transform unit 305. The scaler / inverse transform unit 305 receives, from parser 304, quantized transform coefficients as one or more symbols 313 and control information, including which transform mode to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit can output a block including sample values, which can be input into aggregator 310.

[0041] In some cases, the output samples of the scaler / inverse transform unit 305 can belong to an intra-coded block, i.e., a block that does not use predictive information from a previously reconstructed picture but can use predictive information from a previously reconstructed portion of the current picture. Such predictive information can be provided by the intra-picture prediction unit 307. In some cases, the intra-picture prediction unit 307 generates a block having the same size and shape as the block being reconstructed using surrounding reconstructed information extracted from the current (partially reconstructed) picture 309. In some cases, aggregator 310 adds, based on each sample, the prediction information generated by the intra-picture prediction unit 307 to the output sample information provided by the scaler / inverse transform unit 305.

[0042] In other cases, the output samples of the scaler / inverse transform unit 305 can belong to an inter-coded and potentially motion-compensated block. In this case, the motion-compensation prediction unit 306 can access the reference picture memory 308 to extract samples for prediction. After motion-compensating the extracted samples according to the symbols 313 belonging to the block, these samples can be added by aggregator 310 to the output of the scaler / inverse transform unit (which is referred to as residual samples or a residual signal in this case), thereby generating output sample information. The motion-compensation unit obtaining the prediction samples from an address within the reference picture memory can be controlled by a motion vector, and the motion vector is in the form of the symbol 313 and is provided for use by the motion-compensation unit, and the symbol 313 includes, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of sample values extracted from the reference picture memory when using sub-sample accurate motion vectors, a motion vector prediction mechanism, and so on.

[0043] The output samples of aggregator 310 can be adopted by various loop filtering techniques in loop filter unit 311. Video compression techniques can include in-loop filter techniques, which are controlled by parameters included in the encoded video bitstream, and these parameters can be used in loop filter unit 311 as symbols 313 from parser 304. Video compression techniques can also respond to meta-information obtained during decoding of previous (in decoding order) parts of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0044] The output of loop filter unit 311 can be a sample stream, which can be output to display device 312 and stored in reference picture memory 557 for subsequent inter-picture prediction.

[0045] Once fully reconstructed, some encoded pictures can be used as reference pictures for future prediction. Once an encoded picture is fully reconstructed and the encoded picture is identified (e.g., by parser 304) as a reference picture, the current reference picture 309 can become part of reference picture buffer 308, and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.

[0046] Video decoder 300 can perform decoding operations according to predetermined video compression techniques in standards such as ITU-T Rec. H.265. An encoded video sequence can conform to the syntax specified by the video compression technique or standard used in the sense that the encoded video sequence follows the syntax of the video compression technique or standard specified in the video compression technique document or standard (especially in its profile). For compliance, it is also required that the complexity of the encoded video sequence is within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.

[0047] In an embodiment, receiver 302 can receive additional (redundant) data together with the encoded video. The additional data can be part of one or more encoded video sequences. The additional data can be used by video decoder 300 to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data can be in the form of, for example, temporal, spatial, or signal noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0048] Figure 4 It may be a functional block diagram of a video encoder 400 according to an embodiment of the present disclosure.

[0049] The encoder 400 may receive video samples from a video source 401 (not part of the encoder), and the video source may capture one or more video images to be encoded by the encoder 400.

[0050] The video source 401 may provide a source video sequence in the form of a digital video sample stream to be encoded by the encoder 303. The digital video sample stream may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit...), any color space (e.g., BT.601 Y CrCB, RGB...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source 401 may be a storage device storing previously prepared videos. In a video conferencing system, the video source 401 may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. These pictures themselves may be constructed as a spatial array of pixels, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.

[0051] According to an embodiment, the encoder 400 may encode and compress pictures of the source video sequence into an encoded video sequence 410 in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller 402. The controller controls other functional units as described below and is functionally coupled to these units. For simplicity, the couplings are not labeled in the figure. The parameters set by the controller may include rate control related parameters (picture skip, quantizer, λ value of rate distortion optimization techniques, etc.), picture size, GOP layout, maximum MV search range, etc. Those skilled in the art can easily identify other functions of the controller 402 as they may involve optimizing the video encoder 400 for a specific system design.

[0052] Some video encoders operate in an "encoding loop" that is readily recognizable to those skilled in the art. As a very simplified description, the encoding loop may include an encoding portion of the encoder (hereinafter referred to as the source encoder) (responsible for creating symbols based on the input picture to be encoded and one or more reference pictures) and a (local) decoder 406 embedded in the encoder 400. The decoder 406 reconstructs the symbols in a manner similar to how the (remote) decoder creates sample data to create sample data (since in the video compression techniques contemplated by the disclosed subject matter, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream is input to the reference picture memory 405. Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the content in the reference picture buffer is also bit-exactly corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction portion of the encoder are exactly the same as the sample values that the decoder will "see" when using prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is also known to those skilled in the art.

[0053] The operation of the "local" decoder 406 may be the same as that of the "remote" decoder 300 described above, for example, in combination with Figure 3 which has been described in detail. However, briefly referring to Figure 4 , when the symbols are available and the entropy encoder 408 and the parser 304 can encode / decode the symbols losslessly into the encoded video sequence, the entropy decoding portion of the decoder 300, including the channel 301, the receiver 302, the buffer 303, and the parser 304, may not be fully implementable in the local decoder 406.

[0054] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding present in the decoder needs to exist in the corresponding encoder in substantially the same functional form. The description of the encoder technology can be simplified because these encoder technologies are reciprocal to the decoder technologies described in detail. More detailed descriptions are only needed in certain areas and are provided below.

[0055] As part of its operation, the source encoder 403 may perform motion compensation predictive coding. This motion compensation predictive coding performs predictive coding of the input frame by referring to one or more previously encoded frames designated as "reference frames" in the video sequence. In this way, the coding engine 407 encodes the difference between the pixel blocks of the input frame and the pixel blocks of one or more reference frames, and the one or more reference frames can be selected as one or more predictive references for the input frame.

[0056] The local video decoder 406 may decode the encoded video data of those frames that can be designated as reference frames based on the symbols created by the source encoder 403. The operation of the encoding engine 407 may advantageously be lossy processing. When the encoded video data can be decoded at the video decoder ( Figure 4 not shown), the reconstructed video sequence can generally be a copy of the source video sequence with some errors. The local video decoder 406 replicates the decoding process that can be performed by the video decoder on the reference frames and may cause the reconstructed reference frames to be stored in the reference picture cache 405. In this way, the encoder 400 can locally store a copy of the reconstructed reference frames, which has the same content (without transmission errors) as the reconstructed reference frames that will be obtained by the remote video decoder.

[0057] The predictor 404 may perform a prediction search for the encoding engine 407. That is, for a new frame to be encoded, the predictor 404 may search in the reference picture memory 405 for sample data (as candidate reference pixel blocks) or some metadata that can be used as an appropriate prediction reference for the new picture, such as reference picture motion vectors, block shapes, etc. The predictor 404 may operate on a per-pixel basis for sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor 404, it may be determined that the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory 405.

[0058] The controller 402 may manage the encoding operations of the video encoder 403, which include, for example, setting parameters and subgroup parameters for encoding the video data.

[0059] The outputs of all the above functional units may be entropy encoded in the entropy encoder 408. The entropy encoder 408 performs lossless compression on the symbols generated by various functional units according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting these symbols into an encoded video sequence.

[0060] The transmitter 409 may cache one or more encoded video sequences created by the entropy encoder 408 to prepare for transmission through the communication channel 411, which may be hardware / software leading to a storage device storing the encoded video data. The transmitter 409 may merge the encoded video data from the video encoder 403 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0061] The controller 402 may manage the operation of the encoder 400. During encoding, the controller 405 may assign a specific encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, a picture may typically be assigned to any of the following picture types.

[0062] An intra picture (I picture), which may be a picture that can be encoded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh pictures. Those skilled in the art can understand the variants of I pictures and their corresponding applications and characteristics.

[0063] A predictive picture (P picture), which may be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.

[0064] A bi - predictive picture (B picture), which may be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for reconstructing a single block.

[0065] Source pictures may typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block - by - block. These blocks may be predictive - encoded with reference to other (encoded) blocks, which are determined according to the encoding assignment of the corresponding picture applied to the block. For example, blocks of an I picture may be non - predictive - encoded, or the block may be predictive - encoded with reference to already - encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be non - predictive - encoded with reference to a previously - encoded reference picture either through spatial prediction or through temporal prediction. Blocks of a B picture may be non - predictive - encoded with reference to one or two previously - encoded reference pictures either through spatial prediction or through temporal prediction.

[0066] The video encoder 400 may perform encoding operations according to a predetermined video encoding technique or standard, such as the ITU - T Rec. H.265 recommendation. In operation, the video encoder 400 may perform various compression operations, including predictive - encoding operations that exploit the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.

[0067] In an embodiment, the transmitter 409 may transmit additional data when transmitting the encoded video. The source encoder 403 may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0068] Figure 5 An example 500 of CMAF tracks and CMAF switch sets defined by CMAF is shown according to an exemplary embodiment, such as using standards like ISO / IEC JTC 1 / SC 29 / WG03N00654. A CMAF switch is a set of CMAF tracks with some general constraints. The main purpose of a CMAF switch set is to provide alternative representations of the same content in multiple tracks so that during delivery or playback, the player can switch between tracks to adapt to changes in network bandwidth and other changing attributes. For the track format, the CMAF standard will use ISOBMFF, but it does not provide a standard to signal the presence of a CMAF switch set in an ISOBMFF file.

[0069] Embodiments of the present application introduce a new version of the ISOBMGG track selection box with unique attributes. This new box contains several parameters for signaling the CMAF switch set and its attributes.

[0070] According to some embodiments, there is a version 1 switch group that can (i) identify the switch group using a switch group id, (ii) an alternative group id (for a selection list), (iii) have a switchable group id, (iv) track the group ID for a preselected association, and (v) indicate CMAF parameters.

[0071] Embodiments of the present application provide the following definitions: Box type: 'tsel' Container: UserDataBox of the corresponding TrackBox Mandatory: No Quantity: Zero or one Such a track selection box is contained in the UserDataBox of the track it modifies.

[0072] Figure 6 An example 600 of metadata track samples for an adaptation set segment index, such as for any given adaptation set, is shown. For example, for each adaptation set (AS) expected to signal the instantaneous segment bandwidth, a separate adaptation set may also be included in the file, as Figure 6 shown.

[0073] As Figure 6As shown, for AS i with k media representations whose segments are timed-aligned, a new adaptation set AS index is added to the file containing a single representation. This single representation is a timed metadata track whose segments are also timed-aligned with the segments of AS i.

[0074] Figure 7 An example DASH client processing model 700 is shown, such as an example DASH client processing model for a client example architecture for handling DASH and CMAF events. In this model, client requests for media segments can be based on addresses described in a file that also describes a metadata track. The client can access segments of the metadata track from this metadata track, parse these segments, and send them to an application. Additionally, according to an exemplary embodiment of the addresses of media segments described below, a DASH file can provide addresses for indexing segments. Each index segment can provide information about the duration and size of a segment, and a representation index can provide index information for all segments of a given representation.

[0075] The DASH client downloads the MPD and provides information about existing adaptation sets to the application. The application selects an adaptation set for streaming and provides it to the decoder / renderer.

[0076] After selection, the DASH client selects a representation of any selected adaptation set. The DASH client first downloads the initialization segment. Then, the DASH client requests a specific segment number based on the general timing of the addressing scheme and schedules the next segment request afterwards. A switch to a different representation can be decided due to network bandwidth changes.

[0077] Figure 8 An example 800 of a haptic codec architecture according to an exemplary embodiment is shown. As shown in example 900, the haptic encoder generates a compact and efficient binary distribution format (.hmpg) for distribution. The haptic decoder can decode this format and send it to the renderer.

[0078] According to some embodiments, the haptic data hierarchy includes the following levels: Table 1

[0079] According to some embodiments, as in Example 900, there is signaling for initializing segments, where, at S901, the initialization segments of all the representations of the DASH download adaptation set are initialized and passed to the application. At S902, the application decodes the initialization segments including the Media-Integrated Haptic Stream (MIHS) sample entries, learns about the virtual objects, sensations, channels, and the rest in each AS / representation, and selects one or more adaptation sets to stream based on this information.

[0080] There is also signaling using a primary initialization segment, for example, at S900, where to avoid downloading several initialization segments, for each adaptation set, one representation is selected as the most complete representation. The initialization segment of this representation includes the metadata for the representation of the adaptation set. For example, this initialization segment has the highest profile and tier in the adaptation set. This "primary" initialization segment can be downloaded and passed to the application for each adaptation set. Since the application has information about all the adaptation sets, the application can select one or more adaptation sets to stream. In this case, according to an exemplary embodiment, each adaptation set with primary initialization should be signaled by including @initializationPrincipal.

[0081] According to some embodiments, a selection process is provided at S901 and S902, where the DASH client downloads / accesses the MPD and parses the MPD, and then the DASH client provides the application with information about various adaptation sets including the haptic adaptation set. Then, the application requests to download the initialization segments from the haptic adaptation set (according to some embodiments, for any adaptation set with a primary initialization segment, download the primary initialization segment, otherwise download all the initialization segments of all the representations of the adaptation set, and pass the set of initialization segments to the application). Then, the application parses the initialization segments and exports the metadata information, based on which the application selects one or more sensations to render from the list according to user preferences or the application environment. Then, the application generates a list of the adaptation sets of the selected sensations (according to some embodiments, the application may reduce the list of adaptation sets based on network bandwidth requirements, rendering complexity, and other parameters), and the application provides the list of adaptation sets to the DASH client for streaming. Then, if the adaptation set has multiple representations, the DASH client can adapt dynamically between different representations.

[0082] For a haptic-enabled DASH client, according to some embodiments of the present application, to expose haptic information to an application, a descriptor scheme for MPEG haptics is defined, which can be used with basic attributes at the MPD level. The presence of such a descriptor means that if the DASH client recognizes the descriptor, the DASH client has sufficient interfaces to expose haptic selection information to the application.

[0083] Thus, through embodiments of the present application, the following ways are provided to signal haptic configuration information: in response to determining that the haptic adaptation set does not have a primary initialization segment, providing multiple initialization segments representing the haptic adaptation set to the DASH application; in response to determining that the haptic adaptation set has a primary initialization segment, providing the primary initialization segment of the haptic adaptation set to the DASH application; the DASH application parses the primary initialization segment or multiple initialization segments and information about virtual objects, sensations, channels, and bands; the DASH application then selects a set of virtual objects, sensations, channels, or bands, and selects the corresponding haptic adaptation set or the representation containing the corresponding haptic signals; and based on the primary initialization segment or multiple initialization segments, the DASH application then commands the DASH client to stream from the selected haptic adaptation set or representation.

[0084] Haptic experiences can be delivered using many tracks. One or more tracks can create a haptic experience. The experience can include multiple sensations, various device references, multiple virtual objects (to indicate the body parts of the sensations), and different sensory modalities. A sensation can consist of one or more channels, and a channel can consist of one or more bands. A DASH representation can carry a complete experience, a sensation, one or more channels, or one or more bands. Therefore, providing sufficient information for the selection of adaptation sets / representations is essential for DASH streaming of haptic content. Such signaling is not provided by the DASH specification, nor by the haptic coding specification or the ISOBMFF binding specification.

[0085] According to some embodiments, there is signaling of high-level information, such as at S900, where information about the haptic adaptation set is signaled as any of the following lists: a list of supported virtual objects (body parts): used to identify which body part the AS is addressing and for the application to decide whether to stream the AS based on the available actuators; and a list of sensory modalities and a list of sensory ids: the modality shows the sensory modality addressed by the AS, and if the application finds an actuator for the modality, it streams the AS. The id is used to associate other ASs with the same sensory id (i.e., carrying associated haptic signals).

[0086] That is, in one or more embodiments, the adaptation set includes the following attributes: @haptics avatars, @haptic perception Id, and @haptic perception modality. Each attribute is a blank list of values that defines the body part, the perception Id, the modality, and the Id in the adaptation set. The order of the values in the perception attributes should be the same.

[0087] In one or more embodiments, a complementary and / or base descriptor scheme for haptic information is included. The values of the descriptor contain the following information: a list of virtual objects, a list of rendering modalities, and a list of corresponding perception Ids.

[0088] According to some embodiments, there is provided Figure 10 the selection process 1000 in, where at S1001, the DASH client downloads / accesses the MPD signaled at S1000, and at S1002 parses the MPD. Then, as in the process at S1003, the DASH client provides the application with information about various adaptation sets including the haptic adaptation set. Then the application parses the information about the virtual objects, the perception modalities, and the perception Ids (which group the adaptation sets belonging to the same perception Id into one perception). Then, the application compares the virtual objects and the perception modalities with the virtual objects and the perception modalities supported in the device (selecting the perception modalities that the application can render based on the availability of actuators in the respective body parts). Then, the application selects one or more perceptions to render from the list based on user preferences or the application environment. Then, the application generates a list of adaptation sets of the selected perceptions (the application can reduce the adaptation set list based on network bandwidth requirements, the complexity of rendering (such as the level of battery consumption), and other parameters), and the application provides the list of adaptation sets to the DASH client for streaming. Then, if the adaptation set has multiple representations, the DASH client can adapt dynamically between different representations.

[0089] According to an exemplary embodiment, the following information is additionally signaled at the adaptation set level: a list of channels and their body part masks: for identifying which body part the AS is addressing with the channels, and based on the presence of actuators, the application decides whether to stream the AS.

[0090] The above information of the mid-level information in the adaptation set can be signaled by defining @haptic channels with blank values that define the body part (note that according to some embodiments, if @haptic channels exist, @haptic perception Id should exist), and then adding the channel haptic information to the haptic configuration complementary / base descriptor.

[0091] According to some embodiments, the selection process may be similar to the features described above Figure 10 but with such additional mid-level information at the adaptation set level, the application can access both high-level and mid-level information for selection.

[0092] Thus, to enable a DASH client to expose haptic information to an application, a descriptor scheme for MPEG haptics is defined, which can be used together with the basic attributes at the MPD level. The presence of such a descriptor means that if the DASH client recognizes the descriptor, the DASH client has sufficient interfaces to expose haptic selection information to the application. Accordingly, embodiments of the present application provide a method that includes, such as at S900 and S1000, signaling high-level haptic information and mid-level haptic information by including haptic configuration information in multiple haptic adaptation sets and representations of the DASH MPD file, and the high-level haptic information and the mid-level haptic information are visible to the DASH application together with multiple virtual objects, sensations, channels, and band identifiers. Each haptic adaptation set or representation contains virtual objects, sensations, channels, and band identifiers, and the DASH application selects a subset of haptic adaptation sets and representations from the available haptic adaptation sets and representations. The DASH application streams the selected haptic adaptation sets and representations to render haptic feedback.

[0093] According to some embodiments, such as at S900 and S1000, signaling may be done through an external configuration HJIF file for the adaptation set. For example, according to some embodiments, for each adaptation set having a URL to a partial HJIF file, an @hapticConfiguration property is included. The partial HJIF file may include metadata in HJIF format for the following levels: virtual objects, sensations, device configuration, channels, and bands. And not all levels need to be present in the file.

[0094] The difference between the partial HJIF file and the initialization segment of the representation is that the HJIF is more comprehensive (including multiple representations) and can also include information that will appear in the segments later.

[0095] According to some embodiments, such as at S900 and S1000, signaling can occur over a time period, and in such cases, the time period includes the @hapticConfiguration attribute having a URL to an HJIF file. The HJIF file includes metadata in HJIF format at the following levels for collecting adaptation sets over the time period: virtual object, perception, device configuration, channel, and band. Additionally, in such cases, each adaptation set also has @virtual object Id (@AvatarIds), @perception Id (@perceptionIds), @device configuration Id (@deviceConfigurationIds), @channel Id (@channelIds), and @band Id (@bandIds), which list the ids of each level included in the adaptation set. If different representations in the same adaptation set have different elements, the same attributes as above can be used at the representation level.

[0096] According to some embodiments, such as at S900 and S1000, signaling can occur at the MPD. When at the MPD, the MPD can include the @hapticConfiguration attribute having a URL to an HJIF file. The HJIF file includes metadata in HJIF format at the following levels for collecting adaptation sets over all time periods: virtual object, perception, device configuration, channel, and band. In such cases, each adaptation set also has @avatarIds, @perceptionIds, @deviceConfigurationIds, @channelIds, and @bandIds, which list the ids of each level included in the adaptation set. If different representations in the same adaptation set have different elements, the same attributes as above can be used at the representation level.

[0097] According to some embodiments, such as at S900 and S1000, when a partial HJIF file for a selection process is an HJIF file having all hierarchical data structures except effects, signaling can occur at a partial HJIF file such as according to Table 2 below: Table 2

[0098] Note that according to some embodiments, the content provider can decide what content to include in the HJIF configuration file. For example, the HJIF configuration file may only include virtual object and perception data and not include the remaining data.

[0099] As in Figure 11In Example 1100, according to an exemplary embodiment, the selection process is defined as follows: In S1100, the DASH client downloads / accesses the MPD and parses the MPD in S1101. In S1102, depending on the level that provides a link to the HJIF configuration file, one or more HJIF files are downloaded: once at the MPD level with any MPD update; once at the start of each period if the link is provided at the period level; and once for each adaptation set at the start of the period if the link is provided at the adaptation set level.

[0100] That is, according to some embodiments, the application parses information about the configuration. The application selects the sensations, channels, and bands to be rendered. The application (by the ids listed in each adaptation set and / or representation) finds the relevant adaptation sets and selects those adaptation sets. The application provides a list of the adaptation sets to the DASH client for streaming. If an adaptation set has multiple representations, the DASH client dynamically adapts between the different representations.

[0101] Thus, according to an embodiment of the present application, in order for the DASH client to expose haptic information to the application, a descriptor scheme for MPEG haptics is defined, which can be used together with the basic attributes at the MPD level. The presence of this descriptor means that if the DASH client recognizes the descriptor, the DASH client has sufficient interfaces to expose haptic selection information to the application. Therefore, the present application also provides a method: signaling haptic configuration information by providing a link to the HJIF configuration file, where the HJIF configuration file defines haptic configuration information at multiple different levels in the haptic data hierarchy, and the haptic data hierarchy includes multiple virtual objects, sensations, channel, and band identifiers for haptic representations, haptic adaptation sets, periods, and / or MPDs, and the DASH client downloads the HJIF configuration file and provides it to the DASH application. The DASH application parses the HJIF configuration file received from the DASH client, and based on the parsed information from the HJIF configuration file, the DASH application selects a subset of sensations, channels, or bands. Among them, based on the initialization segment, the DASH application selects the corresponding haptic adaptation set or the representation containing the corresponding haptic signal. According to some embodiments, the DASH application then commands the DASH client to stream from the selected haptic adaptation set and / or representation.

[0102] In addition, signaling can also be done through an embedded HJIF configuration object, which is an HJIF object having all the hierarchical data structures except for the effects as shown in the following table, for the above selection process: Table 3 Note that the content provider can decide what is included in the HJIF configuration object. For example, the HJIF configuration object may include virtual objects and perception data, but not the remaining data.

[0103] Accordingly, embodiments define haptic configuration attributes that can be used at different levels in the MPD. In other words, the @hapticConfig attribute is of type xs:string and this type contains the base64 encoding of the HJIF configuration object.

[0104] According to some embodiments, when used at a representation level, the representation includes an @hapticConf attribute that includes the base64-encoded HJIF configuration object for the adaptation set.

[0105] When used at the adaptation set level, the adaptation set includes an @hapticConf attribute that includes the base64-encoded HJIF configuration object for the adaptation set.

[0106] According to some embodiments, when used at a time period level, the time period includes an @hapticConf attribute. The base64-encoded HJIF configuration object includes haptic configuration information for collecting the adaptation sets within that time period. In this case, according to some embodiments, each adaptation set also includes @avatarIds, @perceptionIds, @deviceConfigurationIds, @channelIds, and @bandIds, which list the ids of each level included in the adaptation set. If different representations within the same adaptation set include different elements, the same attributes as above can be used at the representation level.

[0107] When used at the MPD level, the MPD may include an @hapticConf attribute. The base64-encoded HJIF configuration object includes metadata in HJIF format for the following levels for collecting the adaptation sets for all time periods: virtual objects, perception, device configuration, channels, and bands. In this case, each adaptation set also has @avatarIds, @perceptionIds, @deviceConfigurationIds, @channelIds, and @bandIds, which list the ids of each level included in the adaptation set.

[0108] According to some embodiments, for example, in Figure 12In Example 1200, the selection process is defined as follows: At S1200, the DASH client downloads / accesses the MPD and parses the MPD at S1201. At S1202, based on the hierarchy providing the HJIF configuration object, one or more HJIF objects are base64 decoded: decoded once at the MPD level with any MPD updates; decoded once at the start of each period if the HJIF configuration object is provided at the period level; and decoded once for each adaptation set at the start of the period if the HJIF configuration object is provided at the adaptation set level.

[0109] In other words, the application uses one or more HJIF configuration objects to parse information about the configuration. The application selects the sensations, channels, and bands to be rendered. At S1202, the application (by the id listed in each adaptation set and / or representation) finds the relevant adaptation sets and selects those adaptation sets. The application provides a list of the adaptation sets to the DASH client for streaming. If an adaptation set has multiple representations, the DASH client dynamically adapts between the different representations.

[0110] Thus, for a haptics-enabled DASH client according to an embodiment of the present application, in order for the DASH client to expose haptic information to the application, a descriptor scheme for MPEG haptics is defined, which can be used with the basic attributes at the MPD level. The presence of this descriptor means that if the DASH client recognizes this descriptor, the DASH client has sufficient interfaces to expose haptic selection information to the application. That is, the present application provides a method: to signal haptic configuration information by embedding HJIF configuration objects at different levels in the DASH file or MPD, where the embedded HJIF configuration objects define haptic configurations at multiple different levels in the haptic data hierarchy, and the haptic data hierarchy includes multiple virtual objects, sensations, channels, and bands for haptic representations, haptic adaptation sets, periods, or MPDs. The DASH client downloads the embedded HJIF configuration object as part of the MPD and provides it to the DASH application, and the DASH application parses the MPD received from the DASH client to obtain the HJIF configuration. Then, the DASH application selects a set of sensations, channels, or bands based on the information about the HJIF configuration, and the DASH application selects the corresponding haptic adaptation set or the representation containing the corresponding haptic signal based on the set of initialization segments. Then, the DASH application subsequently commands the DASH client to stream from the selected haptic adaptation set or representation.

[0111] In view of the above, further explanations of the relevant data of these embodiments are provided below. For example, according to an exemplary embodiment, starting from the haptic information at the MPD level, the metadata required by the MPD enables the MPD to have or must have the following information: (a) general high-level information of the haptic, if the general high-level information of the haptic is general within a time period or for the entire MPD, (b) information of the adaptation set, which is differentiated for use in haptic applications, (c) information differentiating the representation of the adaptation set for bandwidth adaptation, and (d) optionally, preselection information combining multiple adaptation sets to create a preselected presentation.

[0112] Therefore, the information carried in the MPD depends on the granularity of the adaptation set. For example, if there is a single haptic adaptation set with a complete experience in the MPD, no selection is required, and based on each adaptation set at the level containing the experience, the information to be exposed is as shown in Table 3: Table 1: Haptic Information of Possible Adaptation Set Levels

[0113] For the format of carrying information in the MPD according to an exemplary embodiment, the metadata information is stored in a hierarchical structure in the HJIF format and the ISOBMFF experience box, while it is stored in a flat manner with cross-references in the MIHS and ISOBMFF configuration boxes. In the case where the complete haptic information can be located in one MPD element (e.g., the adaptation set), it seems that storing in a hierarchical structure is more compact. However, due to the various ways of grouping haptic perceptions, channels, and frequency bands in the adaptation set, storing in a flat manner with references seems more appropriate.

[0114] For flat XML elements, according to some embodiments with this solution, similar to the MIHS and ISOBMFF configurations, the information is stored in flat elements, cross-referenced with each other for association. Then, depending on the scope of the metadata, the shared elements are maintained at the top level, the MPD level, or the time period level, and the adaptation set-specific metadata is signaled at each adaptation set level. Through name extension, haptic elements can exist at the MPD, time period, and adaptation set element levels. The elements that the haptic elements may contain at each level depend on the haptic tracks in the adaptation set. Optionally, the haptic descriptor (using the scheme ID URI) can be defined to include the elements in the MPD, time period, and adaptation set elements.

[0115] Further semantics and syntax are provided in the following table according to an exemplary embodiment. Table 4: Haptic Semantics

[0116] The following shows how to use this element in various transmission scenarios. Table 5: Use of Haptic Elements in Various Scenarios

[0117] The following tables show various haptic element semantics. Table 6: Experience Semantics Table 7: Virtual Object Semantics Table 8: Perception Semantics Table 9: Reference Device Semantics Table 10: Channel Semantics Table 11: Band Semantics Table 12: Effect Library Semantics

[0118] Accordingly, embodiments of the present application provide a method for signaling haptic configuration information by including haptic XML elements and cross - referencing other sub - elements to show associations, where the haptic XML elements are composed of multiple sub - elements for perception, channels, frequency bands, and effect libraries with a flat structure. Based on the characteristics of the adaptation set track, the adaptation set element can carry haptic elements for a complete experience, or for one or more perceptions, or for one or more channels, or for multiple one or more frequency bands. Based on the commonality of the data, haptic metadata information is included in different levels of the MPD. If the common data is invariant for the entire MPD, the common information is included at the MPD level, or if there is no haptic track for some time periods, the common information is included at the time period level. Then, in each adaptation set with haptic content, regardless of whether the haptic content is one or more perceptions, one or more channels, or one or more frequency bands, the haptic elements co - exist with elements indicating the metadata specific to that adaptation set.

[0119] The above - described technology can be implemented as computer software that uses computer - readable instructions and is physically stored in one or more computer - readable media or implemented by one or more specially - configured hardware processors. For example, Figure 13 FIG. 1300 shows a computer system suitable for implementing certain embodiments of the disclosed subject matter.

[0120] The computer software can be encoded in any suitable machine code or computer language, creating code including instructions through mechanisms such as assembly, compilation, and linking, which can be directly executed by a computer's central processing units (CPUs), graphics processing units (GPUs), etc., or executed through decoding, microcode, etc.

[0121] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0122] Figure 13 The components shown for the computer system 1300 are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing embodiments of the present application. Nor should the configuration of the components be construed as having any dependency or requirement on any one component or combination thereof shown in the exemplary embodiments of the computer system 1300.

[0123] The computer system 1300 may include certain human-machine interface input devices. Such human-machine interface input devices may respond to the input of one or more human users through tactile inputs (such as keyboard input, swiping, data glove movement), audio inputs (such as voice, applause), visual inputs (such as gestures), and olfactory inputs (not shown). The human-machine interface device may also be used to capture certain media, which need not be directly related to conscious human input, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0124] The human-machine interface input device may include one or more of the following (only one of which is shown): keyboard 1301, mouse 1302, touchpad 1303, touch screen 1310, joystick 1305, microphone 1306, scanner 1308, camera 1307.

[0125] The computer system 1300 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (such as tactile feedback through the touch screen 1310 or joystick 1305, but there may also be tactile feedback devices that do not function as input devices), audio output devices (such as speakers 1309, headphones (not shown)), visual output devices (such as screens 1310 including cathode ray tube (CRT) screens, liquid crystal screens, plasma screens, organic light-emitting diode screens, each of which may or may not have touch screen input functionality, each of which may or may not have tactile feedback functionality - some of which may output two-dimensional visual output or output above three dimensions through means such as stereoscopic picture output; virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)), and printers (not shown).

[0126] The computer system 1310 may also include human-accessible storage devices and their associated media, such as optical media including high-density read-only / rewritable compact discs (CD / DVD ROM / RW) 1320 with CD / DVD 1311 or similar media, thumb drives 1322, removable hard disk drives or solid state drives 1323, traditional magnetic media such as tapes and floppy disks (not shown), dedicated devices based on ROM / Application-Specific Integrated Circuit (ASIC) / Programmable Logic Device (PLD) such as security software protectors (not shown), and so on.

[0127] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the disclosed subject matter does not include a transmission medium, a carrier wave, or other transient signals.

[0128] Computer system 1300 may also include an interface 1399 to one or more communication networks 1398. For example, the network 1398 may be wireless, wired, optical. The network 1398 may also be a LAN, a wide area network, a metropolitan area network, a vehicular network, and an industrial network, a real-time network, a delay-tolerant network, etc. The network 1398 also includes Ethernet, a wireless local area network, a cellular network (Global System for Mobile communications (GSM), 3G, 4G, 5G, Long-Term Evolution (LTE), etc.), a LAN such as a television cable or wireless wide area digital network (including cable television, satellite television, and terrestrial broadcast television), a vehicular network, and an industrial network (including Controller Area Network Bus (CANBus)). Some networks 1398 typically require an external network interface adapter for connection to certain common data ports or peripheral buses (1350 and 1351) (e.g., the Universal Serial Bus (USB) port of computer system 1300). Other systems are typically integrated into the core of computer system 1300 by connecting to the following system buses (e.g., an Ethernet interface is integrated into a PC computer system or a cellular network interface is integrated into a smart phone computer system). By using any of these networks 1398, computer system 1300 can communicate with other entities. The communication can be one-way, only for receiving (e.g., wireless television), one-way only for sending (e.g., CAN bus to certain CAN bus devices), or two-way (e.g., through a local or wide area digital network to other computer systems). Each of the above networks and network interfaces may use certain protocols and protocol stacks.

[0129] The above-mentioned human-machine interface device, human-accessible storage device, and network interface can be connected to the core 1340 of computer system 1300.

[0130] The core 1340 may include one or more CPUs 1341, GPUs 1342, a graphics adapter 1317, a dedicated programmable processing unit in the form of a Field Programmable Gate Array (FGPA) 1343, a hardware accelerator 1344 for specific tasks, etc. These devices, as well as a read-only memory (ROM) 1345, a random access memory 1346, an internal mass storage (such as an internal non-user-accessible hard disk drive, a solid-state drive, etc.) 1347, etc. can be connected through a system bus 1348. In some computer systems, the system bus 1348 can be accessed in the form of one or more physical plugs so as to be expandable through additional CPUs, GPUs, etc. Peripheral devices can be directly attached to the system bus 1348 of the core or connected through a peripheral bus 1351. The architecture of the peripheral bus includes Peripheral Component Interconnect (PCI), USB, etc.

[0131] The CPU 1341, GPU 1342, FPGA 1343, and accelerator 1344 can execute certain instructions, and these instructions combined can constitute the above-mentioned computer code. This computer code can be stored in the ROM 1345 or the RAM 1346. Transitional data can also be stored in the RAM 1346, while permanent data can be stored in, for example, the internal mass storage 1347. Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with one or more CPUs 1341, GPUs 1342, the mass storage 1347, the ROM 1345, the RAM 1346, etc.

[0132] The computer-readable medium can have computer code for performing various computer-implemented operations. The medium and the computer code can be specially designed and constructed for the purposes of this application, or they can be media and codes well-known and available to those skilled in the computer software field.

[0133] By way of example and not limitation, a computer system having architecture 1300, particularly core 1340, can provide the functionality of a processor (including CPUs, GPUs, FPGAs, accelerators, etc.) to execute software contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with the user-accessible mass storage described above, as well as specific memories of the non-volatile core 1340, such as the core internal mass storage 1347 or ROM 1345. The software implementing the various embodiments of the present application can be stored in such devices and executed by core 1340. Depending on specific needs, the computer-readable media can include one or more storage devices or chips. The software can cause core 1340, particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute specific processes or specific parts of specific processes described herein, including defining data structures stored in Random Access Memory (RAM) 1346 and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system can provide logic hardwired or otherwise included in circuitry (e.g., in accelerator 1344) that can perform specific processes or specific parts of specific processes described herein instead of or in conjunction with the software. In appropriate cases, references to software can include logic, and vice versa. In appropriate cases, references to computer-readable media can include circuitry (such as an IC) that stores the software for execution, circuitry that contains the execution logic, or both. The present application encompasses any suitable combination of hardware and software.

[0134] Although the present application has described multiple exemplary embodiments, various changes, permutations, and various equivalent substitutions of the embodiments are within the scope of the present application. Therefore, it should be understood that those skilled in the art can design various systems and methods that, although not explicitly shown or described herein, embody the principles of the present application and are thus within the spirit and scope of the present application.

Claims

1. A method for processing tactile media, the method being performed by at least one processor, via a Dynamic Adaptive Streaming over Hypertext Transfer Protocol (HTTP) (DASH) application implemented by the at least one processor, the method comprising: receiving a media file including a haptic indicated by an adaptation set of a DASH media presentation description (MPD) file; parsing any of the tactile virtual object, perception, channel, and band identifiers; selecting a subset of a haptic adaptation set comprising the adaptation set based on parsing any of the virtual object, the perception, the channel, and the frequency band identifier; as well as In DASH, haptic feedback for the media file is streamed and rendered based on the selected subset of the haptic adaptation set.

2. The method according to claim 1, wherein: The DASH application is configured to command a DASH client to stream from a selected subset of the haptic adaptation set.

3. The method according to claim 2, wherein: The DASH application receives the media file from the DASH client.

4. The method according to claim 1, wherein: The media file also indicates the haptics via a Haptic Interchange Format (HJIF) profile.

5. The method according to claim 4, wherein: The media file also indicates the haptic by referencing a link to the HJIF file.

6. The method according to claim 1, wherein: The media file also indicates the haptics via a Haptic Interchange Format (HJIF) configuration object.

7. The method according to claim 1, wherein: The media file also indicates the haptic sensation via an extensible markup language (XML) element.

8. An apparatus for haptic processing of a Dynamic Adaptive Streaming over Hypertext Transfer Protocol (HTTP) (DASH) application implemented by at least one processor, the apparatus comprising: at least one memory configured to store computer program code; as well as The at least one processor is configured to access the computer program code and operate according to instructions of the computer program code, wherein the computer program code includes: receiving code configured to cause the at least one processor to receive a media file including a haptic indicated by an adaptation set of a DASH media presentation description (MPD) file; parsing code configured to cause the at least one processor to parse any of the haptic virtual object, perception, channel, and frequency band identifiers; selecting code configured to cause the at least one processor to select a subset of the haptic adaptation set comprising the adaptation set based on parsing any of the virtual object, the perception, the channel, and the frequency band identifier; and The streaming and rendering code is configured to cause the at least one processor to stream and render haptic feedback of the media file based on the selected subset of the haptic adaptation set in DASH.

9. The device according to claim 8, wherein: The DASH application is configured to command a DASH client to stream from a selected subset of the haptic adaptation set.

10. The device according to claim 9, wherein: The DASH application receives the media file from the DASH client.

11. The device according to claim 8, wherein: The media file also indicates the haptics via a Haptic Interchange Format (HJIF) profile.

12. The device according to claim 11, wherein The media file also indicates the haptic by referencing a link to the HJIF file.

13. The device according to claim 8, wherein: The media file also indicates the haptics via a Haptic Interchange Format (HJIF) configuration object.

14. The device according to claim 8, wherein: The media file also indicates the haptic sensation via an extensible markup language (XML) element.

15. A non-transitory computer-readable medium storing a program, the program causing a computer to execute a process, the process being performed by a Dynamic Adaptive Streaming over Hypertext Transfer Protocol (HTTP) (DASH) application implemented by the computer, the process comprising: receiving a media file including a haptic indicated by an adaptation set of a DASH media presentation description (MPD) file; parsing any of the tactile virtual object, perception, channel, and band identifiers; selecting a subset of a haptic adaptation set comprising the adaptation set based on parsing any of the virtual object, the perception, the channel, and the frequency band identifier; as well as In DASH, haptic feedback for the media file is streamed and rendered based on the selected subset of the haptic adaptation set.

16. The non-transitory computer readable medium of claim 15, wherein: The DASH application is configured to command a DASH client to stream from a selected subset of the haptic adaptation set.

17. The non-transitory computer readable medium of claim 16, wherein: The DASH application receives the media file from the DASH client.

18. The non-transitory computer readable medium of claim 15, wherein: The media file also indicates the haptics via a Haptic Interchange Format (HJIF) profile.

19. The non-transitory computer readable medium of claim 18, wherein: The media file also indicates the haptic by referencing a link to the HJIF file.

20. The non-transitory computer readable medium of claim 15, wherein: The media file also indicates the haptics via at least one of a Haptic Interchange Format (HJIF) configuration object and an Extensible Markup Language (XML) element.