Method for signaling type of iso-based media file format haptic sample

By analyzing the ISOBMFF haptic sample syntax of media streams, using four-bit binary form to represent the sample type, the problem of unclear tactile track timing model is solved, and the synchronous decoding of the haptic effect with other media tracks is achieved, which improves the efficiency of multimedia presentation.

CN120239845APending Publication Date: 2025-07-01TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004967.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-16
Filing Date
2024-04-17
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In multimedia presentation, the carrying timing model of the tactile track is unclear, and the existing JSON format fails to effectively carry binary wavelet-encoded streams, resulting in the timing of the tactile signal not matching other media signals.

Method used

By receiving data in a tactile exchange format, the syntax of ISOBMFF haptic samples in the media stream is adopted, and the sample type is represented using four-bit binary form, including the existence of silent samples, time effect data packets and spatial effect data packets, to achieve efficient decoding of the media stream.

Benefits of technology

An efficient timing model is provided to synchronize haptic effects with other media tracks in ISOBMFF files, improving the decoding efficiency of media streams and file navigation speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120239845A_ABST
    Figure CN120239845A_ABST
Patent Text Reader

Abstract

A method, apparatus and system for haptic signal processing are provided, the method comprising receiving a media stream comprising data in a haptic exchange format, obtaining a syntax of ISOBMFF haptic samples from the data in the haptic exchange format of the media stream, the syntax indicating a variable value of a datapackettype of the media stream; and decoding the media stream based on the syntax.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the priority of U.S. Provisional Application No. 63 / 459,930, filed on April 17, 2023, and U.S. Application No. 18 / 636,752, filed on April 16, 2024, the disclosures of which are incorporated herein by reference in their entireties. Technical field

[0003] This disclosure relates to a set of advanced video coding and decoding techniques. More specifically, this disclosure relates to encoding and decoding haptic experiences for multimedia presentations, and methods for carrying binary wavelet streams in a haptic interchange format. Background art

[0004] Haptic experiences have become part of multimedia presentations. In applications where multimedia presentations include aspects of haptic experiences, haptic signals can be delivered to devices or wearable devices, and users can feel haptic sensations in coordination with visual and / or audio media experiences during the use of the application.

[0005] Recognizing the increasing popularity of haptic experiences in multimedia presentations, the Moving Picture Experts Group (MPEG) has started to study compression standards for haptics (for both MPEG - DASH and MPEG - I) and carrying compressed haptic signaling in an ISO - based media file format (ISOBMFF).

[0006] One of the problems to be solved in aspects of multimedia presentations involving haptic experiences is that the timing model for carrying haptic tracks is not clear, i.e., it is not clear how the timing of ISOBMFF tracks is related to the timing of haptic elementary signals. A solution to this problem is needed.

[0007] The haptic committee draft includes a JSON and a binary format. The current JSON format (referred to as the haptic interchange format) carries quantized wavelet coefficients rather than carrying a binary wavelet - encoded stream. A solution to this problem is needed. Summary of the invention

[0008] According to one aspect of the present disclosure, there is provided an apparatus, and similarly, a method and a computer-readable medium. The apparatus includes at least one memory configured to store computer program code; and at least one processor configured to access the computer program code and operate in accordance with what is indicated by the computer program code. The computer program code includes: receiving code configured to cause the at least one processor to receive a media stream including data in a haptic exchange format; obtaining code configured to cause the at least one processor to obtain a syntax of ISO BMFF haptic samples from the data in the haptic exchange format of the media stream, wherein the syntax indicates variable values of data_packet_type of the media stream; and decoding code configured to cause the at least one processor to decode the media stream based on the syntax.

[0009] The syntax can be represented in binary form by four bits of the media stream.

[0010] The first of the four bits can indicate whether the sample in the media stream is a silent sample.

[0011] The second of the four bits can indicate whether the sample includes one or more temporal effect data packets.

[0012] Decoding the media stream can include the following interpretation in binary form: the first bit being 0 indicates that the sample is silent, and the first bit being 1 and the second bit being either 0 or 1 indicates that the sample includes one or more temporal effect data packets.

[0013] The third of the four bits can indicate whether the sample includes one or more spatial effect data packets.

[0014] Decoding the media stream can further include the following interpretation in binary form: the first bit being 0 indicates that the sample is silent, and the first bit being 1 and the third bit being either 0 or 1 indicates that the sample includes one or more spatial effect data packets.

[0015] Decoding the media stream can further include the following interpretation in binary form: the first bit being 1 and the third bit and the second bit being either 0 or 1 indicates that the sample includes one or more spatial effect data packets when the second bit is 0 and when the second bit is 1.

[0016] The binary representation of the syntax can consist of four bits.

[0017] Additional embodiments will be set forth in the following description and will be apparent, in part, from the description, and / or may be learned by practice of the embodiments presented in the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0019] Figure 1 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment of the present disclosure;

[0020] Figure 2 is a schematic diagram of a simplified block diagram of a streaming system according to an embodiment of the present disclosure;

[0021] Figure 3 is an example illustration according to an embodiment of the present disclosure;

[0022] Figure 4 is an example illustration according to an embodiment of the present disclosure;

[0023] Figure 5A is an example illustration according to an embodiment of the present disclosure;

[0024] Figure 5B is an example illustration according to an embodiment of the present disclosure;

[0025] Figure 6 is an exemplary flowchart showing a process for processing tactile media according to an embodiment of the present disclosure;

[0026] Figure 7 is an exemplary diagram showing some aspects according to an embodiment of the present disclosure;

[0027] Figure 8 is an exemplary diagram showing some aspects according to an embodiment of the present disclosure;

[0028] Figure 9 is an exemplary diagram showing some aspects according to an embodiment of the present disclosure; and

[0029] Figure 10 is an exemplary diagram showing some aspects according to an embodiment of the present disclosure. Detailed Description

[0030] According to one aspect of the present disclosure, there are provided methods, systems, and non-transitory storage media for parallel processing of dynamic mesh compression. Embodiments of the present disclosure may also be applied to static meshes.

[0031] Reference Figure 1 and Figure 2 describe embodiments of the present disclosure for implementing the encoding and decoding structures of the present disclosure.

[0032] Figure 1FIG. 0 shows a simplified block diagram of a communication system 100 according to an embodiment of the present disclosure. The system 100 may include at least two terminals 110, 120 interconnected via a network 150. For unidirectional transmission of data, a first terminal 110 may encode video data (which may include mesh data) at a local location for transmission via the network 150 to another terminal 120. The second terminal 120 may receive the encoded video data of another terminal from the network 150, decode the encoded data, and display the recovered video data. Unidirectional data transmission may be common in media service applications and the like.

[0033] Figure 1 FIG. 4 shows a second pair of terminals 130, 140 provided to support two-way transmission of encoded video that may occur, for example, during a video conference. For two-way transmission of data, each terminal 130, 140 may encode video data captured at a local location for transmission via the network 150 to another terminal. Each terminal 130, 140 may also receive the encoded video data sent by another terminal, may decode the encoded data, and may display the recovered video data at a local display device.

[0034] In Figure 1 FIG. 9, the terminals 110 - 140 may be, for example, servers, personal computers, and smart phones and / or any other type of terminal. For example, the terminals (110 - 140) may be laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network 150 represents any number of networks that convey encoded video data between the terminals 110 - 140, including, for example, wired and / or wireless communication networks. The communication network 150 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, unless otherwise explained below, the architecture and topology of the network 150 may be immaterial to the operation of the present disclosure.

[0035] Figure 2 FIG. 13 shows the placement of video encoders and decoders in a streaming environment as an example of an application for the disclosed subject matter. The disclosed subject matter may be used with other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, and the like.

[0036] As Figure 2 shown in FIG. 18, a streaming system 200 may include a capture subsystem 213 that includes a video source 201 and an encoder 203. The streaming system 200 may also include at least one streaming server 205 and / or at least one streaming client 206.

[0037] Video source 201 may create a stream 202 that includes, for example, a 3D mesh and metadata associated with the 3D mesh. The video source 201 may include, for example, a 3D sensor (e.g., a depth sensor) or 3D imaging technology (e.g., a digital camera) and a computing device configured to generate a 3D mesh using data received from the 3D sensor or 3D imaging technology. The sample stream 202, which may have a high data volume when compared to an encoded video bitstream, may be processed by an encoder 203 coupled to the video source 201. The encoder 203 may include hardware, software, or a combination thereof to implement or realize aspects of the disclosed subject matter, as described in more detail below. The encoder 203 may also produce an encoded video bitstream 204. The encoded video bitstream 204, which may have a lower data volume when compared to the uncompressed stream 202, may be stored on a streaming server 205 for future use. One or more streaming clients 206 may access the streaming server 205 to retrieve a video bitstream 209, which may be a copy of the encoded video bitstream 204.

[0038] The streaming client 206 may include a video decoder 210 and a display 212. The video decoder 210 may, for example, decode the video bitstream 209 (which is an input copy of the encoded video bitstream 204) and create an output video sample stream 211 that can be rendered on the display 212 or another rendering device (not depicted). In some streaming systems, the video bitstreams 204, 209 may be encoded according to certain video coding / compression standards.

[0039] Figure 3 May be a functional block diagram of a video decoder 300 according to an embodiment of the present invention.

[0040] A receiver 302 may receive one or more codec video sequences to be decoded by the decoder 300; in the same or another embodiment, one encoded video sequence at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequences may be received from a channel 301, which may be a hardware / software link to a storage device storing the encoded video data. The receiver 302 may receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams, that may be forwarded to their respective using entities (not depicted). The receiver 302 may separate the encoded video sequences from the other data. To counter network jitter, a buffer memory 303 may be coupled between the receiver 302 and an entropy decoder / parser 304 (hereinafter referred to as "parser"). When the receiver 302 receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer 303 may not be needed or may be small. For use on a best-effort packet network such as the Internet, the buffer 303 may be necessary, may be relatively large, and may advantageously have an adaptive size.

[0041] Video decoder 300 may include a parser 304 to reconstruct symbols 313 from an entropy-coded video sequence. The categories of these symbols include information for managing the operation of decoder 300 and potential information for controlling a display device (e.g., display 312), which is not part of the decoder but may be coupled to the decoder. The control information for the display device may be Supplemental Enhancement Information (SEI messages) or a parameter set segment of video usability information (not labeled). The parser 304 may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence may be performed according to video coding techniques or standards and may follow principles well-known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser 304 may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder based on at least one parameter corresponding to a group. Subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), and so on. The entropy encoder / parser may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.

[0042] The parser 304 may perform entropy decoding / parsing operations on the video sequence received from the buffer memory 303 to create symbols 313. The parser 304 may receive the encoded data and selectively decode specific symbols 313. In addition, the parser 304 may determine whether to provide specific symbols 313 to the motion compensation prediction unit 306, the scaler / inverse transform unit 305, the intra prediction unit 307, or the loop filter 311.

[0043] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of symbols 313 may involve multiple different units. Which units are involved and the way they are involved may be controlled by subgroup control information parsed by the parser 304 from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser 304 and multiple units below are not described.

[0044] In addition to the functional blocks already mentioned, decoder 300 can conceptually be subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the functional units below.

[0045] The first unit is the scaler / inverse transform unit 305. The scaler / inverse transform unit 305 receives, from the parser, quantized transform coefficients as symbols 313 and control information, including which transform mode to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit can output a block including sample values, which can be input into the aggregator 310.

[0046] In some cases, the output samples of the scaler / inverse transform unit 305 can belong to intra-coded blocks; that is, blocks that do not use predictive information from previously reconstructed pictures, but can use predictive information from previously reconstructed parts of the current picture. Such predictive information can be provided by the intra-picture prediction unit 307. In some cases, the intra-picture prediction unit 307 generates surrounding blocks of the same size and shape as the block being reconstructed using the reconstructed information extracted from the current partially reconstructed picture 309. In some cases, based on each sample, the aggregator 310 adds the prediction information generated by the intra-picture prediction unit 307 to the output sample information provided by the scaler / inverse transform unit 305.

[0047] In other cases, the output samples of the scaler / inverse transform unit 305 can belong to inter-coded and potentially motion-compensated blocks. In this case, the motion compensation prediction unit 306 can access the reference picture memory 308 to extract samples for prediction. After motion compensation of the extracted samples according to the symbol 313, these samples can be added by the aggregator 310 to the output of the scaler / inverse transform unit (referred to as residual samples or residual signals in this case), thereby generating output sample information. The motion compensation unit obtaining prediction samples from an address within the reference picture memory can be controlled by a motion vector, and the motion vector is in the form of the symbol for use by the motion compensation unit), such as a symbol including X, Y, and reference picture components. Motion compensation can also include interpolation of sample values extracted from the reference picture memory when using sub-sample accurate motion vectors, motion vector prediction mechanisms, and so on.

[0048] The output samples of aggregator 310 can be employed by various loop filtering techniques in loop filter unit 311. Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the encoded bitstream, and the parameters can be available to loop filter unit 311 as symbols 313 from parser 304. However, in other embodiments, video compression techniques can also respond to meta-information obtained during decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0049] The output of loop filter unit 311 can be a sample stream that can be output to display device 312 and stored in reference picture memory 557 for subsequent inter-picture prediction.

[0050] Once fully reconstructed, some encoded pictures can be used as reference pictures for future prediction. Once an encoded picture is fully reconstructed and the encoded picture is identified (e.g., by parser 304) as a reference picture, the current reference picture 309 can become part of reference picture buffer 308, and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.

[0051] Video decoder 300 can perform decoding operations according to a predetermined video compression technique that can be recorded, for example, in the ITU-T H.265 standard. As specified in a video compression technique document or standard (particularly in the profile of this application), an encoded video sequence can conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows the syntax of the video compression technique or standard. For compliance, it is also required that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata for HRD buffer management signaled in the encoded video sequence.

[0052] In an embodiment, receiver 302 can receive additional (redundant) data along with the encoded video. The additional data can be part of the encoded video sequence. The additional data can be used by video decoder 300 to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0053] Figure 4 It may be a functional block diagram of the video encoder 400 according to the embodiments disclosed in the present application.

[0054] The encoder 400 may receive video samples from a video source 401 (not part of the encoder), and the video source may capture video images to be encoded by the encoder 400.

[0055] The video source 401 may provide a source video sequence in the form of a digital video sample stream to be encoded by the encoder 303. The digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, etc.), any color space (e.g., BT.601 YCrCb, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source 401 may be a storage device storing previously prepared videos. In a video conferencing system, the video source 401 may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be constructed as a spatial pixel array, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on the description of samples.

[0056] According to an embodiment, the encoder 400 may encode and compress the pictures of the source video sequence into an encoded video sequence 410 in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller 402. The controller controls other functional units as described below and is functionally coupled to these units. For the sake of simplicity, the couplings are not labeled in the figure. The parameters set by the controller may include rate control related parameters (picture skipping, quantizer, λ value of rate distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. When other suitable functions are involved in the video encoder 400 optimized for a certain system design, those skilled in the art can easily identify these other suitable functions of the controller 402.

[0057] Some video encoders operate in an "encoding loop" that is readily recognizable to those skilled in the art. As a simple description, the encoding loop may include an encoding portion of encoder 402 (referred to hereinafter as the "source encoder") that is responsible for creating symbols based on the input picture to be encoded and reference pictures, and a (local) decoder 406 embedded in encoder 400 that reconstructs the symbols in the same manner as the (remote) decoder creates sample data to create sample data (since in the video compression techniques contemplated in this application, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream is input to reference picture memory 405. Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the reference picture buffer contents are also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction portion of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is well known to those skilled in the art.

[0058] The operation of the "local" decoder 406 may be the same as that of the "remote" decoder 300 described in detail above in connection with Figure 3 However, briefly referring additionally to Figure 4 , when the symbols are available and the entropy encoder 408 and parser 304 can encode / decode the symbols losslessly into the encoded video sequence, the entropy decoding portion of decoder 300 (including channel 301, receiver 302, buffer 303, and parser 304) may not be fully implementable in local decoder 406.

[0059] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding present in the decoder must also exist in the corresponding encoder in substantially the same functional form. The description of the encoder technology can be simplified because the encoder technology is inverse to the decoder technology described comprehensively. More detailed description is only required in certain areas and is provided below.

[0060] As part of its operation, the source encoder 403 may perform motion-compensated predictive coding. Referencing one or more previously encoded frames in the video sequence designated as "reference frames", the motion-compensated predictive coding performs predictive coding on the input frame. In this way, the coding engine 407 encodes the difference between the pixel blocks of the input frame and the pixel blocks of the reference frame, and the reference frame can be selected as the prediction reference for the input frame.

[0061] The local video decoder 406 can decode the encoded video data of the frames that can be designated as reference frames based on the symbols created by the source encoder 403. The operation of the encoding engine 407 can be a lossy process. When the encoded video data can be decoded at the video decoder ( Figure 4 not shown), the reconstructed video sequence can generally be a copy of the source video sequence with some errors. The local video decoder 406 replicates the decoding process that can be performed by the video decoder on the reference frames and can store the reconstructed reference frames in the reference picture cache 405. In this way, the encoder 400 can locally store a copy of the reconstructed reference frames, which has the same content (without transmission errors) as the reconstructed reference frames that will be obtained by the remote video decoder.

[0062] The predictor 404 can perform a prediction search for the encoding engine 407. That is, for a new frame to be encoded, the predictor 404 can search in the reference picture memory 405 for sample data (as candidate reference pixel blocks) or some metadata that can serve as an appropriate prediction reference for the new picture, such as reference picture motion vectors, block shapes, etc. The predictor 404 can operate block by block based on sample blocks to find a suitable prediction reference. In some cases, according to the search results obtained by the predictor 404, it can be determined that the input picture can have prediction references taken from multiple reference pictures stored in the reference picture memory 405.

[0063] The controller 402 can manage the encoding operations of the video encoder 403, including, for example, setting parameters and subgroup parameters for encoding the video data.

[0064] The outputs of all the above functional units can be entropy encoded in the entropy encoder 408. The entropy encoder performs lossless compression on the symbols generated by various functional units according to techniques well known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc., so as to convert the symbols into an encoded video sequence.

[0065] The transmitter 409 can buffer the encoded video sequence created by the entropy encoder 408 to prepare for transmission through the communication channel 411, which can be a hardware / software link leading to a storage device that will store the encoded video data. The transmitter 409 can merge the encoded video data from the video encoder 403 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0066] The controller 402 can manage the operation of the encoder 400. During encoding, the controller 405 can assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques that can be applied to the corresponding picture. For example, pictures can generally be assigned to any of the following frame types:

[0067] An intra picture (I picture) is a picture that can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh pictures. Those skilled in the art are aware of the variations of I pictures and their corresponding applications and characteristics.

[0068] A predictive picture (P picture) is a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.

[0069] A bi - predictive picture (B picture) is a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.

[0070] Source pictures can generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block - by - block. These blocks can be prediction - encoded with reference to other (already - encoded) blocks, which are determined according to the coding assignment of the corresponding picture applied to the 'block'. For example, blocks of an I picture can be non - prediction - encoded, or the blocks can be prediction - encoded with reference to already - encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be non - prediction - encoded by spatial prediction or by temporal prediction with reference to a previously - encoded reference picture. Blocks of a B picture can be non - prediction - encoded by spatial prediction or by temporal prediction with reference to one or two previously - encoded reference pictures.

[0071] Video encoder 400 can perform encoding operations according to a predetermined video coding technique or standard, such as the ITU - T H.265 recommendation. In operation, video encoder 400 can perform various compression operations, including prediction - encoding operations that utilize the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard used.

[0072] In an embodiment, transmitter 409 can transmit additional data when transmitting the encoded video. Source encoder 403 can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures, and slices, supplementary enhancement information (SEI) messages, video usability information (VUI) parameter set fragments, etc.

[0073] Reference Figures 5A to 5B, embodiments of the present disclosure for implementing the haptic encoder 500 and the haptic decoder 550 are described.

[0074] As Figure 5A shown, the haptic encoder 500 can receive descriptive haptic data and waveform haptic data. Thus, the haptic encoder 500 can be capable of processing three types of input files:.ohm metadata files (object haptic metadata - a text file format for haptic metadata), descriptive haptic files (.ivs,.ahap, and.hjif), or waveform PCM files (.wav). Examples of descriptive data can include:.ahap (Apple Haptic and Audio Pattern - a JSON-like file format specifying haptic patterns) from Apple (representing the expected haptic output through a set of modulated continuous signals and a set of parameterized modulation transients);.ivs (representing the expected haptic output through a set of base effects parameterized by a set of parameters) from Immersion; or.hjif (haptic JSON interchange format), i.e., the MPEG format proposed in this application. Examples of waveform pulse code modulation (PCM) signals can include.ohm input files that include metadata information.

[0075] According to an embodiment, the haptic encoder 500 can process two types of input files differently. For descriptive content, the haptic encoder 500 can semantically analyze the input to transcoder (if necessary) the data into the encoding representation proposed in this application.

[0076] According to an embodiment, the.ohm metadata input file can include a description of the haptic system and settings. In particular, the.ohm metadata input file can include the name of each associated haptic file (descriptive or PCM), and a description of the signal. The.ohm metadata input file also provides a mapping between each channel of the signal and the target body part on the user's body. For the.ohm metadata input file, the haptic encoder performs metadata extraction by retrieving the associated haptic file from the URI, encodes the metadata based on the type of metadata, and maps the metadata extracted from the.ohm file to the metadata information of the data model.

[0077] According to an embodiment, descriptive haptic files (e.g.,.ivs,.ahap, and.hjif) can be encoded through a simple process. The haptic encoder 500 first specifically identifies the input format. If the input format is a.hjif file, then no transcoding is required, and the file can be further edited, compressed into a binary format, and finally packaged into a MIHS stream. If an.ahap or.ivs input file is used, transcoding is necessary. The haptic encoder 500 first semantically analyzes the input file information and transcodes it to be formatted into a selected data model. After transcoding, the data can be exported as a.hjif file,.hmpg binary file, or MIHS stream.

[0078] According to an embodiment, the haptic encoder 500 can perform signal analysis to parse the signal structure of a.wav file and convert it into the encoding representation proposed in this application. For waveform PCM content, the signal analysis process can be divided into two sub-processes performed by the haptic encoder 500. After performing band decomposition on the signal, in the first sub-process, the low frequencies can be encoded using a keyframe extraction process. Then the low-frequency band can be reconstructed, and the error between the signal and the original low-frequency signal can be calculated. Then, before encoding using wavelet transform, which is the second sub-process, the residual signal can be added to the original high-frequency band. According to an embodiment, when multiple low-frequency bands are used, the residuals from all low-frequency bands are added to the high-frequency band before encoding. In an embodiment using multiple high-frequency bands, the residuals from the low-frequency band are added to the first high-frequency band before encoding.

[0079] According to an embodiment, keyframe extraction includes obtaining the lower frequency bands from the band decomposition and analyzing their content in the time domain. According to an embodiment, wavelet processing can include obtaining the high-frequency bands from the band decomposition and the low-frequency residuals and dividing them into equally sized blocks. Then these equally sized signal blocks are analyzed in a psychohaptic model. Lossy compression can be applied by performing wavelet transform on the blocks and quantifying them with the help of the psychohaptic model. Finally, each block is saved into a separate effect in a single frequency band, which is done in formatting. Binary compression can apply lossless compression using appropriate coding techniques (e.g., the set partitioning in hierarchical trees (SPIHT) algorithm and arithmetic coding (AC)).

[0080] As Figure 5AAs shown, the tactile encoder 500 can be configured to encode descriptive and quantified tactile data and can output three types of formats - an interchange format (.hjif), a binary compressed format (.hmpg), and a streaming format (e.g., MPEG immersive haptic stream (MIHS)). The.hjif format is a JSON-based human-readable format that can be easily parsed and manually edited, making it an ideal interchange format, especially when designing / creating content. For distribution purposes, the.hjif data can be compressed into a more memory-efficient binary.hmpg bitstream. This compression may be lossy, and different parameters affect the encoding depth of the amplitudes and frequencies that make up the bitstream. For streaming purposes, the data can be compressed and packaged into an MPEG-1 haptic stream (MIHS). The above three formats have complementary purposes and can be lossily converted one-to-one between them.

[0081] As Figure 5B shown, the tactile decoder 550 can take as input a.hmpg compressed binary file format or an MIHS bitstream. The tactile decoder 550 can output the.hjif interchange format that can be directly used for rendering. These two input formats can undergo binary decompression to extract the metadata and the data itself from the file and map it to a selected data structure. Then, the data can be exported to the tactile renderer 580 in the.hjif format.

[0082] As Figure 5B shown, the renderer 580 includes a synthesizer. The synthesizer can render tactile data from the.hjif input file into a PCM output file. The rendering and / or synthesis is illustrative. According to an embodiment, the synthesizer parses the input file and performs high-level synthesis distributions between vectors, wavelets, etc. Then, the synthesis process descends to the band components of the codec where the synthesis process is called. Then, all the bands of a given channel are mixed through a simple addition operator to recreate the desired tactile signal.

[0083] According to an embodiment, the tactile experience defines the root of a hierarchical data model. It provides information about the file date and format version, it describes the tactile experience, it lists the different avatars (i.e., body representations) used throughout the experience, and it defines all tactile perceptions.

[0084] According to an embodiment, a self - contained stream format for transmitting MPEG - I haptic data may use a packetization method and may include two levels of packetization: an MPEG - I Haptic Stream (MIHS) unit, which covers a duration and includes zero or more MIHS packets; and an MIHS packet that includes metadata or haptic effect data. In an embodiment, the MIHS unit may be referred to as a network abstraction layer unit associated with haptic data. In an embodiment, the MIHS unit may be referred to as an MIHS sample associated with haptic data.

[0085] According to an embodiment, the MIHS unit may be a synchronous unit or an asynchronous unit. The synchronous unit resets the previous effect and thus provides a haptic experience independent of the previous MIHS unit. The asynchronous unit is a continuation of the previous MIHS unit and cannot be independently decoded and rendered without decoding the previous MIHS unit.

[0086] According to an embodiment, haptic signals may be encoded on multiple channels. In some embodiments, a haptic channel may define a signal to be rendered at a specific body location using a dedicated actuator / device. Metadata stored at the channel level may include information such as gain associated with the channel, mixing weights, the desired body location of the haptic feedback, and optionally a reference device and / or orientation. Additional information such as a desired sampling frequency or sampling count may also be provided. Finally, the haptic data of a channel is contained in a set of haptic frequency bands defined by its frequency range. The haptic frequency bands describe the haptic signals of a channel within a given frequency range. The bands are defined by a list of types and orders of haptic effects, and each haptic effect includes a set of keyframes. For each type of haptic frequency band, the haptic effect may be defined by at least a position and a type. The position may indicate the temporal or spatial location of the effect. In some embodiments, the value 0 is the relative starting position of the experience, which is a dependent variable of the configured perceptual modality. The default unit for temporal haptic feedback may be milliseconds, while the default unit for spatial haptic feedback may be millimeters. This embodiment discloses the "starting position of the experience" because the binary distribution format does not have the concept of any finite time interval (i.e., frame or sample).

[0087] According to the type of the band and the type of the effect, additional attributes may be specified, including phase, base signal, composition, and multiple consecutive haptic keyframes describing the effect.

[0088] According to an embodiment, a haptic data hierarchy is defined in the present disclosure.

[0089] · Haptic channel

[0090] o Haptic frequency band

[0091] ■ Haptic effect

[0092] Embodiments of the present disclosure describe two anchor points for the position of haptic effects relative to an ISOBMFF track.

[0093] Figure 6 A first embodiment 600 is shown. As Figure 6 shown, each MIHS unit (also referred to as a MIHS sample, an ISOBMFF haptic sample, or a sample in the embodiments) includes one or more haptic channel information and one or more haptic band information. As described above, each MIHS unit consists of one or more channels, and each channel consists of one or more bands. Then each band can have one or more effects.

[0094] In the first embodiment, the temporal position of an effect can be defined as an offset relative to the start timing of the sample carrying the effect (e.g., the MIHS unit start time). In a second embodiment or the same embodiment, the offset is based on the start time and / or presentation time of the media or haptic track.

[0095] According to an embodiment, the first embodiment enables manipulation of the track without affecting the position of the haptic effect, because any change in the ISOBMFF sampling timing does not affect the relative position of the effect. According to an embodiment, in the case of a basic haptic stream (e.g., a high-level syntax stream), the second embodiment can be used when the basic stream is used without ISOBMFF.

[0096] According to an embodiment, multiple types of haptic tracks can be used. In an embodiment, a sample or MIHS unit can be used in a haptic track, and the temporal position of the effect of the sample or MIHS unit is defined as an offset relative to the start timing of the sample. According to another embodiment, for example Figure 7 in example 700, a sample or MIHS unit is used in a haptic track, and the effect of the sample or MIHS unit has a temporal position relative to the start time of the track. In another embodiment, a mixed MIHS unit or sample can be used.

[0097] Embodiments of the present disclosure provide a timing model that can be used to synchronize haptic effects with other media tracks in the same or related ISOBMFF file. In the case where the timing model of the haptic track is relative to the timing model of the related ISOBMFF file, the manipulation and processing of the media track become more efficient.

[0098] As Figure 8 shown, process 800 shows an exemplary process for decoding haptic data.

[0099] At operation 805, a media stream including one or more haptic tracks and one or more video tracks can be received.

[0100] At operation 810, one or more Moving Picture Experts Group (MPEG) Immersive Haptic Streams (MIHS) units can be obtained from a media stream. In some embodiments, the MIHS unit can include one or more haptic effects. The MIHS unit can also include a start time of the MIHS unit.

[0101] In an embodiment, the MIHS unit is associated with at least one haptic channel, the at least one haptic channel includes one or more haptic bands, and each of the one or more haptic bands has at least one haptic effect.

[0102] At operation 815, timing information associated with one or more haptic effects can be obtained. In an embodiment, the timing information can include at least one time position of the one or more haptic effects.

[0103] In an embodiment, the time position of the haptic effect indicates an effect start time of the haptic effect, wherein the effect start time of the haptic effect is an offset based on the start time of the corresponding MIHS unit. The effect start time can indicate a start time of the haptic effect relative to the start time of the corresponding MIHS unit.

[0104] In an embodiment, the effect start time of the haptic effect is an absolute time based on the start time of at least one haptic track or at least one video track.

[0105] At operation 820, the media stream is rendered based on the obtained timing information.

[0106] According to an embodiment, manipulation of the order of one or more MIHS units does not affect at least one time position of one or more haptic effects because the one or more MIHS units correspond to one or more ISO-based media file format (ISOBMFF) samples associated with at least one video track.

[0107] In some embodiments, synchronized MIHS units can be obtained from the media stream. In an embodiment, the synchronized MIHS unit is a special type of MIHS unit configured to provide a reset point in the bitstream. In an embodiment, the synchronized MIHS unit is mapped to synchronization samples in the video bitstream corresponding to one or more haptic channels.

[0108] As in Figure 9 Example 900, a haptic encoder according to an embodiment herein generates a compact and efficient binary distribution format (.hmpg) for distribution. A haptic decoder can decode this format and send it to a renderer.

[0109] The ISOBMFF haptic binding Working draft defines the following for haptic samples in ISOBMFF tracks:

[0110]

[0111]

[0112] As described above, the data_packet_type is set to a fixed value.

[0113] An ISOBMFF haptic sample according to an embodiment of the present document may have the following data groupings:

[0114] 1. A silent grouping with no data, i.e., a zero-data grouping.

[0115] 2. One or more time data groupings.

[0116] 3. One or more spatial data groupings.

[0117] 4. A combination of 2 and 3

[0118] 5. A possible combination of 2 and 3, without any commitment, i.e., zero or more time and / or data groupings.

[0119] Embodiments of the present document may use the data_packet_type to signal the above conditions, such that the data_packet_type (i.e., b3b2b1b0) may be signaled as follows:

[0120] Table 1 - Data grouping / sample type

[0121]

[0122] Where, according to an embodiment, data_packet_type = 0 indicates that the sample is a silent sample, i.e., equivalent to one or more consecutive MIHS silent units.

[0123] Therefore, the following values of data_packet_type in Table 2 have the following meanings:

[0124] Table 2

[0125]

[0126] And, according to an embodiment, the file format parser may use the data_packet_type in the following manner:

[0127] 1. Identify silent samples when processing a file and skip further parsing of the silent samples, thereby accelerating random access and fast forward / rewind navigation through the file.

[0128] 2. Identify which samples have only time data groups for extracting time effects.

[0129] 3. Identify which samples have only spatial data groups for extracting spatial effects.

[0130] 4. When decomposing an orbit into multiple orbits, it is easier to separate time samples and spatial samples.

[0131] 5. When combining multiple orbits into a single orbit, determine the sample structure of the new orbit and whether the samples should combine time groups and spatial groups or should they be kept in separate samples.

[0132] Accordingly, according to an embodiment, a method for signaling the sample type in an ISOBMFF haptic sample is presented, wherein the sample is identified as a silent sample, a non - silent sample, a non - silent sample having only time effects, a non - silent sample having only spatial effects, a non - silent sample having a combination of time and spatial effects, wherein signaling is implemented using bit - based flags for different features, wherein file format parsing can utilize the information provided by the sample type and navigate through the file faster due to skipping silent samples, or use the sample type information to extract only spatial information, only time information, or only silent information, or use the information for bitstream manipulation, single - track to multi - track conversion, and multi - track to single - track conversion.

[0133] As in Figure 9 Example 900, a haptic encoder according to an embodiment of the present disclosure generates a compact and efficient binary distribution format (.hmpg) for distribution. A haptic decoder can decode this format and send it to a renderer. However, on the other hand, in the absence of the embodiments of the present disclosure, the haptic exchange format (.hjif) does not have binary compression, and thus, wavelet coefficients are stored in the.hjif instead of in a compressed bitstream. And in view of this, also according to an embodiment, the embodiments herein extend the band type in the.hjif format to include a new option: binary wavelet.

[0134] Those skilled in the art will understand that the techniques described herein can be implemented on both the encoder side and the decoder side. The above - mentioned techniques can be implemented as computer software using computer - readable instructions and physically stored on one or more computer - readable media. For example, Figure 10 FIG. shows a computer system 1000 suitable for implementing certain embodiments of the present disclosure.

[0135] Computer software can be encoded using any suitable machine code or computer language that can be subjected to assembly, compilation, linking, or similar mechanisms to create code that includes instructions that can be executed directly by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or executed through interpretation, microcode execution, etc.

[0136] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0137] Figure 10 The components shown for computer system 1000 are examples and are not intended to imply any limitation on the scope of use or functionality of the computer software implementing embodiments of the present disclosure. The configuration of the components should also not be construed as having any dependencies or requirements related to any one or combination of the components shown in the non-limiting embodiments of computer system 1000.

[0138] Computer system 1000 may include certain human-machine interface input devices. Such human-machine interface input devices can respond to input by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not depicted). The human-machine interface device can also be used to capture certain media that is not necessarily directly related to conscious human input, such as, for example, audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0139] The input human-machine interface devices can include one or more of the following (only one of each is depicted): keyboard 1001, mouse 1002, touchpad 1003, touch screen 1010, data glove, joystick 1005, microphone 1006, scanner 1007, camera 1008.

[0140] The computer system 1000 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, haptic output, sound, light, and smell / taste. Such human-machine interface output devices may include haptic output devices (e.g., haptic feedback provided by the touch screen 1010, data glove, or joystick 1005, although there may also be haptic feedback devices that do not function as input devices). For example, such devices may be audio output devices (e.g., speakers 1009, headphones (not shown)), visual output devices (e.g., screen 1010, including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without haptic feedback capabilities - some of which may be capable of outputting two-dimensional visual output or more than three-dimensional output through means such as stereoscopic output, virtual reality glasses (not shown), holographic displays, and fog machines (not shown)), and printers (not shown).

[0141] The computer system 1000 may also include human-accessible storage devices and their associated media, such as optical media, including CD / DVD ROM / RW 1020 with CD / DVD or similar media 1021, thumb drives 1022, removable hard disk drives or solid state drives 1023, traditional magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), etc.

[0142] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.

[0143] The computer system 1000 may also include an interface to one or more communication networks. For example, the network can be wireless, wired, optical. The network can also be local, wide area, metropolitan area network, in-vehicle and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks (such as Ethernet), wireless LANs, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), TV wired or wireless wide area networks (including cable TV, satellite TV, and terrestrial broadcast TV), in-vehicle and industrial networks (including CANBus), etc. Some networks typically require an external network interface adapter, which is attached to some general-purpose data ports or peripheral buses 1049 (e.g., the USB port of the computer system 1000); other network interface adapters are usually integrated into the core of the computer system 1000 by attaching to the system bus as described below (e.g., the Ethernet interface of a PC computer system or the cellular network interface of a smart phone computer system). Using any of these networks, the computer system 1000 can communicate with other entities. Such communication can be one-way only receiving (e.g., broadcast TV), one-way only sending (e.g., CANbus to certain CANbus devices), or two-way, such as to other computer systems using local or wide area networks. Such communication can include communication to a cloud computing environment 1055. Certain protocols and protocol stacks can be used on each of the networks and network interfaces described above.

[0144] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interface 1054 can be attached to the core 1040 of the computer system 1000.

[0145] The core 1040 can include one or more central processing units (CPUs) 1041, graphics processing units (GPUs) 1042, dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) 1043, hardware accelerators 1044 for certain tasks, etc. These devices, as well as read-only memory (ROM) 1045, random access memory 1046, internal mass storage such as internal non-user-accessible hard disk drives, SSDs, etc. 1047, can be connected via a system bus 1048. In some computer systems, the system bus 1048 can be accessible in the form of one or more physical plugs to enable the expansion of additional CPUs, GPUs, etc. Peripheral devices can be directly attached to the system bus 1048 of the core or attached to the system bus 1048 of the core via a peripheral bus 1049. The architecture of the peripheral bus includes PCI, USB, etc. A graphics adapter 1050 can be included in the core 1040.

[0146] The CPU 1041, GPU 1042, FPGA 1043, and accelerator 1044 can execute certain instructions that, when combined, can constitute the aforementioned computer code. The computer code can be stored in the ROM 1045 or RAM 1046. Transitional data can also be stored in the RAM 1046, while permanent data can be stored, for example, in the internal mass storage 1047. Fast storage and retrieval of any memory device can be achieved by using a cache memory that can be closely associated with one or more of the CPU 1041, GPU 1042, mass storage 1047, ROM 1045, RAM 1046, etc.

[0147] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be those specially designed and constructed for the purposes of this disclosure, or they can be of the type well-known and available to those skilled in the art of computer software.

[0148] By way of example and not limitation, a computer system having the architecture of computer system 1000, and in particular the core 1040, can provide functions as a result of a processor (including CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with the user-accessible mass storage as described above and certain storage of a non-transitory nature of the core 1040, such as the core internal mass storage 1047 or ROM 1045. The software implementing the various embodiments of this disclosure can be stored in such devices and executed by the core 1040. Depending on specific needs, the computer-readable media can include one or more memory devices or chips. The software can cause the core 1040 and in particular the processors therein (including CPU, GPU, FPGA, etc.) to execute the specific processes or specific portions of the specific processes described herein, including defining data structures stored in the RAM 1046 and modifying such data structures in accordance with processes defined by the software. Additionally or alternatively, the computer system can provide functions as a result of logic hardwired or otherwise embodied in a circuit (e.g., accelerator 1044) that can operate in place of or in conjunction with the software to execute the specific processes or specific portions of the specific processes described herein. In appropriate cases, references to software can encompass logic and vice versa. In appropriate cases, references to computer-readable media can encompass a circuit (e.g., an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or both. This disclosure encompasses any suitable combination of hardware and software.

[0149] Although the present disclosure has described a number of non-limiting embodiments, there are changes, permutations and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.

Claims

1. A method for decoding video data, the method being performed by at least one processor, the method comprising: receiving a media stream including data in a haptic exchange format; obtaining a syntax of an ISOBMFF haptic sample from the data in the haptic exchange format of the media stream, wherein the syntax indicates a variable value of a data_packet_type of the media stream; as well as The media stream is decoded based on the syntax.

2. The method according to claim 1, wherein: The syntax is represented in binary form by four bits of the media stream.

3. The method according to claim 2, wherein: The first bit of the four bits indicates whether the sample in the media stream is a silence sample.

4. The method according to claim 3, wherein: The second bit of the four bits indicates whether the sample includes one or more temporal effects data packets.

5. The method according to claim 4, wherein: Decoding the media stream includes interpreting in binary form: The first bit is 0, indicating that the sample is silent, and The first bit is 1 and the second bit is any one of 0 and 1, indicating that the sample includes the one or more temporal effects data packets.

6. The method according to claim 4, wherein: The third bit of the four bits indicates whether the sample includes one or more spatial effects data packets.

7. The method according to claim 6, wherein: Decoding the media stream further includes interpreting in binary form: The first bit is 0, indicating that the sample is silent, and The first bit is 1 and the third bit is any one of 0 and 1, indicating that the sample includes the one or more spatial effects data packets.

8. The method according to claim 7, wherein: Decoding the media stream further includes interpreting in binary form: The first bit is 1 and the third bit is any one of 0 and 1, indicating that the sample includes the one or more spatial effect data packets when the second bit is 0 and when the second bit is 1.

9. The method according to claim 8, wherein: The representation of the syntax in binary form consists of the four bits.

10. The method according to claim 2, wherein: The representation of the syntax in binary form consists of the four bits.

11. An apparatus for decoding video data, the apparatus comprising: at least one memory configured to store program code; and At least one processor is configured to read the program code and perform operations according to instructions of the program code, wherein the program code includes: receiving code configured to cause the at least one processor to receive a media stream including data in a haptic exchange format; obtaining code configured to cause the at least one processor to obtain a syntax of an ISOBMFF haptic sample from the data in the haptic exchange format of the media stream, wherein the syntax indicates a variable value of a data_packet_type of the media stream; and A decoding code is configured to cause the at least one processor to decode the media stream based on the grammar.

12. The device according to claim 11, wherein: The syntax is represented in binary form by four bits of the media stream.

13. The device according to claim 12, wherein: The first bit of the four bits indicates whether the sample of the media stream is a silence sample.

14. The device according to claim 13, wherein: The second bit of the four bits indicates whether the sample includes one or more temporal effects data packets.

15. The device according to claim 14, wherein: Decoding the media stream includes interpreting in binary form: The first bit is 0, indicating that the sample is silent, and The first bit is 1 and the second bit is any one of 0 and 1, indicating that the sample includes the one or more temporal effects data packets.

16. The device according to claim 14, wherein: The third bit of the four bits indicates whether the sample includes one or more spatial effects data packets.

17. The device according to claim 16, wherein: Decoding the media stream further includes interpreting in binary form: The first bit is 0, indicating that the sample is silent, and The first bit is 1 and the third bit is any one of 0 and 1, indicating that the sample includes the one or more spatial effects data packets.

18. The device according to claim 17, wherein: Decoding the media stream further includes interpreting in binary form: The first bit is 1 and the third bit is any one of 0 and 1, indicating that the sample includes the one or more spatial effect data packets when the second bit is 0 and when the second bit is 1.

19. The device according to claim 18, wherein: The representation of the syntax in binary form consists of the four bits.

20. A non-transitory computer-readable medium storing instructions, the instructions comprising: One or more instructions that, when executed by one or more processors of a device for decoding video data, cause the one or more processors to: receiving a media stream including data in a haptic exchange format; obtaining a syntax of an ISOBMFF haptic sample from the data in the haptic exchange format of the media stream, wherein the syntax indicates a variable value of a data_packet_type of the media stream; and The media stream is decoded based on the syntax.