Method for carrying binary wavelet stream in haptic switched format
By designing a device configured to decode data in haptic exchange format, the problem of unclear tactile track timing model in multimedia presentation is solved, and the stream of binary wavelet encoding is supported, efficient tactile signal encoding and decoding is achieved.
Patent Information
- Application Number
- CN202480004208.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-16
- Filing Date
- 2024-04-17
- Publication Date
- 2025-05-23
AI Technical Summary
In multimedia presentation, the carrying timing model of the tactile track is unclear, especially the problem of how the timing of the ISOBMFF track is related to the timing of the tactile basic signal. In addition, the current JSON format is used to carry quantized wavelet coefficients rather than streams that carry binary wavelet codes.
By designing an apparatus and method, the apparatus comprises at least one memory and at least one processor configured to receive a media stream including data in a haptic exchange format, obtain wavelet effect values, and decode the media stream based on these values. The tactile exchange format can be in .hjif format, and the wavelet effect value can be obtained from the "band_type" attribute of the data in .hjif format, indicating that the frequency band of the media stream is in a binary encoded wavelet stream, and entropy decoding and inverse wavelet transformation are required for decoding.
It realizes effective encoding and decoding of tactile signals in multimedia presentation, solves the problem of unclear timing model, and supports binary wavelet encoding streams, improving the quality and efficiency of tactile experience.
Smart Images

Figure CN120035994A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Application No. 63 / 459,929, filed on April 17, 2023, and U.S. Application No. 18 / 636,856, filed on April 16, 2024, the disclosures of which are incorporated herein by reference in their entirety. Technical Field
[0003] The present disclosure relates to a set of advanced video coding and decoding techniques. More specifically, the present disclosure relates to encoding and decoding haptic experiences for multimedia presentations, and methods for carrying binary wavelet streams in a haptic exchange format. Background Art
[0004] Haptic experience has become part of multimedia presentations. In applications where multimedia presentations include aspects of haptic experience, haptic signals may be delivered to a device or wearable device, and the user may experience haptic sensations in coordination with the visual and / or audio media experience during use of the application.
[0005] Recognizing the increasing popularity of tactile experience in multimedia presentations, the motion picture experts group (MPEG) has begun studying compression standards for tactile (for both MPEG-DASH and MPEG-I) as well as carrying compressed tactile signaling in the ISO based media file format (ISOBMFF).
[0006] One of the problems to be solved in aspects related to haptic experience in multimedia presentations is that the timing model carried by the haptic track is unclear, ie, it is unclear how the timing of the ISOBMFF track relates to the timing of the haptic base signal. A solution to this problem is needed.
[0007] The Haptic Committee draft includes both a JSON and a binary format. The current JSON format (called the Haptic Interchange Format) carries quantized wavelet coefficients rather than a stream of binary wavelet encodings. A solution to this problem is needed. Summary of the invention
[0008] According to one aspect of the present disclosure, there is an apparatus, and similarly, a method and a computer-readable medium, the apparatus comprising at least one memory configured to store computer program code; and at least one processor, the at least one processor being configured to access the computer program code and to operate as instructed by the computer program code, the computer program code comprising: a receiving code configured to cause the at least one processor to receive a media stream comprising data in a tactile exchange format; an obtaining code configured to cause the at least one processor to obtain a wavelet effect value from the data in the tactile exchange format of the media stream; and a decoding code configured to cause the at least one processor to decode the media stream based on the wavelet effect value.
[0009] The haptic exchange format may be in .hjif format, and the data may be in .hjif format.
[0010] The wavelet effect value can be obtained from the "band_type" attribute of the .hjif format data.
[0011] The wavelet effect value may be indicated as "BinaryWavelet" in the "band_type" attribute of data in the .hjif format.
[0012] The wavelet effect value may indicate that the frequency band of the media stream is in a binary coded wavelet stream and that entropy decoding along with an inverse wavelet transform is required to decode the waves of the frequency band of the media stream.
[0013] The wavelet effect value may indicate a frequency band of the media stream that is in a binary coded wavelet stream.
[0014] The wavelet effect value may indicate a frequency band wave that requires entropy decoding along with an inverse wavelet transform to decode the media stream.
[0015] Decoding the media stream may include decoding binary wavelet keyframes in .hjif format by running a base64 decode of the media stream.
[0016] Decoding the media stream may include running "readWaveletEffect()" on the haptic data of the media stream.
[0017] Additional embodiments will be set forth in the description which follows, and in part will be apparent from the description, and / or may be learned by practice of the embodiments presented in the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Other features, properties and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:
[0019] Figure 1is a schematic diagram of a simplified block diagram of a communication system according to an embodiment of the present disclosure;
[0020] Figure 2 is a schematic diagram of a simplified block diagram of a streaming system according to an embodiment of the present disclosure;
[0021] Figure 3 A is a schematic diagram of a simplified block diagram of a tactile encoder according to an embodiment of the present disclosure;
[0022] Figure 3 B is a schematic diagram of a simplified block diagram of a haptic decoder and a haptic renderer according to an embodiment of the present disclosure;
[0023] Figure 4 is a schematic diagram of a process for determining relative timing of MIHS units according to an embodiment of the present disclosure;
[0024] 5 is a schematic diagram of a process for determining relative timing of MIHS units according to an embodiment of the present disclosure;
[0025] Figure 6 is an exemplary flow chart illustrating a process for processing tactile media according to an embodiment of the present disclosure;
[0026] Figure 7 is an exemplary diagram illustrating some aspects of an embodiment according to the present disclosure;
[0027] Figure 8 is an exemplary diagram illustrating some aspects of an embodiment according to the present disclosure;
[0028] Fig. 9 are exemplary diagrams illustrating some aspects of embodiments according to the present disclosure; and
[0029] Fig.10 is an exemplary diagram illustrating some aspects of an embodiment according to the present disclosure. DETAILED DESCRIPTION
[0030] According to one aspect of the present disclosure, a method, system and non-transitory storage medium for parallel processing of dynamic mesh compression are provided. Embodiments of the present disclosure may also be applied to static meshes.
[0031] refer to Figure 1 and Figure 2 , describes an embodiment of the present disclosure for implementing the encoding and decoding structures of the present disclosure.
[0032] Figure 1A simplified block diagram of a communication system 100 according to an embodiment of the present disclosure is shown. The system 100 may include at least two terminals 110, 120 interconnected via a network 150. For unidirectional transmission of data, a first terminal 110 may encode video data (which may include mesh data) at a local location for transmission to another terminal 120 via the network 150. The second terminal 120 may receive the encoded video data of the other terminal from the network 150, decode the encoded data, and display the recovered video data. Unidirectional data transmission may be common in media service applications and the like.
[0033] Figure 1 A second pair of terminals 130, 140 is shown, which are provided to support bidirectional transmission of encoded video that may occur, for example, during a video conference. For bidirectional transmission of data, each terminal 130, 140 may encode video data captured at a local location for transmission to the other terminal via the network 150. Each terminal 130, 140 may also receive encoded video data sent by the other terminal, may decode the encoded data, and may display the recovered video data at a local display device.
[0034] exist Figure 1 In the embodiment of the present invention, the terminals 110-140 can be, for example, servers, personal computers, and smart phones and / or any other type of terminals. For example, the terminals (110-140) can be laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network 150 represents any number of networks that transmit encoded video data between the terminals 110-140, including, for example, wired and / or wireless communication networks. The communication network 150 can exchange data in circuit switching and / or packet switching channels. Representative networks include telecommunication networks, local area networks, wide area networks, and / or the Internet. For the purpose of this discussion, unless explained below, the architecture and topology of the network 150 may be unimportant to the operation of the present disclosure.
[0035] Figure 2 The placement of a video encoder and decoder in a streaming environment as an example of an application for the disclosed subject matter is shown. The disclosed subject matter can be used with other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc., and the like.
[0036] like Figure 2 As shown, the streaming system 200 may include a capture subsystem 213, which includes a video source 201 and an encoder 203. The streaming system 200 may also include at least one streaming server 205 and / or at least one streaming client 206.
[0037] Video source 201 can create a stream 202 that includes, for example, a 3D mesh and metadata associated with the 3D mesh. Video source 201 can include, for example, a 3D sensor (e.g., a depth sensor) or a 3D imaging technology (e.g., a digital camera) and a computing device configured to generate a 3D mesh using data received from the 3D sensor or the 3D imaging technology. Sample stream 202, which can have a high amount of data when compared to an encoded video stream, can be processed by an encoder 203 coupled to video source 201. Encoder 203 can include hardware, software, or a combination thereof to implement or realize various aspects of the disclosed subject matter, as described in more detail below. Encoder 203 can also generate an encoded video stream 204. The encoded video stream 204, which can have a lower amount of data when compared to the uncompressed stream 202, can be stored on a streaming server 205 for future use. One or more streaming clients 206 can access streaming server 205 to retrieve a video stream 209, which can be a copy of the encoded video stream 204.
[0038] The streaming client 206 may include a video decoder 210 and a display 212. The video decoder 210 may, for example, decode a video stream 209 (which is an input copy of an encoded video stream 204) and create an output video sample stream 211 that may be rendered on the display 212 or another rendering device (not depicted). In some streaming systems, the video streams 204, 209 may be encoded according to some video encoding / compression standard.
[0039] Figure 3 may be a functional block diagram of a video decoder 300 according to an embodiment of the present invention.
[0040] Receiver 302 may receive one or more codec video sequences to be decoded by decoder 300; in the same or another embodiment, one coded video sequence at a time, wherein the decoding of each coded video sequence is independent of the other coded video sequences. The coded video sequence may be received from channel 301, which may be a hardware / software link to a storage device storing the coded video data. Receiver 302 may receive the coded video data and other data that may be forwarded to its corresponding use entity (not depicted), such as coded audio data and / or auxiliary data streams. Receiver 302 may separate the coded video sequence from the other data. To combat network jitter, a buffer memory 303 may be coupled between receiver 302 and entropy decoder / parser 304 (hereinafter referred to as "parser"). Buffer 303 may not be needed or may be small when receiver 302 receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network. For use over a best-effort packet network such as the Internet, buffer 303 may be necessary, may be relatively large, and may advantageously be of an adaptive size.
[0041] The video decoder 300 may include a parser 304 to reconstruct symbols 313 from an entropy coded video sequence. The categories of these symbols include information for managing the operation of the decoder 300, and potential information for controlling a display device (e.g., display 312), which is not part of the decoder but may be coupled to the decoder. The control information for the display device may be a parameter set fragment (not shown) of Supplemental Enhancement Information (SEI message) or video usability information. The parser 304 may parse / entropy decode the received coded video sequence. The encoding of the coded video sequence may be performed according to a video coding technique or standard, and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 304 may extract a subgroup parameter set for at least one of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. Subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CU), blocks, Transform Units (TU), Prediction Units (PU), etc. The entropy encoder / parser may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.
[0042] The parser 304 may perform an entropy decoding / parsing operation on the video sequence received from the buffer memory 303, thereby creating the symbol 313. The parser 304 may receive the encoded data and selectively decode a specific symbol 313. In addition, the parser 304 may determine whether to provide the specific symbol 313 to the motion compensation prediction unit 306, the scaler / inverse transform unit 305, the intra prediction unit 307, or the loop filter 311.
[0043] Depending on the type of the coded video picture or a portion of the coded video picture (e.g., inter- and intra-pictures, inter- and intra-blocks), and other factors, the reconstruction of the symbol 313 may involve a plurality of different units. Which units are involved and how they are involved may be controlled by subgroup control information parsed from the coded video sequence by the parser 304. For the sake of brevity, such subgroup control information flow between the parser 304 and the plurality of units below is not described.
[0044] In addition to the functional blocks already mentioned, decoder 300 may be conceptually subdivided into several functional units as described below. In a practical embodiment operating under commercial constraints, many of these units closely interact with each other and may be integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the functional units below.
[0045] The first unit is the sealer / inverse transform unit 305. The sealer / inverse transform unit 305 receives the quantized transform coefficients as symbols 313 from the parser, as well as control information including which transform method to use, block size, quantization factor, quantization scaling matrix, etc. The sealer / inverse transform unit may output a block including sample values, which may be input into the aggregator 310.
[0046] In some cases, the output samples of the sealer / inverse transform unit 305 may belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the intra-picture prediction unit 307. In some cases, the intra-picture prediction unit 307 uses reconstructed information extracted from the current partial reconstructed picture 309 to generate surrounding blocks of the same size and shape as the block being reconstructed. In some cases, the aggregator 310 adds the prediction information generated by the intra-prediction unit 307 to the output sample information provided by the sealer / inverse transform unit 305 on a per-sample basis.
[0047] In other cases, the output samples of the sealer / inverse transform unit 305 may belong to an inter-frame coded and potentially motion compensated block. In this case, the motion compensated prediction unit 306 may access the reference picture memory 308 to extract samples for prediction. After the extracted samples are motion compensated according to the symbols 313, these samples may be added to the output of the sealer / inverse transform unit (in this case referred to as residual samples or residual signals) by the aggregator 310 to generate output sample information. The acquisition of prediction samples by the motion compensation unit from the address in the reference picture memory may be controlled by a motion vector, and the motion vector is provided to the motion compensation unit in the form of the symbols, for example, including X, Y and reference picture components. Motion compensation may also include interpolation of sample values extracted from the reference picture memory when using sub-sample accurate motion vectors, motion vector prediction mechanisms, and the like.
[0048] The output samples of aggregator 310 may be employed by various loop filtering techniques in loop filter unit 311. The video compression techniques may include in-loop filter techniques controlled by parameters included in the encoded bitstream and available to loop filter unit 311 as symbols 313 from parser 304. However, in other embodiments, the video compression techniques may also be responsive to meta-information obtained during decoding of a previous (in decoding order) portion of an encoded picture or encoded video sequence, as well as to previously reconstructed and loop filtered sample values.
[0049] The output of the loop filter unit 311 may be a sample stream, which may be output to the display device 312 and stored in the reference picture memory 557 for subsequent inter-picture prediction.
[0050] Once fully reconstructed, certain coded pictures may be used as reference pictures for future prediction. Once a coded picture is fully reconstructed, and the coded picture is identified as a reference picture (e.g., by parser 304), current reference picture 309 may become part of reference picture buffer 308, and new current picture memory may be reallocated before starting reconstruction of a subsequent coded picture.
[0051] The video decoder 300 may perform decoding operations according to a predetermined video compression technique that may be recorded, for example, in the ITU-T H.265 standard. As specified in a video compression technology document or standard (particularly in the configuration file of the present application), the coded video sequence may conform to the syntax specified by the video compression technology or standard used in the sense that the coded video sequence follows the syntax of the video compression technology or standard. For compliance, it is also required that the complexity of the coded video sequence is within the range defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the coded video sequence.
[0052] In an embodiment, the receiver 302 may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by the video decoder 300 to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0053] Figure 4 It may be a functional block diagram of a video encoder 400 according to an embodiment disclosed in the present application.
[0054] The encoder 400 may receive video samples from a video source 401 (not a part of the encoder), which may capture video images to be encoded by the encoder 400 .
[0055] The video source 401 may provide a source video sequence in the form of a digital video sample stream to be encoded by the encoder 303, wherein the digital video sample stream may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source 401 may be a storage device storing previously prepared videos. In a video conferencing system, the video source 401 may be a camera that collects local image information as a video sequence. The video data may be provided as a plurality of separate pictures that are given motion when viewed sequentially. The picture itself may be constructed as a spatial pixel array, wherein each pixel may include one or more samples depending on the sampling structure, color space, etc. used. The relationship between pixels and samples may be easily understood by those skilled in the art. The following focuses on describing samples.
[0056] According to an embodiment, the encoder 400 may encode and compress pictures of a source video sequence into an encoded video sequence 410 in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller 402. The controller controls other functional units as described below and is functionally coupled to these units. For the sake of brevity, couplings are not shown in the figure. The parameters set by the controller may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization technology, etc.), picture size, picture group (group of pictures, GOP) layout, maximum motion vector search range, etc. When other suitable functions are related to the video encoder 400 optimized for a certain system design, those skilled in the art can easily identify these other suitable functions of the controller 402.
[0057] Some video encoders operate in what is readily recognizable to those skilled in the art as a "coding loop". As a simplified description, the coding loop may include the encoding portion of an encoder 402 (hereinafter referred to as a "source encoder") (responsible for creating symbols based on the input picture to be encoded and the reference picture) and a (local) decoder 406 embedded in the encoder 400, which reconstructs the symbols to create sample data in the same way as the (remote) decoder created the sample data (because in the video compression techniques considered in this application, any compression between the symbols and the encoded video code stream is lossless). The reconstructed sample stream is input to a reference picture memory 405. Since the decoding of the symbol stream produces bit-accurate results that are independent of the decoder location (local or remote), the reference picture buffer contents are also bit-accurately corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction portion of the encoder are exactly the same sample values that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, for example due to channel errors) is well known to those skilled in the art.
[0058] The operation of the "local" decoder 406 can be combined with the above Figure 3 The same is true of the "remote" decoder 300 described in detail. However, additional brief reference is made to Figure 4 , when symbols are available and the entropy encoder 408 and the parser 304 are capable of losslessly encoding / decoding the symbols into an encoded video sequence, the entropy decoding portion of the decoder 300 (including the channel 301, the receiver 302, the buffer 303 and the parser 304) may not be fully implemented in the local decoder 406.
[0059] At this point it can be observed that any decoder techniques other than parsing / entropy decoding present in the decoder must also be present in the corresponding encoder in substantially the same functional form. The description of encoder techniques can be simplified because encoder techniques are reciprocal to the decoder techniques described comprehensively. A more detailed description is only required in certain areas and is provided below.
[0060] As part of its operation, source encoder 403 may perform motion compensated predictive coding. Motion compensated predictive coding predictively encodes an input frame with reference to one or more previously encoded frames from a video sequence designated as "reference frames." In this manner, encoding engine 407 encodes the differences between pixel blocks of an input frame and pixel blocks of a reference frame that may be selected as a prediction reference for the input frame.
[0061] The local video decoder 406 may decode the encoded video data of the frame that may be designated as the reference frame based on the symbol created by the source encoder 403. The operation of the encoding engine 407 may be a lossy process. When the encoded video data is available at the video decoder ( Figure 4 When decoded at a remote video decoder (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder 406 replicates the decoding process that may be performed by the video decoder on the reference frame and may cause the reconstructed reference frame to be stored in the reference picture cache 405. In this way, the encoder 400 may locally store a copy of the reconstructed reference frame that has common content (absent transmission errors) with the reconstructed reference frame to be obtained by the remote video decoder.
[0062] The predictor 404 may perform a prediction search for the encoding engine 407. That is, for a new frame to be encoded, the predictor 404 may search the reference picture memory 405 for sample data (as a candidate reference pixel block) or certain metadata, such as a reference picture motion vector, block shape, etc., that may serve as a suitable prediction reference for the new picture. The predictor 404 may operate pixel-by-pixel based on a sample block to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor 404, it may be determined that the input picture may have a prediction reference taken from a plurality of reference pictures stored in the reference picture memory 405.
[0063] The controller 402 may manage encoding operations of the video encoder 403 , including, for example, setting parameters and subgroup parameters for encoding video data.
[0064] The outputs of all the above functional units may be entropy encoded in the entropy encoder 408. The entropy encoder performs lossless compression on the symbols generated by the various functional units according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into a coded video sequence.
[0065] Transmitter 409 may buffer the encoded video sequence created by entropy encoder 408 in preparation for transmission over communication channel 411, which may be a hardware / software link to a storage device where the encoded video data will be stored. Transmitter 409 may combine the encoded video data from video encoder 403 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0066] The controller 402 may manage the operation of the encoder 400. During encoding, the controller 405 may assign a certain coded picture type to each coded picture, but this may affect the coding techniques that may be applied to the corresponding picture. For example, a picture may generally be assigned to any of the following frame types:
[0067] An intra picture (I picture) may be a picture that can be encoded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh pictures. Those skilled in the art are aware of the variations of I pictures and their corresponding applications and features.
[0068] A predictive picture (P picture) may be a picture that can be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict sample values for each block.
[0069] Bidirectional predictive pictures (B pictures), which can be pictures that can be encoded and decoded using intra prediction or inter prediction, which uses up to two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstruction of a single block.
[0070] The source picture may typically be spatially subdivided into blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, which are determined according to the coding allocation of the corresponding picture applied to the 'block'. For example, blocks of an I picture may be non-predictively coded, or the blocks may be predictively coded (spatial prediction or intra-frame prediction) with reference to already coded blocks of the same picture. Blocks of pixels of a P picture may be non-predictively coded by spatial prediction with reference to one previously coded reference picture or by temporal prediction. Blocks of a B picture may be non-predictively coded by spatial prediction with reference to one or two previously coded reference pictures or by temporal prediction.
[0071] The video encoder 400 may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder 400 may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.
[0072] In an embodiment, the transmitter 409 may transmit additional data when transmitting the encoded video. The source encoder 403 may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, auxiliary enhancement information (SEI) messages, video usability information (VUI) parameter set fragments, etc.
[0073] refer to FIG. 5A to FIG. 5B, describes an embodiment of the present disclosure for implementing a haptic encoder 500 and a haptic decoder 550 .
[0074] like Figure 5A As shown, the haptic encoder 500 can receive descriptive haptic data and waveform haptic data. Therefore, the haptic encoder 500 may be able to process three types of input files: .ohm metadata files (object haptic metadata-text file format of haptic metadata), descriptive haptic files (.ivs, .ahap and .hjif) or waveform PCM files (.wav). Examples of descriptive data may include: .ahap (Apple Haptic and AudioPattern-JSON-like file format for specifying haptic patterns) from Apple (representing the expected haptic output through a set of modulated continuous signals and a set of parameterized modulation transients); .ivs from Immersion (representing the expected haptic output through a set of basic effects parameterized by a set of parameters); or .hjif (haptic JSON interchange format), that is, the MPEG format proposed in this application. An example of a waveform pulse code modulation (PCM) signal may include an .ohm input file, which includes metadata information.
[0075] According to an embodiment, the haptic encoder 500 may handle the two types of input files differently. For descriptive content, the haptic encoder 500 may semantically analyze the input to transcode the data (if necessary) into the encoded representation proposed by the application.
[0076] According to an embodiment, the .ohm metadata input file may include a description of the haptic system and settings. In particular, the .ohm metadata input file may include the name of each associated haptic file (descriptive or PCM), and a description of the signal. The .ohm metadata input file also provides a mapping between each channel of the signal and a target body part on the user's body. For the .ohm metadata input file, the haptic encoder performs metadata extraction by retrieving the associated haptic file from the URI, and encoding the metadata based on the type of metadata, and by extracting the metadata from the .ohm file, mapping it to the metadata information of the data model.
[0077] According to an embodiment, descriptive tactile files (e.g., .ivs, .ahap, and .hjif) can be encoded by a simple process. The tactile encoder 500 first specifically identifies the input format. If the input format is a .hjif file, then transcoding is not required and the file can be further edited, compressed into a binary format, and ultimately packaged into a MIHS stream. If a .ahap or .ivs input file is used, transcoding is necessary. The tactile encoder 500 first semantically analyzes the input file information and transcodes it to format it into a selected data model. After transcoding, the data can be exported as a .hjif file, a .hmpg binary file, or a MIHS stream.
[0078] According to an embodiment, the tactile encoder 500 can perform signal analysis to parse the signal structure of the .wav file and convert it into the coded representation proposed in the present application. For waveform PCM content, the signal analysis process can be divided into two sub-processes performed by the tactile encoder 500. After performing frequency band decomposition on the signal, in the first sub-process, the low frequency can be encoded using a key frame extraction process. The low frequency band can then be reconstructed, and the error between the signal and the original low frequency signal can be calculated. The residual signal can then be added to the original high frequency band before encoding using a wavelet transform, and encoding using a wavelet transform is a second sub-process. According to an embodiment, when multiple low frequency bands are used, the residuals from all low frequency bands are added to the high frequency bands before encoding. In an embodiment using multiple high frequency bands, the residuals from the low frequency bands are added to the first high frequency band before encoding.
[0079] According to an embodiment, key frame extraction includes obtaining lower frequency bands from frequency band decomposition and analyzing their contents in the time domain. According to an embodiment, wavelet processing may include obtaining high frequency bands from frequency band decomposition and low frequency residuals and dividing them into blocks of equal size. These equal-sized signal blocks are then analyzed in a psychohaptic model. Lossy compression may be applied by wavelet transforming the blocks and quantizing them with the help of a psychohaptic model. Finally, each block is saved into a separate effect in a single frequency band, which is done in formatting. Binary compression may apply lossless compression using appropriate coding techniques (e.g., a multi-level tree set splitting (SPIHT, Set partitioning in hierarchicaltrees) algorithm and arithmetic coding (AC)).
[0080] like Figure 5AAs shown, the tactile encoder 500 can be configured to encode descriptive and quantized tactile data, and can output three types of formats - an interchange format (.hjif), a binary compressed format (.hmpg), and a streaming format (e.g., MPEG immersive haptic stream (MIHS)). The .hjif format is a human-readable format based on JSON that can be easily parsed and manually edited, making it an ideal interchange format, especially when designing / creating content. For distribution purposes, the .hjif data can be compressed into a more memory-efficient binary .hmpg code stream. This compression can be lossy, and different parameters affect the encoding depth of the amplitude and frequency that make up the code stream. For streaming purposes, the data can be compressed and packaged into an MPEG-1 tactile stream (MIHS). The above three formats have complementary purposes, and lossy one-to-one conversions can be performed between them.
[0081] like Figure 5B As shown, the tactile decoder 550 can take a .hmpg compressed binary file format or a MIHS code stream as input. The tactile decoder 550 can output a .hjif interchange format that can be used directly for rendering. These two input formats can undergo binary decompression to extract metadata and the data itself from the file and map it to the selected data structure. The data can then be exported to the tactile renderer 580 in .hjif format.
[0082] like Figure 5B As shown, the renderer 580 includes a synthesizer. The synthesizer can render tactile data from a .hjif input file into a PCM output file. Rendering and / or synthesis is informative. According to an embodiment, the synthesizer parses the input file and performs high-level synthesis distribution between vectors, wavelets, etc. The synthesis process then descends to the frequency band components of the codec in which the synthesis process is called. All frequency bands of a given channel are then mixed by a simple addition operator to recreate the desired tactile signal.
[0083] According to an embodiment, the haptic experience definition is the root of the hierarchical data model. It provides information about the file date and format version, it describes the haptic experience, it lists the different avatars (i.e., body representations) used throughout the experience, and it defines all haptic perceptions.
[0084] According to an embodiment, a self-contained stream format for transmitting MPEG-I tactile data may use a packetization method and may include two levels of packetization: an MPEG-I tactile stream (MIHS) unit, which covers a duration and includes zero or more MIHS packets; and a MIHS packet including metadata or tactile effect data. In an embodiment, a MIHS unit may be referred to as a network abstraction layer unit associated with tactile data. In an embodiment, a MIHS unit may be referred to as a MIHS sample associated with tactile data.
[0085] According to an embodiment, a MIHS unit may be a synchronous unit or an asynchronous unit. A synchronous unit resets the previous effect, thus providing a tactile experience independent of the previous MIHS unit. An asynchronous unit is a continuation of the previous MIHS unit and cannot be independently decoded and rendered without decoding the previous MIHS unit.
[0086] According to an embodiment, a tactile signal may be encoded on multiple channels. In some embodiments, a tactile channel may define a signal to be rendered at a specific body position using a dedicated actuator / device. Metadata stored at the channel level may include information such as gain associated with the channel, mixing weights, desired body position of tactile feedback, and optionally reference to a device and / or direction. Additional information such as a desired sampling frequency or sampling count may also be provided. Finally, the tactile data of the channel is contained in a set of tactile frequency bands defined by its frequency range. The tactile frequency band describes the tactile signal of the channel within a given frequency range. The frequency band is defined by a list of types and sequences of tactile effects, each tactile effect including a set of key frames. For each type of tactile frequency band, the tactile effect may be defined at least by position and type. The position may indicate the temporal or spatial position of the effect. In some embodiments, a value of 0 is a relative starting position of an experience that depends on the dependent variable of the configured perceptual modality. The default unit for temporal tactile feedback may be milliseconds, while the default unit for spatial tactile feedback may be millimeters. This embodiment discloses a "starting location of the experience" because the binary distribution format does not have any concept of finite time intervals (ie, frames or samples).
[0087] Depending on the type of band and the type of effect, additional properties can be specified, including phase, base signal, composition, and multiple consecutive haptic keyframes that describe the effect.
[0088] According to an embodiment, a haptic data hierarchy is defined in the present disclosure.
[0089] ●Tactile channel
[0090] oTactile bands
[0091] ■Tactile effects
[0092] Embodiments of the present disclosure describe two anchor points for the location of a haptic effect relative to an ISOBMFF track.
[0093] Figure 6 A first embodiment 600 is shown in FIG. Figure 6 As shown, each MIHS unit (also referred to as MIHS sample, ISOBMFF tactile sample or sample in the embodiment) includes one or more tactile channel information and one or more tactile frequency band information. As described above, each MIHS unit is composed of one or more channels, and each channel is composed of one or more frequency bands. Then each frequency band can have one or more effects.
[0094] In a first embodiment, the temporal position of an effect can be defined as an offset relative to the start timing (e.g., MIHS unit start time) of a sample carrying the effect. In a second or the same embodiment, the offset is based on the start time and / or presentation time of a media or haptic track.
[0095] According to an embodiment, the first embodiment enables manipulation of the track without affecting the position of the haptic effect, since any changes in the ISOBMFF sampling timing do not affect the relative position of the effect. According to an embodiment, in the case of a basic haptic stream (e.g., a high-level syntax stream), the second embodiment may be used when the basic stream is used without ISOBMFF.
[0096] According to an embodiment, multiple types of haptic tracks may be used. In an embodiment, a sample or MIHS cell may be used in the haptic track, the temporal position of the effect having the sample or MIHS cell being defined as an offset relative to the start timing of the sample. According to another embodiment, for example Figure 7 In example 700, a sample or MIHS unit is used in a haptic track, the effect of which has a time position relative to the start time of the track. In another embodiment, a mixture of MIHS units or samples may be used.
[0097] Embodiments of the present disclosure provide a timing model that can be used to synchronize haptic effects with other media tracks in the same or related ISOBMFF files. Where the timing model of a haptic track is relative to the timing model of a related ISOBMFF file, manipulation and processing of the media track becomes more efficient.
[0098] like Figure 8 As shown, process 800 illustrates an exemplary process for decoding haptic data.
[0099] At operation 805 , a media stream including one or more haptic tracks and one or more video tracks may be received.
[0100] At operation 810, one or more Moving Picture Experts Group (MPEG) Immersive Haptic Stream (MIHS) units may be obtained from a media stream. In some embodiments, a MIHS unit may include one or more haptic effects. A MIHS unit may also include a start time of the MIHS unit.
[0101] In an embodiment, the MIHS unit is associated with at least one haptic channel, the at least one haptic channel includes one or more haptic frequency bands, and each of the one or more haptic frequency bands has at least one haptic effect.
[0102] At operation 815, timing information associated with the one or more haptic effects may be obtained. In an embodiment, the timing information may include at least one temporal position of the one or more haptic effects.
[0103] In an embodiment, the temporal position of the haptic effect indicates an effect start time of the haptic effect, wherein the effect start time of the haptic effect is an offset based on a start time of a corresponding MIHS unit.The effect start time may indicate a start time of the haptic effect relative to a start time of the corresponding MIHS unit.
[0104] In an embodiment, the effect start time of the haptic effect is an absolute time based on a start time of the at least one haptic track or the at least one video track.
[0105] At operation 820, the media stream is rendered based on the obtained timing information.
[0106] According to an embodiment, manipulation of the order of one or more MIHS units does not affect at least one temporal position of one or more haptic effects because the one or more MIHS units correspond to one or more ISO-based media file format (ISOBMFF) samples associated with at least one video track.
[0107] In some embodiments, a synchronization MIHS unit may be obtained from a media stream. In an embodiment, a synchronization MIHS unit is a special type of MIHS unit configured to provide a reset point in a bitstream. In an embodiment, a synchronization MIHS unit is mapped to synchronization samples corresponding to one or more haptic channels in a video bitstream.
[0108] As in Fig. 9In example 900, a tactile encoder according to an embodiment of the present invention generates a compact and efficient binary distribution format (.hmpg) for distribution. A tactile decoder can decode this format and send it to a renderer. However, without the embodiments of the present invention, on the other hand, the tactile interchange format (.hjif) does not have binary compression, and therefore, the wavelet coefficients are stored in the .hjif instead of in a compressed bitstream. And in view of this, the embodiments of the present invention extend the band types in the .hjif format to include a new option: binary wavelets.
[0109] For example, see Table 1 below, which provides an example description of an MPEG_haptics.band object:
[0110] Table 1 - Description of MPEG_haptics.band object
[0111]
[0112]
[0113]
[0114] In an extended solution of such an embodiment herein, band_type is "BinaryWavelet" indicating that the band is in a binary-coded wavelet stream, and entropy decoding together with inverse wavelet transform is required to decode the wave, as shown in Table 1.
[0115] In addition, the embodiments of this article also include extended key frame features. For example, the embodiments of this article extend the MPEG_haptics key frame to include a binary wavelet coded stream. In this case, there is a separate key frame object, which includes one item: a base64-encoded string of a binary arithmetic coded stream, as shown in Table 2 below:
[0116] Table 2 - Description of the MPEG_haptics.effect object when used with band_type="BinaryWavelet"
[0117]
[0118] In addition, the embodiments of this document also provide a decoding process, so that in order to decode the binary wavelet key frame of .hjif, the decoder runs the base64 decoding of based64_arithmetic_stream and calculates its size:
[0119] arithmetic_stream=base64 decoding(based64_arithmetic_stream);
[0120] streamsize=length of(arithmetic_stream);
[0121] Then run MIHS's readWaveletEffect(), for example according to Table 3 below:
[0122] Table 3 - Syntax of readWaveletEffect()
[0123]
[0124]
[0125] As shown in Table 3, the arithmetic_stream is decoded using the SPIHT algorithm, and the BinaryWavelet type can be replaced with the WaveletWave type, as shown in Table 4 below:
[0126] Table 4 - Syntax for decoding basis effect in wavelet band (MPEG_haptics_decodeWaveletEffect())
[0127]
[0128] Therefore, according to the embodiments of the present invention, the benefits of binary wavelet mode are provided so that in the new band type described herein, the wavelet stream is binary encoded and stored in the HJIF file. This mode has the following advantages, for example: 1. It is easy to convert from HJIF to MIHS stream because the wavelet stream is already binary encoded and only requires a simple base64 conversion. 2. The MIHS stream can be easily converted to HJIF without arithmetic decoding. The MIHS stream can be converted to HJIF by running only base64 encoding. No wavelet decoder is required.
[0129] Therefore, this article provides a method for storing wavelet effects in binary in a haptic interchange format, wherein a wavelet stream is binary encoded and then converted to a string using base64 encoding and stored in a single keyframe in the effect, wherein the same stream stored in MIHS format is also stored in .hjif format, and thus conversion from MIHS to .hjif and from .hjif to MIHS does not require any arithmetic encoding / decoding, and any general purpose processor that supports base64 encoding / decoding can be used for this purpose.
[0130] Those skilled in the art will appreciate that the techniques described herein can be implemented on both the encoder side and the decoder side. The above techniques can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 7 A computer system 700 suitable for implementing certain embodiments of the present disclosure is shown.
[0131] Computer software may be encoded using any suitable machine code or computer language that may be subjected to assembly, compilation, linking or similar mechanisms to create code comprising instructions that may be executed directly by a computer central processing unit (CPU), graphics processing unit (GPU), etc., or through interpretation, microcode execution, etc.
[0132] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0133] Fig.10 The components shown for computer system 1000 are examples and are not intended to imply any limitation on the scope of use or functionality of computer software implementing embodiments of the present disclosure. Nor should the configuration of components be interpreted as having any dependency or requirement related to any one or combination of components shown in the non-limiting embodiment of computer system 1000.
[0134] Computer system 1000 may include certain human interface input devices. Such human interface input devices may be responsive to input from one or more human users, for example, through tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not depicted). Human interface devices may also be used to capture certain media that are not necessarily directly related to a person's conscious input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0135] Input human interface devices may include one or more of the following (only one of each is depicted): keyboard 1001 , mouse 1002 , trackpad 1003 , touch screen 1010 , data gloves, joystick 1005 , microphone 1006 , scanner 1007 , camera 1008 .
[0136] The computer system 1000 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users by, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback performed by touch screen 1010, data gloves, or joystick 1005, but there may also be tactile feedback devices that are not used as input devices). For example, such devices may be audio output devices (e.g., speakers 1009, headphones (not shown)), visual output devices (e.g., screen 1010, including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities-some of which may be able to output two-dimensional visual outputs or more than three-dimensional outputs by devices such as stereo output, virtual reality glasses (not shown), holographic displays, and smoke cans (not shown))), and printers (not shown).
[0137] The computer system 1000 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 1020 with CD / DVD or similar media 1021, thumb drive 1022, removable hard drive or solid state drive 1023, traditional magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD based devices such as security software dogs (not shown), and the like.
[0138] Those skilled in the art will also appreciate that the term "computer-readable media" used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transient signals.
[0139] The computer system 1000 may also include an interface to one or more communication networks. For example, the network may be wireless, wired, optical. The network may also be local, wide area, metropolitan area network, vehicle-mounted and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks (e.g., Ethernet), wireless LANs, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), TV wired or wireless wide area networks (including cable TV, satellite TV, and terrestrial broadcast TV), vehicle-mounted and industrial networks (including CANBus), etc. Certain networks typically require an external network interface adapter that attaches to certain general data ports or peripheral buses 1049 (e.g., a USB port of the computer system 1000; other network interface adapters are typically integrated into the core of the computer system 1000 by attaching to a system bus as described below (e.g., an Ethernet interface for a PC computer system or a cellular network interface for a smart phone computer system). Using any of these networks, the computer system 1000 can communicate with other entities. Such communication can be one-way receive-only (e.g., broadcast TV), one-way send-only (e.g., CANbus to certain CANbus devices), or two-way, such as to other computer systems using local or wide area networks. Such communication can include communication to a cloud computing environment 1055. Certain protocols and protocol stacks can be used on each of those networks and network interfaces as described above.
[0140] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interface 1054 may be attached to the core 1040 of the computer system 1000 .
[0141] The core 1040 may include one or more central processing units (CPUs) 1041, graphics processing units (GPUs) 1042, dedicated programmable processing units in the form of field programmable gate areas (FPGAs) 1043, hardware accelerators 1044 for certain tasks, and the like. These devices, as well as read-only memory (ROM) 1045, random access memory 1046, internal mass storage 1047 such as internal non-user accessible hard disk drives, SSDs, and the like, may be connected via a system bus 1048. In some computer systems, the system bus 1048 may be accessible in the form of one or more physical plugs to enable expansion of additional CPUs, GPUs, and the like. Peripheral devices may be attached directly to the system bus 1048 of the core, or may be attached to the system bus 1048 of the core via a peripheral bus 1049. The architecture of the peripheral bus includes PCI, USB, and the like. A graphics adapter 1050 may be included in the core 1040.
[0142] The CPU 1041, GPU 1042, FPGA 1043, and accelerator 1044 may execute certain instructions, which in combination may constitute the aforementioned computer code. The computer code may be stored in ROM 1045 or RAM 1046. Transitional data may also be stored in RAM 1046, while permanent data may be stored, for example, in internal mass storage 1047. Fast storage and retrieval of any memory device may be achieved by using a cache memory, which may be closely associated with one or more CPUs 1041, GPU 1042, mass storage 1047, ROM 1045, RAM 1046, and the like.
[0143] The computer readable medium may have thereon computer codes for performing various computer-implemented operations. The media and computer codes may be those specially designed and constructed for the purposes of the present disclosure, or they may be of a type well known and available to those skilled in the art of computer software.
[0144] As an example and not by way of limitation, a computer system having the architecture of computer system 1000, and specifically core 1040, can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such a computer-readable medium can be a medium associated with certain storage of the non-transitory nature of the core 1040 and the user-accessible mass storage as described above, such as core internal mass storage 1047 or ROM 1045. Software implementing various embodiments of the present disclosure can be stored in such a device and executed by core 1040. Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can enable the core 1040 and specifically the processor therein (including CPU, GPU, FPGA, etc.) to perform a specific process or a specific part of a specific process described herein, including defining a data structure stored in RAM 1046 and modifying such a data structure according to a process defined by the software. Additionally or alternatively, the computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator 1044) that may replace software or operate in conjunction with software to perform specific processes or specific portions of specific processes described herein. Where appropriate, references to software may encompass logic and vice versa. Where appropriate, references to computer-readable media may encompass circuits (e.g., integrated circuits (ICs)) storing software for execution, circuits embodying logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0145] Although the present disclosure has described a number of non-limiting embodiments, there are changes, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it should be understood that those skilled in the art will be able to devise many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.
Claims
1. A method for decoding video data, the method being performed by at least one processor, the method comprising: receiving a media stream including data in a haptic exchange format; obtaining a wavelet effect value from the data in the haptic exchange format of the media stream; as well as The media stream is decoded based on the wavelet effect value.
2. The method according to claim 1, wherein: The haptic exchange format is a .hjif format, and the data is in the .hjif format.
3. The method according to claim 2, wherein: The wavelet effect value is obtained from the "band_type" attribute of the data in the .hjif format.
4. The method according to claim 3, wherein: The wavelet effect value is indicated as "binary wavelet" in the "band_type" attribute of the data in the .hjif format.
5. The method according to claim 4, wherein: The wavelet effect value indicates: The frequency bands of the media stream are in a binary coded wavelet stream, and Entropy decoding together with inverse wavelet transform is required to decode the waves of the frequency band of the media stream.
6. The method according to claim 1, wherein: The wavelet effect value indicates a frequency band of the media stream in a binary-coded wavelet stream.
7. The method according to claim 6, wherein: The wavelet effect value indicates waves of a frequency band that requires entropy decoding together with an inverse wavelet transform to decode the media stream.
8. The method according to claim 1, wherein: Decoding the media stream includes decoding binary wavelet keyframes in .hjif format by running a base64 decode of the media stream.
9. The method according to claim 8, wherein: Decoding the media stream also includes running "readWaveletEffect()" on the haptic data of the media stream.
10. The method according to claim 9, wherein: The haptic data of the media stream adopts an MPEG-I Haptic Stream (MIHS) format.
11. An apparatus for decoding video data, the apparatus comprising: at least one memory configured to store program code; and At least one processor is configured to read the program code and perform operations according to instructions of the program code, wherein the program code includes: receiving code configured to cause the at least one processor to receive a media stream including data in a haptic exchange format; obtaining code configured to cause the at least one processor to obtain a wavelet effect value from the data in the haptic exchange format of the media stream; and A decoding code is configured to cause the at least one processor to decode the media stream based on the wavelet effect value.
12. The device according to claim 11, wherein The haptic exchange format is a .hjif format, and the data is in the .hjif format.
13. The device according to claim 12, wherein: The wavelet effect value is obtained from the "band_type" attribute of the data in the .hjif format.
14. The device according to claim 13, wherein: The wavelet effect value is indicated as "binary wavelet" in the "band_type" attribute of the data in the .hjif format.
15. The device according to claim 14, wherein: The wavelet effect value indicates: The frequency bands of the media stream are in a binary coded wavelet stream, and Entropy decoding together with inverse wavelet transform is required to decode the waves of the frequency band of the media stream.
16. The device according to claim 11, wherein The wavelet effect value indicates a frequency band of the media stream in a binary-coded wavelet stream.
17. The device according to claim 16, wherein: The wavelet effect value indicates waves of a frequency band that requires entropy decoding together with an inverse wavelet transform to decode the media stream.
18. The device according to claim 11, wherein Decoding the media stream includes decoding binary wavelet keyframes in .hjif format by running a base64 decode of the media stream.
19. The device according to claim 18, wherein: Decoding the media stream also includes running "readWaveletEffect()" on the haptic data of the media stream.
20. A non-transitory computer-readable medium storing instructions, the instructions comprising: One or more instructions that, when executed by one or more processors of a device for decoding haptic data, cause the one or more processors to: receiving a media stream including data in a haptic exchange format; obtaining a wavelet effect value from the data in the haptic exchange format of the media stream; and The media stream is decoded based on the wavelet effect value.