Method for compliance check of haptic files, streams and haptic decoders
By receiving and verifying the pattern and semantics of haptic media streams, compliance checks for haptic exchange formats and binary streams in multimedia presentation are solved to ensure the effectiveness and consistency of haptic experiences.
Patent Information
- Application Number
- CN202480004966.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-16
- Filing Date
- 2024-04-17
- Publication Date
- 2025-07-04
AI Technical Summary
The lack of compliance checking process for haptic exchange formats and binary flows in existing multimedia presentation technologies has led to the inability to ensure the effectiveness and consistency of haptic experiences.
Provides a method and apparatus to ensure that the media stream complies with haptic specifications and output compliance reports by receiving and verifying its mode and semantics, decoding haptic exchange formats and streaming formats.
It realizes effective compliance checks for haptic media streams, ensures the effectiveness and consistency of haptic experiences, and supports compliance verification of haptic exchange formats and binary streams.
Smart Images

Figure BDA0005411363170000181 
Figure BDA0005411363170000182 
Figure BDA0005411363170000191
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the priority of U.S. Provisional Application No. 63 / 459,934, filed on April 17, 2023, and U.S. Application No. 18 / 636,700, filed on April 16, 2024. The disclosures of these applications are hereby incorporated by reference in their entirety. Technical field
[0003] This disclosure relates to a set of advanced video coding and decoding techniques. More specifically, this disclosure relates to methods for encoding and decoding haptic experiences for multimedia presentations, as well as conformance checking for haptic interchange formats and binary streams. Background art
[0004] Haptic experiences have become part of multimedia presentations. In applications where multimedia presentations include aspects of haptic experiences, haptic signals can be delivered to devices or wearables, and users can feel haptic sensations in coordination with visual and / or audio media experiences during the use of the application.
[0005] Recognizing the increasing popularity of haptic experiences in multimedia presentations, the Motion Picture Experts Group (MPEG) has started researching compression standards for haptics (for both MPEG - DASH and MPEG - I) and carrying compressed haptic signaling in an ISO - based media file format (ISO BMFF).
[0006] One of the problems to be solved in aspects of multimedia presentations involving haptic experiences is that the haptic committee draft does not include any process for conformance checking of haptic interchange formats or MIHS binary streams. In view of this, since a solution to this problem is needed, aspects of the present invention provide methods for conformance checking of such files and streams. Summary of the invention
[0007] According to one aspect of the present disclosure, there is provided an apparatus, and similarly, a method and a computer-readable medium. The apparatus includes at least one memory configured to store computer program code; and at least one processor configured to access the computer program code and operate in accordance with the instructions of the computer program code. The computer program code includes: receiving code configured to cause the at least one processor to receive a media stream including data in one of a haptic exchange format and a haptic streaming format; verification code configured to cause the at least one processor to, in the case where the data is in the haptic exchange format: obtain the pattern of the input document of the media stream and the semantics of the input document; and verify at least one of the pattern against a reference pattern and the semantics against a rule table; and in the case where the data is in the haptic streaming format: decode a bitstream from at least a part of the media stream; and verify at least one of the syntax and conditions of the bitstream; and control code configured to cause the at least one processor to control the decoding of the media stream based on any one of: verifying at least one of the syntax and conditions of the bitstream; and verifying at least one of the pattern against a reference pattern and the semantics against a rule table.
[0008] The haptic exchange format may be the.hjif format.
[0009] The data may be in the haptic exchange format, and wherein the rule table may refer to haptic specifications.
[0010] In the case where the data is in the haptic exchange format, verifying at least one of the pattern against a reference pattern and the semantics against a rule table includes verifying the pattern against a reference pattern and also verifying the semantics against a rule table.
[0011] Verifying the pattern against a reference pattern and also verifying the semantics against a rule table may include verifying the pattern against a reference pattern before verifying the semantics against a rule table.
[0012] Verifying the pattern against a reference pattern and also verifying the semantics against a rule table may include verifying the semantics against a rule table in response to having previously verified the pattern against a reference pattern.
[0013] The haptic streaming format may be the MPEG-I Haptic Stream (MIHS) format.
[0014] Decoding the bitstream may include decoding a plurality of MIHS units and packets of the media stream.
[0015] In the case where the data is in the haptic exchange format, verifying at least one of the pattern against a reference pattern and the semantics against a rule table may include outputting a report indicating the result of verifying at least one of the pattern against a reference pattern and the semantics against a rule table.
[0016] In the case where the data is in a haptic streaming format, at least one of the syntax and conditions of the verification code stream may include an output report that indicates the result of at least one of the syntax and conditions of the verification code stream.
[0017] Additional embodiments will be set forth in the following description and, to the extent that they are understandable from the description, and / or can be learned by practicing the embodiments presented in the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0019] Figure 1 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment of the present disclosure;
[0020] Figure 2 is a schematic diagram of a simplified block diagram of a streaming system according to an embodiment of the present disclosure;
[0021] Figure 3 is an example illustration according to an embodiment of the present disclosure;
[0022] Figure 4 is an example illustration according to an embodiment of the present disclosure;
[0023] Figure 5A is an example illustration according to an embodiment of the present disclosure;
[0024] Figure 5B is an example illustration according to an embodiment of the present disclosure;
[0025] Figure 6 is an exemplary flowchart showing a process for processing haptic media according to an embodiment of the present disclosure;
[0026] Figure 7 is an exemplary diagram showing some aspects according to an embodiment of the present disclosure;
[0027] Figure 8 is an exemplary diagram showing some aspects according to an embodiment of the present disclosure;
[0028] Figure 9 is an exemplary diagram showing some aspects according to an embodiment of the present disclosure;
[0029] Figure 10 is an exemplary diagram showing some aspects according to an embodiment of the present disclosure;
[0030] Figure 11 is an exemplary diagram showing some aspects according to an embodiment of the present disclosure;
[0031] Figure 12 is an exemplary diagram showing some aspects of an embodiment according to the present disclosure; and
[0032] Figure 13 is an exemplary diagram showing some aspects of an embodiment according to the present disclosure. Detailed Description
[0033] According to one aspect of the present disclosure, there are provided a method, a system, and a non-transitory storage medium for parallel processing of dynamic mesh compression. Embodiments of the present disclosure may also be applied to static meshes.
[0034] Reference Figure 1 and Figure 2 describe embodiments of the present disclosure for implementing the encoding and decoding structures of the present disclosure.
[0035] Figure 1 FIG. shows a simplified block diagram of a communication system 100 according to an embodiment of the present disclosure. The system 100 may include at least two terminals 110, 120 interconnected via a network 150. For unidirectional transmission of data, a first terminal 110 may encode video data (which may include mesh data) at a local location for transmission via the network 150 to another terminal 120. The second terminal 120 may receive the encoded video data of another terminal from the network 150, decode the encoded data, and display the restored video data. Unidirectional data transmission may be common in media service applications and the like.
[0036] Figure 1 FIG. shows a second pair of terminals 130, 140 provided to support bidirectional transmission of encoded video that may occur, for example, during a video conference. For bidirectional transmission of data, each terminal 130, 140 may encode video data captured at a local location for transmission via the network 150 to another terminal. Each terminal 130, 140 may also receive the encoded video data sent by another terminal, may decode the encoded data, and may display the restored video data at a local display device.
[0037] In Figure 1Among them, the terminals 110-140 can be, for example, servers, personal computers, smart phones, and / or any other type of terminal. For example, the terminals (110-140) can be laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network 150 represents any number of networks that transfer the encoded video data between the terminals 110-140, including, for example, wired and / or wireless communication networks. The communication network 150 can exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, unless otherwise explained below, the architecture and topology of the network 150 may be unimportant for the operation of the present disclosure.
[0038] Figure 2 Shows the placement of a video encoder and decoder as an example of an application for the disclosed subject matter in a streaming environment. The disclosed subject matter can be used with other video-supported applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0039] As Figure 2 As shown, the streaming system 200 can include a capture subsystem 213, and the capture subsystem 213 includes a video source 201 and an encoder 203. The streaming system 200 can also include at least one streaming server 205 and / or at least one streaming client 206.
[0040] The video source 201 can create a stream 202 that includes, for example, a 3D mesh and metadata associated with the 3D mesh. The video source 201 can include, for example, a 3D sensor (such as a depth sensor) or 3D imaging technology (such as a digital camera) and a computing device configured to generate a 3D mesh using the data received from the 3D sensor or 3D imaging technology. The sample stream 202, which may have a high data volume when compared with an encoded video bitstream, can be processed by an encoder 203 coupled to the video source 201. The encoder 203 can include hardware, software, or a combination thereof to implement or realize aspects of the disclosed subject matter, as described in more detail below. The encoder 203 can also produce an encoded video bitstream 204. The encoded video bitstream 204, which may have a lower data volume when compared with the uncompressed stream 202, can be stored on the stream server 205 for future use. One or more streaming clients 206 can access the streaming server 205 to retrieve the video bitstream 209, and the video bitstream 209 can be a copy of the encoded video bitstream 204.
[0041] The streaming client 206 may include a video decoder 210 and a display 212. The video decoder 210 may, for example, decode a video bitstream 209 (which is an input copy of the encoded video bitstream 204) and create an output video sample stream 211 that can be rendered on the display 212 or another rendering device (not depicted). In some streaming systems, the video bitstreams 204, 209 may be encoded according to certain video coding / compression standards.
[0042] Figure 3 May be a functional block diagram of a video decoder 300 according to an embodiment of the present invention.
[0043] The receiver 302 may receive one or more codec video sequences to be decoded by the decoder 300; in the same or another embodiment, one encoded video sequence at a time, where the decoding of each encoded video sequence is independent of the other encoded video sequences. The encoded video sequences may be received from a channel 301, which may be a hardware / software link to a storage device storing the encoded video data. The receiver 302 may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams, that may be forwarded to their respective using entities (not depicted). The receiver 302 may separate the encoded video sequences from the other data. To counteract network jitter, a buffer memory 303 may be coupled between the receiver 302 and the entropy decoder / parser 304 (hereinafter referred to as "parser"). When the receiver 302 receives data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer 303 may be unnecessary or may be small. For use on a best-effort packet network such as the Internet, the buffer 303 may be necessary, may be relatively large, and may advantageously have an adaptive size.
[0044] Video decoder 300 may include a parser 304 to reconstruct symbols 313 from an entropy-coded video sequence. The categories of these symbols include information for managing the operation of decoder 300 and potential information for controlling a display device (e.g., display 312), which is not a part of the decoder but may be coupled to the decoder. The control information for the display device may be Supplemental Enhancement Information (SEI messages) or a parameter set segment of video usability information (not labeled). Parser 304 may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence may be performed according to a video coding technology or standard and may follow principles well-known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. Parser 304 may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to a group. Subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), and so on. The entropy encoder / parser may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.
[0045] Parser 304 may perform entropy decoding / parsing operations on the video sequence received from buffer memory 303 to create symbols 313. Parser 304 may receive the encoded data and selectively decode specific symbols 313. In addition, parser 304 may determine whether to provide specific symbols 313 to motion compensation prediction unit 306, scaler / inverse transform unit 305, intra prediction unit 307, or loop filter 311.
[0046] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of symbols 313 may involve multiple different units. Which units are involved and the way they are involved may be controlled by subgroup control information parsed by parser 304 from the encoded video sequence. For the sake of brevity, such subgroup control information flows between parser 304 and multiple units below are not described.
[0047] In addition to the functional blocks already mentioned, decoder 300 can be conceptually divided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually divide into the functional units below.
[0048] The first unit is the scaler / inverse transform unit 305. The scaler / inverse transform unit 305 receives, from the parser, the quantized transform coefficients as symbols 313 and control information, including which transform mode to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit can output a block including sample values, and the sample values can be input into the aggregator 310.
[0049] In some cases, the output samples of the scaler / inverse transform unit 305 can belong to intra-coded blocks; that is, blocks that do not use predictive information from previously reconstructed pictures but can use predictive information from previously reconstructed parts of the current picture. Such predictive information can be provided by the intra-picture prediction unit 307. In some cases, the intra-picture prediction unit 307 generates surrounding blocks of the same size and shape as the block being reconstructed, using the reconstructed information extracted from the currently partially reconstructed picture 309. In some cases, based on each sample, the aggregator 310 adds the prediction information generated by the intra-picture prediction unit 307 to the output sample information provided by the scaler / inverse transform unit 305.
[0050] In other cases, the output samples of the scaler / inverse transform unit 305 can belong to inter-coded and potentially motion-compensated blocks. In this case, the motion-compensation prediction unit 306 can access the reference picture memory 308 to extract samples for prediction. After motion-compensating the extracted samples according to the symbol 313, these samples can be added by the aggregator 310 to the output of the scaler / inverse transform unit (which is called the residual sample or residual signal in this case), thereby generating output sample information. The motion-compensation unit obtaining the prediction samples from the address within the reference picture memory can be controlled by a motion vector, and the motion vector is in the form of the symbol for use by the motion-compensation unit), such as a symbol including X, Y, and reference picture components. Motion compensation can also include interpolation of the sample values extracted from the reference picture memory when using sub-sample accurate motion vectors, a motion vector prediction mechanism, and so on.
[0051] The output samples of the aggregator 310 can be employed by various loop filtering techniques in the loop filter unit 311. Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the encoded bitstream, and the parameters can be available to the loop filter unit 311 as symbols 313 from the parser 304. However, in other embodiments, the video compression techniques can also respond to meta-information obtained during the decoding of the previous (in decoding order) parts of the encoded picture or the encoded video sequence, and to the previously reconstructed and loop-filtered sample values.
[0052] The output of the loop filter unit 311 can be a sample stream that can be output to the display device 312 and stored in the reference picture memory 557 for subsequent inter-picture prediction.
[0053] Once fully reconstructed, some encoded pictures can be used as reference pictures for future prediction. Once an encoded picture is fully reconstructed and the encoded picture is identified as a reference picture (e.g., by the parser 304), the current reference picture 309 can become part of the reference picture buffer 308, and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.
[0054] The video decoder 300 can perform decoding operations according to a predetermined video compression technique that can be recorded, for example, in the ITU-T H.265 standard. As specified in the video compression technique document or standard (especially in the profile of this application), in the sense that the encoded video sequence follows the syntax of the video compression technique or standard, the encoded video sequence can conform to the syntax specified by the video compression technique or standard being used. For compliance, it is also required that the complexity of the encoded video sequence is within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata for HRD buffer management signaled in the encoded video sequence.
[0055] In an embodiment, the receiver 302 can receive additional (redundant) data along with the encoded video. The additional data can be part of the encoded video sequence. The additional data can be used by the video decoder 300 to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0056] Figure 4 It may be a functional block diagram of a video encoder 400 according to an embodiment disclosed in the present application.
[0057] The encoder 400 may receive video samples from a video source 401 (not part of the encoder), and the video source may capture video images to be encoded by the encoder 400.
[0058] The video source 401 may provide a source video sequence in the form of a digital video sample stream to be encoded by the encoder 303. The digital video sample stream may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit...), any color space (e.g., BT.601 Y CrCb, RGB...), and any suitable sampling structure (e.g., YCrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source 401 may be a storage device storing previously prepared videos. In a video conferencing system, the video source 401 may be a camera capturing local image information as a video sequence. The video data may be provided as a plurality of individual pictures, which are given motion when viewed in sequence. The pictures themselves may be constructed as a spatial pixel array, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.
[0059] According to an embodiment, the encoder 400 may encode and compress pictures of the source video sequence into an encoded video sequence 410 in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller 402. The controller controls other functional units as described below and is functionally coupled to these units. For the sake of simplicity, the couplings are not labeled in the figure. The parameters set by the controller may include rate control related parameters (picture skip, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. When other suitable functions are involved in the video encoder 400 optimized for a certain system design, those skilled in the art can easily identify these other suitable functions of the controller 402.
[0060] Some video encoders operate in an "encoding loop" that is readily recognizable to those skilled in the art. As a simple description, the encoding loop may include an encoding portion of encoder 402 (hereinafter referred to as the "source encoder") that is responsible for creating symbols based on an input picture to be encoded and reference pictures, and a (local) decoder 406 embedded in encoder 400 that reconstructs the symbols in the same manner as the (remote) decoder creates sample data to create sample data (since in the video compression techniques contemplated in this application, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream is input to reference picture memory 405. Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the reference picture buffer content is also bit-exactly corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction portion of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is well known to those skilled in the art.
[0061] The operation of the "local" decoder 406 may be the same as that of the "remote" decoder 300 described in detail above in connection with Figure 3 However, briefly referring additionally to Figure 4 , when the symbols are available and the entropy encoder 408 and the parser 304 can encode / decode the symbols losslessly into the encoded video sequence, the entropy decoding portion of the decoder 300 (including the channel 301, the receiver 302, the buffer 303, and the parser 304) may not be fully implemented in the local decoder 406.
[0062] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding present in the decoder must also exist in the corresponding encoder in substantially the same functional form. The description of the encoder technology can be simplified because the encoder technology is reciprocal to the decoder technology described comprehensively. More detailed descriptions are only needed in certain areas and are provided below.
[0063] As part of its operation, the source encoder 403 may perform motion compensation predictive coding. Referring to one or more previously encoded frames in the video sequence designated as "reference frames", the motion compensation predictive coding performs predictive coding on the input frame. In this way, the coding engine 407 encodes the difference between the pixel blocks of the input frame and the pixel blocks of the reference frame, and the reference frame can be selected as the prediction reference for the input frame.
[0064] The local video decoder 406 may decode the encoded video data of a frame that may be designated as a reference frame based on the symbols created by the source encoder 403. The operation of the encoding engine 407 may be a lossy process. When the encoded video data may be decoded at a video decoder ( Figure 4 not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder 406 replicates the decoding process that may be performed by the video decoder on the reference frame and may cause the reconstructed reference frame to be stored in the reference picture cache 405. In this way, the encoder 400 may locally store a copy of the reconstructed reference frame that has the same content (without transmission errors) as the reconstructed reference frame that will be obtained by the remote video decoder.
[0065] The predictor 404 may perform a prediction search for the encoding engine 407. That is, for a new frame to be encoded, the predictor 404 may search the reference picture memory 405 for sample data (as a candidate reference pixel block) or some metadata, such as a reference picture motion vector, block shape, etc., that may serve as an appropriate prediction reference for the new picture. The predictor 404 may operate block by block based on sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor 404, it may be determined that the input picture may have a prediction reference taken from multiple reference pictures stored in the reference picture memory 405.
[0066] The controller 402 may manage the encoding operations of the video encoder 403, including, for example, setting parameters and subgroup parameters for encoding the video data.
[0067] The outputs of all the above functional units may be entropy encoded in the entropy encoder 408. The entropy encoder performs lossless compression on the symbols generated by the various functional units according to techniques well known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.
[0068] The transmitter 409 may buffer the encoded video sequence created by the entropy encoder 408 to prepare for transmission over a communication channel 411, which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter 409 may merge the encoded video data from the video encoder 403 with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).
[0069] The controller 402 may manage the operation of the encoder 400. During encoding, the controller 405 may assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques that may be applied to the corresponding picture. For example, a picture may typically be assigned to any one of the following frame types:
[0070] An intra picture (I picture) is a picture that can be encoded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh pictures. Those skilled in the art are aware of the variations of I pictures and their corresponding applications and characteristics.
[0071] A predictive picture (P picture) is a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.
[0072] A bi-predictive picture (B picture) is a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.
[0073] Source pictures can generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, which are determined according to the coding assignment of the corresponding picture applied to the 'block'. For example, blocks of an I picture can be non-predictively encoded, or the blocks can be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be non-predictively encoded by spatial prediction or by temporal prediction with reference to one previously encoded reference picture. Blocks of a B picture can be non-predictively encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference pictures.
[0074] Video encoder 400 can perform encoding operations according to a predetermined video coding technique or standard such as the ITU-T H.265 recommendation. In operation, video encoder 400 can perform various compression operations, including predictive coding operations that exploit the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0075] In an embodiment, transmitter 409 can transmit additional data when transmitting the encoded video. Source encoder 403 can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures, and slices, supplementary enhancement information (SEI) messages, video usability information (VUI) parameter set fragments, etc.
[0076] Reference Figures 5A to 5B, embodiments of the present disclosure for implementing a haptic encoder 500 and a haptic decoder 550 are described.
[0077] As Figure 5A shown, the haptic encoder 500 can receive descriptive haptic data and waveform haptic data. Thus, the haptic encoder 500 can be capable of processing three types of input files:.ohm metadata files (object haptic metadata - a text file format for haptic metadata), descriptive haptic files (.ivs,.ahap, and.hjif), or waveform PCM files (.wav). Examples of descriptive data can include:.ahap (Apple Haptic and Audio Pattern - a JSON-like file format specifying haptic patterns) from Apple (representing the expected haptic output through a set of modulated continuous signals and a set of parameterized modulation transients);.ivs (representing the expected haptic output through a set of base effects parameterized by a set of parameters) from Immersion; or.hjif (Haptic JSON Interchange Format), i.e., the MPEG format proposed in this application. Examples of waveform pulse code modulation (PCM) signals can include.ohm input files that include metadata information.
[0078] According to an embodiment, the haptic encoder 500 can process two types of input files differently. For descriptive content, the haptic encoder 500 can semantically analyze the input to transcoding (if necessary) the data into the encoding representation proposed in this application.
[0079] According to an embodiment, the.ohm metadata input file can include a description of the haptic system and settings. In particular, the.ohm metadata input file can include the name of each associated haptic file (descriptive or PCM), and a description of the signal. The.ohm metadata input file also provides a mapping between each channel of the signal and the target body part on the user's body. For the.ohm metadata input file, the haptic encoder performs metadata extraction by retrieving the associated haptic file from the URI, encodes the metadata based on the type of metadata, and maps the metadata extracted from the.ohm file to the metadata information of the data model.
[0080] According to an embodiment, descriptive haptic files (e.g.,.ivs,.ahap, and.hjif) can be encoded through a simple process. The haptic encoder 500 first specifically identifies the input format. If the input format is a.hjif file, then no transcoding is required, and the file can be further edited, compressed into a binary format, and finally packaged into a MIHS stream. If an.ahap or.ivs input file is used, transcoding is necessary. The haptic encoder 500 first semantically analyzes the input file information and transcodes it to be formatted into a selected data model. After transcoding, the data can be exported as a.hjif file,.hmpg binary file, or MIHS stream.
[0081] According to an embodiment, the haptic encoder 500 can perform signal analysis to parse the signal structure of a.wav file and convert it into the encoded representation proposed in this application. For waveform PCM content, the signal analysis process can be divided into two sub-processes performed by the haptic encoder 500. After performing band decomposition on the signal, in the first sub-process, the low frequency can be encoded using a keyframe extraction process. Then the low frequency band can be reconstructed, and the error between the signal and the original low frequency signal can be calculated. Then, before encoding using wavelet transform, the residual signal can be added to the original high frequency band, and encoding using wavelet transform is the second sub-process. According to an embodiment, when multiple low frequency bands are used, the residuals from all low frequency bands are added to the high frequency band before encoding. In an embodiment using multiple high frequency bands, the residuals from the low frequency band are added to the first high frequency band before encoding.
[0082] According to an embodiment, keyframe extraction includes obtaining the lower frequency band from band decomposition and analyzing its content in the time domain. According to an embodiment, wavelet processing can include obtaining the high frequency band from band decomposition and the low frequency residuals, and dividing it into blocks of equal size. Then these signal blocks of equal size are analyzed in a psychohaptic model. Lossy compression can be applied by performing wavelet transform on the blocks and quantifying them with the help of the psychohaptic model. Finally, each block is saved into a separate effect in a single frequency band, which is done in formatting. Binary compression can apply lossless compression using appropriate coding techniques (e.g., the set partitioning in hierarchical trees (SPIHT) algorithm and arithmetic coding (AC)).
[0083] As Figure 5AAs shown, the tactile encoder 500 can be configured to encode descriptive and quantified tactile data and can output three types of formats - an interchange format (.hjif), a binary compressed format (.hmpg), and a streaming format (e.g., MPEG immersive haptic stream (MIHS)). The.hjif format is a JSON-based human-readable format that can be easily parsed and manually edited, making it an ideal interchange format, especially when designing / creating content. For distribution purposes, the.hjif data can be compressed into a more memory-efficient binary.hmpg bitstream. This compression may be lossy, and different parameters affect the encoding depth of the amplitudes and frequencies that make up the bitstream. For streaming purposes, the data can be compressed and packaged into an MPEG-1 haptic stream (MIHS). The above three formats have complementary purposes and can be lossily converted one-to-one between them.
[0084] As Figure 5B shown, the tactile decoder 550 can take as input a.hmpg compressed binary file format or an MIHS bitstream. The tactile decoder 550 can output the.hjif interchange format that can be directly used for rendering. These two input formats can undergo binary decompression to extract the metadata and the data itself from the file and map it to the selected data structure. Then, the data can be exported to the tactile renderer 580 in the.hjif format.
[0085] As Figure 5B shown, the renderer 580 includes a synthesizer. The synthesizer can render tactile data from the.hjif input file into a PCM output file. The rendering and / or synthesis is illustrative. According to an embodiment, the synthesizer parses the input file and performs advanced synthesis distributions between vectors, wavelets, etc. Then, the synthesis process descends to the band components of the codec where the synthesis process is called. Then, all the bands of a given channel are mixed through a simple addition operator to recreate the desired tactile signal.
[0086] According to an embodiment, the tactile experience defines the root of a hierarchical data model. It provides information about the file date and format version, it describes the tactile experience, it lists the different avatars (i.e., body representations) used throughout the experience, and it defines all tactile perceptions.
[0087] According to an embodiment, a self - contained stream format for transmitting MPEG - I haptic data may use a packetization method and may include two levels of packetization: an MPEG - I Haptic Stream (MIHS) unit that covers a duration and includes zero or more MIHS packets; and an MIHS packet that includes metadata or haptic effect data. In an embodiment, the MIHS unit may be referred to as a network abstraction layer unit associated with haptic data. In an embodiment, the MIHS unit may be referred to as an MIHS sample associated with haptic data.
[0088] According to an embodiment, the MIHS unit may be a synchronous unit or an asynchronous unit. The synchronous unit resets the previous effect and thus provides a haptic experience independent of the previous MIHS unit. The asynchronous unit is a continuation of the previous MIHS unit and cannot be independently decoded and rendered without decoding the previous MIHS unit.
[0089] According to an embodiment, haptic signals may be encoded on multiple channels. In some embodiments, a haptic channel may define a signal to be rendered at a specific body location using a dedicated actuator / device. Metadata stored at the channel level may include information such as gain associated with the channel, mixing weights, desired body location for haptic feedback, and optionally reference device and / or orientation. Additional information such as desired sampling frequency or sampling count may also be provided. Finally, the haptic data of a channel is contained within a set of haptic frequency bands defined by its frequency range. The haptic frequency bands describe the haptic signals of the channel within a given frequency range. The bands are defined by a list of types and orders of haptic effects, each haptic effect including a set of keyframes. For each type of haptic band, the haptic effect may be defined by at least position and type. The position may indicate the temporal or spatial position of the effect. In some embodiments, the value 0 is the relative start position of the experience, which is a dependent variable of the configured perceptual modality. The default unit for temporal haptic feedback may be milliseconds, while the default unit for spatial haptic feedback may be millimeters. This embodiment discloses the "start position of the experience" because the binary distribution format does not have the concept of any finite time interval (i.e., frame or sample).
[0090] According to the type of the band and the type of the effect, additional attributes may be specified, including phase, base signal, composition, and multiple consecutive haptic keyframes describing the effect.
[0091] According to an embodiment, a haptic data hierarchy is defined in the present disclosure.
[0092] · Haptic channel
[0093] o Haptic frequency band
[0094] ■ Haptic effect
[0095] Embodiments of the present disclosure describe two anchor points for the position of haptic effects relative to an ISOBMFF track.
[0096] Figure 6 The first embodiment 600 is shown. As Figure 6 shown, each MIHS unit (also referred to as an MIHS sample, an ISOBMFF haptic sample, or a sample in the embodiment) includes one or more haptic channel information and one or more haptic band information. As described above, each MIHS unit consists of one or more channels, and each channel consists of one or more bands. Then each band can have one or more effects.
[0097] In the first embodiment, the temporal position of an effect can be defined as an offset relative to the start timing (e.g., the MIHS unit start time) of the sample carrying the effect. In a second embodiment or the same embodiment, the offset is based on the start time and / or presentation time of the media or haptic track.
[0098] According to an embodiment, the first embodiment enables the manipulation of the track without affecting the position of the haptic effect, because any change in the ISOBMFF sampling timing does not affect the relative position of the effect. According to an embodiment, in the case of a basic haptic stream (e.g., a high-level syntax stream), the second embodiment can be used when the basic stream is used without ISOBMFF.
[0099] According to an embodiment, multiple types of haptic tracks can be used. In an embodiment, samples or MIHS units can be used in a haptic track, and the temporal position of the effect of the sample or MIHS unit is defined as an offset relative to the start timing of the sample. According to another embodiment, for example Figure 7 in the example 700, samples or MIHS units are used in a haptic track, and the effect of the sample or MIHS unit has a temporal position relative to the start time of the track. In another embodiment, a mixed MIHS unit or sample can be used.
[0100] Embodiments of the present disclosure provide a timing model that can be used to synchronize haptic effects with other media tracks in the same or related ISOBMFF file. In the case of the timing model of the haptic track relative to the timing model of the related ISOBMFF file, the manipulation and processing of the media track become more efficient.
[0101] As Figure 8 shown, the process 800 shows an exemplary process for decoding haptic data.
[0102] At operation 805, a media stream including one or more haptic tracks and one or more video tracks can be received.
[0103] At operation 810, one or more Moving Picture Experts Group (MPEG) Immersive Haptic Streams (MIHS) units can be obtained from a media stream. In some embodiments, the MIHS units can include one or more haptic effects. The MIHS units can also include the start time of the MIHS units.
[0104] In an embodiment, the MIHS unit is associated with at least one haptic channel, the at least one haptic channel includes one or more haptic bands, and each of the one or more haptic bands has at least one haptic effect.
[0105] At operation 815, timing information associated with one or more haptic effects can be obtained. In an embodiment, the timing information can include at least one time position of the one or more haptic effects.
[0106] In an embodiment, the time position of the haptic effect indicates the effect start time of the haptic effect, wherein the effect start time of the haptic effect is an offset based on the start time of the corresponding MIHS unit. The effect start time can indicate the start time of the haptic effect relative to the start time of the corresponding MIHS unit.
[0107] In an embodiment, the effect start time of the haptic effect is an absolute time based on the start time of at least one haptic track or at least one video track.
[0108] At operation 820, the media stream is rendered based on the obtained timing information.
[0109] According to an embodiment, manipulation of the order of one or more MIHS units does not affect at least one time position of one or more haptic effects because the one or more MIHS units correspond to one or more ISO-based media file format (ISOBMFF) samples associated with at least one video track.
[0110] In some embodiments, synchronized MIHS units can be obtained from the media stream. In an embodiment, the synchronized MIHS unit is a special type of MIHS unit configured to provide a reset point in the bitstream. In an embodiment, the synchronized MIHS unit is mapped to synchronization samples in the video bitstream corresponding to one or more haptic channels.
[0111] As in Figure 9 Example 900, a haptic encoder according to an embodiment herein generates a compact and efficient binary distribution format (.hmpg) for distribution. A haptic decoder can decode this format and send it to a renderer.
[0112] The ISOBMFF haptic binding Working draft defines the following for haptic samples in an ISOBMFF track:
[0113]
[0114] As described above, the data_packet_type is set to a fixed value.
[0115] An ISOBMFF haptic sample according to an embodiment of this document may have the following data packets:
[0116] 1. A silent packet with no data, i.e., a zero-data packet.
[0117] 2. One or more time data packets.
[0118] 3. One or more spatial data packets.
[0119] 4. A combination of 2 and 3
[0120] 5. A possible combination of 2 and 3, without any commitment, i.e., zero or more time and / or data packets.
[0121] Embodiments of this document may use the data_packet_type to signal the above conditions such that the data_packet_type (i.e., b3b2b1b0) may be signaled as follows:
[0122] Table 1 - Data Packet / Sample Type
[0123]
[0124]
[0125] Where, according to an embodiment, data_packet_type = 0 indicates that the sample is a silent sample, i.e., equivalent to one or more consecutive MIHS silent units.
[0126] Therefore, the following values of data_packet_type in Table 2 have the following meanings:
[0127] Table 2
[0128]
[0129]
[0130] And, according to an embodiment, the file format parser may use the data_packet_type in the following manner:
[0131] 1. Identify silent samples when processing a file and skip further parsing of the silent samples, thereby accelerating random access and fast forward / rewind navigation through the file.
[0132] 2. Identify which samples have only time data groups for extracting time effects.
[0133] 3. Identify which samples have only spatial data groups for extracting spatial effects.
[0134] 4. When decomposing an orbit into multiple orbits, it is easier to separate time samples and spatial samples.
[0135] 5. When combining multiple orbits into a single orbit, determine the sample structure of the new orbit and whether the samples should combine time groups and spatial groups or should keep them in separate samples.
[0136] Accordingly, according to an embodiment, a method for signaling the sample type in an ISOBMFF haptic sample is proposed, wherein the sample is identified as a silent sample, a non - silent sample, a non - silent sample having only time effects, a non - silent sample having only spatial effects, a non - silent sample having a combination of time and spatial effects, wherein signaling is implemented using bit - based flags for different features, wherein file format parsing can utilize the information provided by the sample type and navigate through the file faster due to skipping silent samples, or use the sample type information to extract only spatial information, only time information, or only silent information, or use the information for bitstream manipulation, single - track to multi - track conversion, and multi - track to single - track conversion.
[0137] As in Figure 9 Example 900, a haptic encoder according to an embodiment herein generates a compact and efficient binary distribution format (.hmpg) for distribution. A haptic decoder can decode this format and send it to a renderer. Embodiments herein extend the band types in the.hjif format to include binary wavelets.
[0138] Figure 10 Example 1000 of an HJIF compliance tool for haptic exchange format compliance checking is shown. For example, as Figure 10 shown, there is a two - step process for checking the compliance of an HJIF file. For example, in Figure 10In this case, the HJIF file 1001 is input as the HJIF input document compliance web service. As step 1, at the schema validator 1003, schema validation is achieved by using the relevant schema 1002 to validate the schema of the HJIF file 1001. And as step 2, there is semantic validation performed by the semantic validator 1006 using a set of rule tables 1005, and through this rule table 1005, the semantics of the HJIF file 1001 are validated. Each step generates a report, such as report 1004 and report 1007, which outline which items in the document (HJIF file 1001) fail (if any).
[0139] The following information is provided to the validators, such as the schema validator 1003 and the semantic validator 1006. For step 1, the schema 1002 of the HJIF file 1001 is provided to the schema validator 1003. And for step 2, what is provided is the rule table 1005 that reflects the "shall" of the haptic specification (such as the range of parameters).
[0140] Figure 11 An example 1100 of the MIHS compliance check process is shown. That is, the embodiments herein also define a process for checking the compliance of the MIHS stream, as Figure 11 shown.
[0141] As Figure 11 shown, according to the embodiment, the MIHS input stream 1101 is input to the haptic MIHS bitstream validator 1102, so that the compliance of the MIHS input stream 1101 can be verified in a single step, such that: the haptic MIHS bitstream validator 1102 decodes the MIHS input stream 1101, decodes each MIHS unit and packet, and checks whether the syntax and conditions of the bitstream are valid according to the haptic specification. The output of the validator 1102 is the report 1103. According to the embodiment, the report 1103 can be in the form of the HJIF format. However, according to one or more embodiments, additional information, including silent MIHS units, independent effects, etc., can be provided by extending the HJIF format.
[0142] Figure 12 An example 1200 of the haptic decoder compliance check is shown. That is, the example 1200 shows a compliance check process for the haptic decoder. According to the embodiment, Figure 12Illustrated is that the MIHS reference stream 1201 can be used for haptic decoder compliance checking or the compliance checking process of a haptic decoder, such that in sequence: 1. The MIHS reference stream 1201 is fed into the haptic decoder 1202, and then 2. The output of the haptic decoder 1202 and the HJIF reference file 1203 corresponding to the MIHS reference stream 1201 are provided to the HJIF comparator 1204, and then 3. The HJIF comparator 1204 compares the items in the two HJIF files. Considering that the order of the items in the two HJIF files may be different, the HJIF file comparator 1204 compares each item in one HJIF file and finds an equivalent item in the other file. If there is any item in one file that does not have an equivalent in the other file, the HJIF file comparator 1203 generates an error that will be indicated in the output report 1205. According to an embodiment, an error is not a prerequisite for generating the report 1205.
[0143] Thus, through the embodiments herein, a method for checking a haptic exchange format is provided, wherein, first, the mode of the haptic exchange format is checked, and then the rules of the specification are checked regarding the values of the parameters in the file. A method for checking a haptic streaming format is also provided, wherein, the MIHS stream is decoded, and the syntax and rules regarding the parameter values are checked using compliance stream checker software. And a method for compliance checking of a haptic decoder is also provided, wherein, a reference MIHS stream is fed into the decoder, and the HJIF output of the decoder is compared with the corresponding reference HJIF file representing the equivalent of the MIHS stream, wherein, the comparator module compares the objects and items in the two HJIF files and finds a one-to-one equivalent for each object and item, and if an object or item present in one file does not have an equivalent in the other file, a compliance error report is generated.
[0144] Those skilled in the art will understand that the techniques described herein can be implemented on both the encoder side and the decoder side. The above techniques can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 13 Illustrated is a computer system 1300 suitable for implementing certain embodiments of the present disclosure.
[0145] Computer software can be encoded using any suitable machine code or computer language that can be subject to mechanisms such as assembly, compilation, linking, or the like to create code that includes instructions that can be directly executed by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or executed through interpretation, microcode execution, etc.
[0146] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0147] Figure 13 The components shown for computer system 1300 are examples and are not intended to imply any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should also not be construed as having any dependencies or requirements related to any one or combination of the components shown in the non-limiting embodiments of computer system 1300.
[0148] Computer system 1300 may include certain human-machine interface input devices. Such human-machine interface input devices can respond to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not depicted). The human-machine interface devices can also be used to capture certain media that are not necessarily directly related to conscious human input, such as, for example, audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0149] The input human-machine interface devices may include one or more of the following (only one of each is depicted): keyboard 1301, mouse 1302, touchpad 1303, touch screen 1310, data glove, joystick 1305, microphone 1306, scanner 1307, camera 1308.
[0150] Computer system 1300 may also include certain human-machine interface output devices. Such human-machine interface output devices can stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback provided by touch screen 1310, data glove, or joystick 1305, but there may also be tactile feedback devices that do not function as input devices). For example, such devices may be audio output devices (e.g., speaker 1309, headphones (not shown)), visual output devices (e.g., screen 1310, including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capability, each with or without tactile feedback capability - some of which may be capable of outputting two-dimensional visual output or more than three-dimensional output through means such as stereoscopic output, virtual reality glasses (not shown), holographic displays, and fog machines (not shown)), and printers (not shown).
[0151] The computer system 1300 may also include human-accessible storage devices and their associated media, such as optical media, including CD / DVD ROM / RW 1320 with CD / DVD or similar media 1321, thumb drives 1322, removable hard disk drives or solid state drives 1323, conventional magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), etc.
[0152] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.
[0153] The computer system 1300 may also include an interface to one or more communication networks. For example, the network may be wireless, wired, optical. The network may also be local, wide area, metropolitan area, vehicular and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks (such as Ethernet), wireless LAN, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), TV wired or wireless wide area networks (including cable TV, satellite TV, and terrestrial broadcast TV), vehicular and industrial networks (including CANBus), etc. Certain networks typically require an external network interface adapter, which is attached to certain general-purpose data ports or peripheral buses 1349 (e.g., the USB port of the computer system 1300; other network interface adapters are typically integrated into the core of the computer system 1300 by attaching to the system bus described below (e.g., the Ethernet interface of a PC computer system or the cellular network interface of a smart phone computer system). Using any of these networks, the computer system 1300 can communicate with other entities. Such communication may be unidirectional only receive (e.g., broadcast TV), unidirectional only send (e.g., CANbus to certain CANbus devices), or bidirectional, such as to other computer systems using local or wide area networks. Such communication may include communication to a cloud computing environment 1355. Certain protocols and protocol stacks may be used on each of the networks and network interfaces described above.
[0154] The above-mentioned human-machine interface device, human-accessible storage device, and network interface 1354 may be attached to the core 1340 of the computer system 1300.
[0155] The core 1340 may include one or more central processing units (CPUs) 1341, a graphics processing unit (GPU) 1342, a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) 1343, a hardware accelerator 1344 for certain tasks, etc. These devices, as well as a read-only memory (ROM) 1345, a random access memory 1346, an internal mass storage 1347 such as an internal non-user-accessible hard disk drive, SSD, etc., may be connected via a system bus 1348. In some computer systems, the system bus 1348 may be accessible in the form of one or more physical plugs to enable the expansion of additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the system bus 1348 of the core or attached to the system bus 1348 of the core via a peripheral bus 1349. The architecture of the peripheral bus includes PCI, USB, etc. A graphics adapter 1350 may be included in the core 1340.
[0156] The CPU 1341, GPU 1342, FPGA 1343, and accelerator 1344 may execute certain instructions that, when combined, may constitute the aforementioned computer code. The computer code may be stored in the ROM 1345 or the RAM 1346. Transitional data may also be stored in the RAM 1346, while permanent data may be stored, for example, in the internal mass storage 1347. Fast storage and retrieval of any memory device may be achieved by using a cache memory that may be closely associated with one or more CPUs 1341, GPUs 1342, mass storage 1347, ROM 1345, RAM 1346, etc.
[0157] A computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and the computer code may be those specially designed and constructed for the purposes of this disclosure, or they may be of the type well known and available to those of ordinary skill in the computer software art.
[0158] As an example and not by way of limitation, a computer system having the architecture of computer system 1300, and in particular core 1340, can provide functions that are the result of software embodied in one or more tangible computer-readable media being executed by a processor (including a CPU, GPU, FPGA, accelerator, etc.). Such computer-readable media can be media associated with user-accessible mass storage as described above and certain storage of a non-transitory nature of core 1340, such as on-chip mass storage 1347 or ROM 1345 of the core. The software implementing the various embodiments of the present disclosure can be stored in such devices and executed by core 1340. Depending on specific requirements, the computer-readable media can include one or more memory devices or chips. The software can cause core 1340 and in particular the processors therein (including the CPU, GPU, FPGA, etc.) to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM 1346 and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system can provide functions that are the result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator 1344), which can operate in place of or in conjunction with the software to execute specific processes or specific portions of specific processes described herein. In appropriate instances, references to software can encompass logic and vice versa. In appropriate instances, references to computer-readable media can encompass circuitry (e.g., an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0159] Although the present disclosure has described multiple non-limiting embodiments, there are changes, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.
Claims
1. A method for decoding video data, the method being executed by at least one processor, the method comprising: Receiving a media stream, the media stream including data in one of a haptic exchange format and a haptic streaming format; In the case where the data is in the haptic exchange format: Obtaining a mode of an input document of the media stream and a semantics of the input document; And Performing at least one of the following: verifying the mode against a reference mode and verifying the semantics against a rule table; In the case where the data is in the haptic streaming format: Decoding a bitstream from at least a part of the media stream; And Verifying at least one of a syntax and a condition of the bitstream; and Controlling decoding of the media stream based on any one of: verifying at least one of a syntax and a condition of the bitstream; and at least one of the performed verifying the mode against a reference mode and verifying the semantics against a rule table.
2. The method according to claim 1, wherein, The haptic exchange format is the.hjif format.
3. The method according to claim 2, wherein The data is in the haptic exchange format, and wherein the rule table refers to a haptic specification.
4. The method according to claim 3, wherein In the case where the data is in the haptic exchange format, performing at least one of the verifying the mode against the reference mode and verifying the semantics against the rule table includes: verifying the mode against the reference mode and also verifying the semantics against the rule table.
5. The method according to claim 4, wherein Verifying the mode against the reference mode and also verifying the semantics against the rule table includes: verifying the mode against the reference mode before verifying the semantics against the rule table.
6. The method according to claim 4, wherein, Verifying the mode against the reference mode and also verifying the semantics against the rule table includes: verifying the semantics against the rule table in response to previously verifying the mode against the reference mode.
7. The method according to claim 1, wherein Verifying the mode against the reference mode and also verifying the semantics against the rule table includes: verifying the semantics against the rule table in response to previously verifying the mode against the reference mode.
8. The method according to claim 7, wherein, Decoding the bitstream includes decoding a plurality of MIHS units and packets of the media stream.
9. The method according to claim 1, Among them, In the case where the data is in the haptic exchange format, performing at least one of the verifying the mode against a reference mode and verifying the semantics against a rule table includes: outputting a report indicating a result of at least one of the verifying the mode against a reference mode and verifying the semantics against a rule table.
10. The method according to claim 1, Among them, In the case where the data is in the haptic streaming format, verifying at least one of a syntax and a condition of the bitstream includes outputting a report indicating a result of verifying at least one of a syntax and a condition of the bitstream.
11. An apparatus for decoding video data, the apparatus comprising: At least one memory configured to store program code; And At least one processor configured to read the program code and operate according to what is indicated by the program code, the program code including: Receiving code configured to cause the at least one processor to receive a media stream, the media stream including data in one of a haptic exchange format and a haptic streaming format; Verification code configured to cause the at least one processor to: In the case where the data is in the haptic exchange format: Obtain a pattern of an input document of the media stream and a semantics of the input document; and Verify at least one of the pattern against a reference pattern and the semantics against a rule table; and In the case where the data is in the haptic exchange format: Obtain a pattern of an input document of the media stream and a semantics of the input document; and Verify at least one of the pattern against a reference pattern and the semantics against a rule table; In the case where the data is in the haptic streaming format: Decode a bitstream from at least a portion of the media stream; and Verify at least one of a syntax and a condition of the bitstream; and Control code configured to cause the at least one processor to control decoding of the media stream based on any one of: verifying at least one of a syntax and a condition of the bitstream; and verifying at least one of the pattern against a reference pattern and the semantics against a rule table.
12. The apparatus according to claim 11, wherein, The haptic exchange format is the.hjif format.
13. The device according to claim 12, wherein, The data is in the haptic exchange format, and wherein the rule table refers to a haptic specification.
14. The device according to claim 13, wherein, In the case where the data is in the haptic exchange format, verifying at least one of the pattern against the reference pattern and the semantics against the rule table includes verifying the pattern against the reference pattern and also verifying the semantics against the rule table.
15. The apparatus according to claim 14, wherein Verifying the pattern against the reference pattern and also verifying the semantics against the rule table includes verifying the pattern against the reference pattern before verifying the semantics against the rule table.
16. The device according to claim 14, wherein Verifying the pattern against the reference pattern and also verifying the semantics against the rule table includes verifying the semantics against the rule table in response to previously verifying the pattern against the reference pattern.
17. The apparatus according to claim 11, wherein, The haptic streaming format is the MPEG-I haptic stream MIHS format.
18. The apparatus according to claim 17, wherein Decoding the bitstream includes decoding a plurality of MIHS units and packets of the media stream.
19. The apparatus according to claim 11, Among them, In the case where the data is in the haptic exchange format, verifying at least one of the pattern against a reference pattern and the semantics against a rule table includes outputting a report indicating a result of verifying at least one of the pattern against a reference pattern and the semantics against a rule table.
20. A non-transitory computer-readable medium storing instructions, the instructions comprising: One or more instructions that, when executed by one or more processors of an apparatus for decoding video data, cause the one or more processors to: Receive a media stream, the media stream including data in one of a haptic exchange format and a haptic streaming format; In the case where the data is in the haptic exchange format: Obtain a pattern of an input document of the media stream and a semantics of the input document; And Verify at least one of the pattern by the reference pattern and the semantics by the control rule table; In the case where the data adopts the haptic streaming format: Decode at least a part of the bitstream from the media stream; And Verify at least one of the syntax and conditions of the bitstream; and Control the decoding of the media stream based on any one of the following: verify at least one of the syntax and conditions of the bitstream; and verify at least one of the pattern by the reference pattern and the semantics by the control rule table.