Method for embedding haptic effect semantics in haptic streaming
By employing syntax-based semantic description methods in haptic exchange and streaming formats, the timing model for haptic tracks is clarified, allowing efficient and flexible haptic data transmission in multimedia presentations.
Patent Information
- Application Number
- CN202480005368.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-16
- Filing Date
- 2024-04-17
- Publication Date
- 2025-07-15
AI Technical Summary
In the prior art, the transmission timing model of the tactile track is not clear, the timing relationship between the timing of the ISOBMFF track and the tactile basic signal is unclear, and the semantic dictionary has not been added to the standard, and the flexible semantic description method is not supported, resulting in the addition of semantics in the binary format may affect compactness.
By introducing semantic description schemes in tactile exchange formats and tactile streaming formats, using MPEG_haptics.perception object and readMetadataPerception() syntax to indicate semantic descriptions of effects, adding scheme identifiers and semantic keyword codes, to achieve flexible and compact encoding and decoding of tactile effects.
The haptic effect is synchronized with other media tracks in the ISOBMFF file, improving the efficiency and flexibility of media track processing, and ensuring the compactness and scalability of semantic descriptions in binary formats.
Smart Images

Figure BDA0005438386340000161 
Figure BDA0005438386340000171 
Figure BDA0005438386340000172
Abstract
Description
Cross - Reference to Related Applications
[0001] This application claims priority to U.S. Provisional Application No. 63 / 459,947, filed on April 17, 2023, U.S. Provisional Application No. 63 / 459,945, also filed on April 17, 2023, and U.S. Application No. 18 / 636,607, filed on April 16, 2024. The disclosures of the U.S. Provisional Applications and the U.S. Application are hereby incorporated by reference in their entireties. Technical Field
[0002] This disclosure relates to a set of advanced video coding and decoding techniques. More specifically, this disclosure relates to encoding and decoding haptic experiences for multimedia presentation and to methods for transmitting binary wavelet streams in a haptic interchange format. Background Art
[0003] Haptic experiences have become part of multimedia presentations. In applications where multimedia presentations include aspects of haptic experiences, haptic signals can be transmitted to a device or a wearable device, and a user can feel haptic sensations that are coordinated with visual and / or audio media experiences during the use of the application.
[0004] Recognizing the increasing popularity of haptic experiences in multimedia presentations, the Moving Picture Experts Group (MPEG) has started researching compression standards for haptics (for both MPEG - DASH and MPEG - I) and transmitting compressed haptic signaling in an ISO (International Organization for Standardization, IOS) - based media file format (ISO Based Media File Format, ISOBMFF).
[0005] One of the problems to be solved in aspects of multimedia presentations that involve haptic experiences is that the timing model for the transmission of haptic tracks is not clear, i.e., it is not clear how the timing of an ISOBMFF track relates to the timing of haptic elementary signals. A solution to this problem is needed.
[0006] Since the haptic Committee Draft includes a JSON (JavaScript Object Notation, JSON) and a binary format. Recently, there has been discussion about adding semantic descriptions of effects in these two formats. However, semantic dictionaries have not been added to the standard, and flexible ways to signal different semantic dictionaries or extended semantic dictionaries have not been supported. Additionally, while it may be easy to add semantics to the exchange format, it is a challenge to add a large number of keywords to the binary format while maintaining compactness while achieving flexibility. A solution to address these issues is needed. Summary of the Invention
[0007] According to one aspect of the present disclosure, there is provided an apparatus, and similarly a method and a computer-readable medium. The apparatus includes at least one memory configured to store computer program code; and at least one processor configured to access the computer program code and operate as indicated by the computer program code. The computer program code includes: a receiving code configured to cause the at least one processor to receive a media stream including data in one of a haptic exchange format and a haptic streaming format; an obtaining code configured to cause the at least one processor to, when the data is in the haptic exchange format, obtain a first syntax from the media stream and determine a scheme for semantic description of effects for the media stream according to the first syntax; and when the data is in the haptic streaming format, obtain a second syntax from the media stream and determine a scheme for semantic description of effects for the media stream and the number of characters in the scheme; and a control code configured to cause the at least one processor to control decoding of the media stream based on the determined scheme.
[0008] The first syntax may indicate "MPEG_haptics.perception object".
[0009] When the data is in the haptic exchange format, the scheme may be indicated by the "effect_semantic" syntax in the "MPEG_haptics.perception object".
[0010] The second syntax may indicate "readMetadataPerception()".
[0011] In the case where the data is in the haptic streaming format, the "readMetadataPerception()" can indicate the scheme through the "effectSemantic" syntax, and indicate the number of characters in the scheme through the "schemeLength" syntax.
[0012] "readMetadataPerception()" can also indicate that at least one character in the effect semantic scheme identifier is in the form of a uniform resource name (URN).
[0013] In the case where the data is in the haptic streaming format, the media stream can indicate whether the effect has any semantic keywords and indicate how many keywords are included in the effect.
[0014] In the case where the data is in the haptic streaming format, the media stream can indicate whether the effect has any semantic keywords through the "hasSemantics" flag.
[0015] In the case where the data is in the haptic streaming format, the media stream can indicate how many keywords are included in the effect through the "semanticNumber" flag.
[0016] In the case where the data is in the haptic streaming format, the media stream can also indicate the semantic keyword code according to the semantic code table through the "effectSemantic" flag.
[0017] Additional embodiments will be set forth in the following description, and will be apparent in part from the description, and / or may be learned by practice of the presented embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Additional features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0019] Figure 1 is a schematic illustration of a simplified block diagram of a communication system according to an embodiment of the present disclosure;
[0020] Figure 2 is a schematic illustration of a simplified block diagram of a streaming system according to an embodiment of the present disclosure.
[0021] Figure 3 is an example illustration according to an embodiment of the present disclosure;
[0022] Figure 4 is an example illustration according to an embodiment of the present disclosure;
[0023] Figure 5A is an exemplary illustration according to an embodiment of the present disclosure;
[0024] Figure 5B is an exemplary illustration according to an embodiment of the present disclosure;
[0025] Figure 6 is an exemplary flowchart showing processing for handling haptic media according to an embodiment of the present disclosure;
[0026] Figure 7 is an exemplary diagram showing aspects according to an embodiment of the present disclosure;
[0027] Figure 8 is an exemplary diagram showing aspects according to an embodiment of the present disclosure;
[0028] Figure 9 is an exemplary diagram showing aspects according to an embodiment of the present disclosure; and
[0029] Figure 10 is an exemplary diagram showing aspects according to an embodiment of the present disclosure. Detailed Description
[0030] According to one aspect of the present disclosure, there are provided a method, a system, and a non-transitory storage medium for parallel processing of dynamic mesh compression. Embodiments of the present disclosure may also be applied to static meshes.
[0031] Referring to Figure 1 and Figure 2 embodiments of the present disclosure for implementing an encoding structure and a decoding structure of the present disclosure are described.
[0032] Figure 1 shows a simplified block diagram of a communication system 100 according to an embodiment of the present disclosure. The system 100 may include at least two terminals 110, 120 interconnected via a network 150. For unidirectional data transmission, a first terminal 110 may encode video data that may include mesh data at a local location for transmission to another terminal 120 via the network 150. A second terminal 120 may receive the encoded video data of another terminal from the network 150, decode the encoded data, and display the restored video data. Unidirectional data transmission may be common in media service applications and the like.
[0033] Figure 1A second pair of terminals 130, 140 is shown, and the second pair of terminals 130, 140 are provided to support two-way transmission of encoded video that may occur, for example, during a video conference. For two-way transmission of data, each of the terminals 130, 140 may encode video data captured at a local location for transmission via the network 150 to the other terminal. Each of the terminals 130, 140 may also receive encoded video data transmitted by the other terminal, may decode the encoded data, and may display the recovered video data at a local display device.
[0034] In Figure 1 , the terminals 110 to 140 may be, for example, servers, personal computers, and smart phones and / or any other type of terminal. For example, the terminals (110 to 140) may be laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network 150 represents any number of networks that convey encoded video data among the terminals 110 to 140, including, for example, wired communication networks and / or wireless communication networks. The communication network 150 may exchange data in circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, unless otherwise stated below, the architecture and topology of the network 150 may be immaterial to the operation of the present disclosure.
[0035] Figure 2 Placement of video encoders and decoders in a streaming environment is shown as an example of an application of the disclosed subject matter. The disclosed subject matter may be used with other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CD (Compact Disc), DVD (Digital Versatile Disc), memory sticks, and the like.
[0036] As Figure 2 shown, the streaming system 200 may include a capture subsystem 213 that includes a video source 201 and an encoder 203. The streaming system 200 may also include at least one streaming server 205 and / or at least one streaming client 206.
[0037] Video source 201 may create a stream 202 that includes, for example, a 3D (Three Dimensional) mesh and metadata associated with the 3D mesh. Video source 201 may include, for example, a 3D sensor (e.g., a depth sensor) or 3D imaging technology (e.g., a digital camera device), and a computing device configured to generate a 3D mesh using data received from the 3D sensor or 3D imaging technology. Sample stream 202, which may have a high data volume when compared to an encoded video bitstream, may be processed by an encoder 203 coupled to video source 201. Encoder 203 may include hardware, software, or a combination thereof to implement or carry out aspects of the disclosed subject matter described in more detail below. Encoder 203 may also generate an encoded video bitstream 204. Encoded video bitstream 204, which may have a lower data volume when compared to uncompressed stream 202, may be stored on a streaming server 205 for future use. One or more streaming clients 206 may access streaming server 205 to retrieve a video bitstream 209 that may be a copy of encoded video bitstream 204.
[0038] Streaming client 206 may include a video decoder 210 and a display 212. Video decoder 210 may, for example, decode video bitstream 209, which is an incoming copy of encoded video bitstream 204, and create an outgoing video sample stream 211 that may be rendered on display 212 or another rendering device (not depicted). In some streaming systems, video bitstreams 204, 209 may be encoded according to certain video coding / compression standards.
[0039] Figure 3 May be a functional block diagram of a video decoder 300 according to an embodiment of the present invention.
[0040] The receiver 302 may receive one or more codec video sequences to be decoded by the decoder 300; in the same or another embodiment, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequences may be received from the channel 301, which may be a hardware / software link to a storage device storing the encoded video data. The receiver 302 may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective consuming entities (not depicted). The receiver 302 may separate the encoded video sequences from the other data. To prevent network jitter, a buffer memory 303 may be coupled between the receiver 302 and the entropy decoder / parser 304 (hereinafter referred to as the "parser"). When the receiver 302 is receiving data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer 303 may not be required or the buffer 303 may be small. For use on a best-effort packet network such as the Internet, the buffer 303 may be required, and the buffer 303 may be relatively large and may advantageously have an adaptive size.
[0041] The video decoder 300 may include a parser 304 to reconstruct symbols 313 from the entropy-encoded video sequence. The categories of these symbols include: information for managing the operation of the decoder 300; and information potentially for controlling a rendering device such as the display 312, which is not part of the decoder but may be coupled to the decoder. The control information for the rendering device may be in the form of supplementary enhancement information (SEI (Supplementary Enhancement Information, SEI) messages) or video usability information parameter set fragments (not depicted). The parser 304 may parse / entropy-decode the received encoded video sequence. The encoding of the encoded video sequence may be according to a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 304 may extract subgroup parameter sets from the encoded video sequence for at least one subgroup among at least one subgroup of pixels in the video decoder based on at least one parameter corresponding to the subgroup. The subgroups may include group of pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The entropy decoder / parser may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.
[0042] The parser 304 can perform entropy decoding / parsing operations on the video sequence received from the buffer 303 to create symbols 313. The parser 304 can receive the encoded data and selectively decode specific symbols 313. In addition, the parser 304 can determine whether to provide a specific symbol 313 to the motion compensation prediction unit 306, the scaler / inverse transform unit 305, the intra prediction unit 307, or the loop filter 311.
[0043] Depending on the type of the encoded video picture or a part thereof (e.g., inter picture and intra picture, inter block and intra block) and other factors, the reconstruction of the symbol 313 may involve multiple different units. Which units are involved and the way they are involved can be controlled by subgroup control information parsed by the parser 304 from the encoded video sequence. For clarity, this subgroup control information flow between the parser 304 and the multiple units below is not depicted.
[0044] In addition to the functional blocks already mentioned, the decoder 300 can conceptually be subdivided into multiple functional units as described below. In actual implementations operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the following functional units.
[0045] The first unit is the scaler / inverse transform unit 305. The scaler / inverse transform unit 305 receives the quantized transform coefficients as symbols 313 and control information from the parser 304, including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit 305 can output a block including sample values, and the sample values can be input into the aggregator 310.
[0046] In some cases, the output samples of the scaler / inverse transform 305 can belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture but can use predictive information from a previously reconstructed part of the current picture. Such predictive information can be provided by the intra picture prediction unit 307. In some cases, the intra picture prediction unit 307 uses the already reconstructed surrounding information extracted from the current (partially reconstructed) picture 309 to generate a block having the same size and shape as the size and shape of the block being reconstructed. In some cases, the aggregator 310 adds the prediction information already generated by the intra prediction unit 307 to the output sample information provided by the scaler / inverse transform unit 305 based on each sample.
[0047] In other cases, the output samples of the scaler / inverse transform unit 305 may belong to a block that has been inter-coded and possibly motion-compensated. In such cases, the motion compensation prediction unit 306 may access the reference picture memory 308 to extract samples for prediction. After motion-compensating the extracted samples according to the symbols 313 belonging to the block, these samples may be added by the aggregator 310 to the output of the scaler / inverse transform unit (referred to as residual samples or residual signals in this case) to generate output sample information. The address in the reference picture memory from which the motion compensation unit obtains the prediction samples may be controlled by a motion vector, which is available to the motion compensation unit in the form of the symbols 313, and the symbols 313 may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include interpolation of sample values obtained from the reference picture memory when using sub-sample accurate motion vectors, motion vector prediction mechanisms, and the like.
[0048] The output samples of the aggregator 310 may undergo various loop filtering techniques in the loop filter unit 311. Video compression techniques may include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream and available to the loop filter unit 311 as symbols 313 from the parser 304, but video compression techniques may also respond to meta-information obtained during decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0049] The output of the loop filter unit 311 may be a sample stream that may be output to the rendering device 312 and stored in the reference picture memory 557 for future inter-picture prediction.
[0050] Certain encoded pictures may be used as reference pictures for future prediction once they have been fully reconstructed. Once an encoded picture has been fully reconstructed and the encoded picture has been identified as a reference picture (e.g., by the parser 304), the current reference picture 309 may become part of the reference picture buffer 308, and a new current picture memory may be reallocated before starting to reconstruct subsequent encoded pictures.
[0051] Video decoder 300 may perform decoding operations according to a predetermined video compression technique that may be recorded in a standard such as ITU-T (International Telecommunication Union - Telecommunication Standardization Sector) H.265 recommendation. In the sense that an encoded video sequence conforms to the syntax of the video compression technique or standard, the encoded video sequence may comply with the syntax specified by the video compression technique or standard being used, as specified in the video compression technique document or standard and explicitly in the profile document therein. For compliance, it is also required that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstructed sample rate (measured in, for example, megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level may be further restricted by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.
[0052] In an embodiment, receiver 302 may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the (one or more) encoded video sequences. The additional data may be used by video decoder 300 for proper decoding of the data and / or more accurate reconstruction of the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0053] Figure 4 May be a functional block diagram of a video encoder 400 according to an embodiment of the present disclosure.
[0054] Encoder 400 may receive video samples from video source 401 (which is not part of the encoder), and the video source 401 may capture video images to be encoded by encoder 400.
[0055] The video source 401 may provide a source video sequence in the form of a stream of digital video samples to be encoded by an encoder (303). The digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, …), any color space (e.g., BT.601 Y CrCb, RGB, …), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source 401 may be a storage device storing previously prepared video. In a video conferencing system, the video source 401 may be a camera device that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be organized as a spatial pixel array, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can readily understand the relationship between pixels and samples. The following description focuses on samples.
[0056] According to an embodiment, the encoder 400 may encode and compress pictures of the source video sequence into an encoded video sequence 410 in real time or according to any other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller 402. The controller controls the other functional units as described below and is functionally coupled to these units. For clarity, the couplings are not depicted. The parameters set by the controller may include: rate control related parameters (picture skip, quantizer, λ value of rate distortion optimization techniques, …), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can readily identify other functions of the controller 402 since these functions may be specific to the video encoder 400 optimized for a particular system design.
[0057] Some video encoders operate in a manner that is readily recognizable to those skilled in the art as a "coding / decoding loop". As a gross oversimplification, the coding / decoding loop can include: an encoding portion of encoder 402 (hereinafter "source encoder") (responsible for creating symbols based on the input picture and reference pictures to be encoded); and a (local) decoder 406 embedded in encoder 400, where the (local) decoder 406 reconstructs the symbols to create sample data that the (remote) decoder will also create (since in the video compression techniques contemplated in the disclosed subject matter, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream is input to reference picture memory 405. Since decoding the symbol stream results in a bit-exact result independent of the decoder location (local or remote), the reference picture buffer content is also bit-exact between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction portion of the encoder are exactly the same as the sample values that the decoder will "see" when using prediction during decoding. The rationale for this reference picture synchronization (and the resulting drift if synchronization cannot be maintained, e.g., due to channel errors) is well known to those skilled in the art.
[0058] The operation of the "local" decoder 406 can be the same as that of the "remote" decoder 300 already described in detail above. However, also briefly referring to Figure 3 , when symbols are available and the entropy encoder 408 and parser 304 can encode / decode the symbols losslessly into the encoded / decoded video sequence, the entropy decoding portion of decoder 300 including channel 301, receiver 302, buffer 303, and parser 304 may not be fully implemented in local decoder 406. Figure 4
[0059] What can be observed at this point is that any decoder technology other than the parsing / entropy decoding present in the decoder must also necessarily exist in the corresponding encoder in substantially the same functional form. Since the encoder technology is reciprocal to the decoder technology that has been fully described, the description of the encoder technology can be simplified. More detailed descriptions will only be needed and provided in certain places below.
[0060] As part of its operation, source encoder 403 can perform motion-compensated predictive coding that predicts the input frame by referring to one or more previously encoded frames in the video sequence designated as "reference frames". In this way, encoding engine 407 encodes the difference between a pixel block of the input frame and a pixel block of the reference frame, where the reference frame can be selected as the prediction reference for the input frame.
[0061] The local video decoder 406 can decode the encoded video data of the frames that can be designated as reference frames based on the symbols created by the source encoder 403.
[0062] The operation of the encoding engine 407 can advantageously be lossy processing. When the encoded video data can be decoded at a video decoder ( Figure 4 (not shown in the figure)), the reconstructed video sequence is generally a copy of the source video sequence with some errors. The local video decoder 406 replicates the decoding process that can be performed by the video decoder on the reference frames, and can store the reconstructed reference frames in the reference picture buffer 405. In this way, the encoder 400 can locally store a copy of the reconstructed reference frames, which has the same content (without transmission errors) as the reconstructed reference frames that will be obtained by the remote video decoder.
[0063] The predictor 404 can perform a prediction search for the encoding engine 407. That is, for a new frame to be encoded, the predictor 404 can search the reference picture memory 405 for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc. that can be used as an appropriate prediction reference for the new picture. The predictor 404 can operate block by block based on sample blocks to find an appropriate prediction reference. In some cases, as determined by the search results obtained by the predictor 404, the input picture can have prediction references extracted from multiple reference pictures stored in the reference picture memory 405.
[0064] The controller 402 can manage the encoding operations of the video encoder 403, including, for example, setting parameters and subgroup parameters for encoding the video data.
[0065] The outputs of all the above-mentioned functional units can undergo entropy encoding in the entropy encoder 408. The entropy encoder converts these symbols into an encoded video sequence by losslessly compressing the symbols generated by various functional units according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.
[0066] The transmitter 409 can buffer the encoded video sequence(s) created by the entropy encoder 408 in preparation for transmission via the communication channel 411, which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter 409 can merge the encoded video data from the video encoder 403 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0067] The controller 402 may manage the operation of the encoder 400. During encoding, the controller 405 may assign a certain coded picture type to each coded picture, which may affect the encoding techniques that may be applied to the corresponding picture. For example, pictures may typically be assigned to one of the following frame types:
[0068] An intra picture (I picture), which may be a picture that can be encoded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example independent decoder refresh pictures. Those skilled in the art are aware of those variations of I pictures and their corresponding applications and characteristics.
[0069] A predictive picture (P picture), which may be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.
[0070] A bi - predictive picture (B picture), which may be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indexes to predict the sample values of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0071] Source pictures may typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples respectively), and encoded on a block - by - block basis. These blocks may be predictively encoded with reference to other (coded) blocks, which are determined by the encoding assignment applied to the corresponding picture of the block. For example, blocks of an I picture may be non - predictively encoded, or may be predictively encoded (spatial prediction or intra prediction) with reference to coded blocks of the same picture. Pixel blocks of a P picture may be non - predictively encoded with reference to one previously coded reference picture via spatial prediction or via temporal prediction. Blocks of a B picture may be non - predictively encoded with reference to one or two previously coded reference pictures via spatial prediction or via temporal prediction.
[0072] The video encoder 400 may perform encoding operations according to a predetermined video encoding technique or standard such as the ITU - T H.265 recommendation. In the operation of the video encoder 400, the video encoder 400 may perform various compression operations, including predictive encoding operations that exploit the temporal redundancy and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard being used.
[0073] In an embodiment, the transmitter 409 may transmit additional data along with the encoded video. The source encoder 403 may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, visual usability information (VUI) parameter set segments, etc.
[0074] Referring Figures 5A to 5B to, embodiments of the present disclosure for implementing a haptic encoder 500 and a haptic decoder 550 are described.
[0075] As Figure 5A shown, the haptic encoder 500 may receive both descriptive haptic data and waveform haptic data. Thus, the haptic encoder 500 may be capable of processing three types of input files:.ohm metadata files (Object Haptic Metadata - a text file format for haptic metadata), descriptive haptic files (.ivs,.ahap, and.hjif), or waveform PCM (Pulse Code Modulation, PCM) files (.wav). Examples of descriptive data may include:.ahap from Apple (Apple Haptic and Audio Pattern - a JSON-like file format that specifies haptic patterns) (representing the expected haptic output through a parameterized set of modulated transients and a set of modulated continuous signals),.ivs from Immersion (representing the expected haptic output through a set of basic effects parameterized by a set of parameters), or.hjif (Haptics JSON Interchange Format), the proposed MPEG format. Examples of waveform pulse code modulation (PCM) signals may include:.ohm input files that include metadata information.
[0076] According to an embodiment, the haptic encoder 500 may process two types of input files differently. For descriptive content, the haptic encoder 500 may semantically analyze the input to transcoding (if necessary) the data into the proposed encoded representation.
[0077] According to an embodiment, the.ohm metadata input file may include a description of the tactile system and settings. In particular, the.ohm metadata input file may include the name of each associated tactile file (descriptive or PCM) and a description of the signal. The.ohm metadata input file also provides a mapping between each channel of the signal and a target body part on the user's body. For the.ohm metadata input file, the tactile encoder performs metadata extraction by: retrieving the associated tactile file from a URI (Uniform Resource Identifier, URI) and encoding the associated tactile file based on the type of the associated tactile file, and extracting metadata from the.ohm file and mapping the metadata to the metadata information of the data model.
[0078] According to an embodiment, descriptive tactile files (e.g.,.ivs,.ahap, and.hjif) may be encoded by simple processing. The tactile encoder 500 first explicitly identifies the input format. If the input format is a.hjif file, no transcoding is required, and the file can be further edited, compressed into a binary format, and finally grouped into an MIHS (MPEG Immersive Haptic Stream, MIHS) stream. If an.ahap or.ivs input file is used, transcoding is required. The tactile encoder 500 first semantically analyzes the input file information and transcodes it to format it into the selected data model. After transcoding, the data can be exported as a.hjif file,.hmpg binary file, or MIHS stream.
[0079] According to an embodiment, the tactile encoder 500 may perform signal analysis to account for the signal structure of a.wav file and convert it into the proposed encoded representation. For waveform PCM content, the signal analysis process may be divided into two sub-processes by the tactile encoder 500. After performing band decomposition on the signal, at the first sub-process, the low frequencies may be encoded using a keyframe extraction process. Then the low frequency band may be reconstructed, and the error between the signal and the original low frequency signal may be calculated. Then, the residual signal may be added to the original high frequency band, and then encoded using wavelet transform, which is the second sub-process. According to an embodiment, when several low frequency bands are used, the residuals from all low frequency bands are added to the high frequency band before encoding. In an embodiment, when several high frequency bands are used, the residuals from the low frequency band are added to the first high frequency band before encoding.
[0080] According to an embodiment, key frame extraction includes obtaining a lower frequency band from band decomposition and analyzing the content of the lower frequency band in the time domain. According to an embodiment, wavelet processing may include obtaining a high frequency band from band decomposition and a low frequency residue, and dividing the high frequency band into equally sized blocks. These equally sized signal blocks are then analyzed in a haptic model. With the assistance of the haptic model, lossy compression can be applied by performing wavelet transform on the blocks and quantizing them. Finally, each block is then saved as a separate effect in a single frequency band, which is done in formatting. Binary compression can use appropriate coding techniques, such as the Set Partitioning In Hierarchical Trees (SPIHT) algorithm and Arithmetic Coding (AC), to apply lossless compression.
[0081] As Figure 5A shown, the haptic encoder 500 may be configured to encode descriptive haptic data and quantized haptic data, and may output three types of formats - an interchange format (.hjif), a binary compressed format (.hmpg), and a streaming format (e.g., MPEG Immersive Haptic Stream (MIHS)). The.hjif format is a JSON-based human-readable format and can be easily parsed and manually edited, making it an ideal interchange format, especially when designing / creating content. For distribution purposes, the.hjif data can be compressed into a more memory-efficient binary.hmpg bitstream. This compression may be lossy, where different parameters affect the encoding depth of the amplitude and frequency that make up the bitstream. For streaming purposes, the data can be compressed and grouped into an MPEG-I Haptic Stream (MIHS). The three formats mentioned above have complementary purposes and can perform lossy one-to-one conversion between them.
[0082] As Figure 5B shown, the haptic decoder 550 may take a.hmpg compressed binary file format or an MIHS bitstream as input. The haptic decoder 550 may output an.hjif interchange format that can be directly used for rendering. Both input formats can be binary decompressed to extract both the metadata and the data itself from the file, and map the data to a selected data structure. Then, the data can be exported to the haptic renderer 580 in the.hjif format.
[0083] As Figure 5BAs shown, renderer 580 includes a synthesizer. The synthesizer can render haptic data from a.hjif input file into a PCM output file. The rendering and / or synthesis is informative. According to an embodiment, the synthesizer parses the input file and performs advanced synthesis distribution among vectors, wavelets, etc. Then, the synthesis process continues to the band components of the codec where the synthesis process is called. Then, all bands of a given channel are mixed by a simple adder to recreate the desired haptic signal.
[0084] According to an embodiment, the haptic experience defines the base of a hierarchical data model. The base provides information about the file date and format version, describes the haptic experience, lists the different visualizations (i.e., body representations) used throughout the experience, and defines all haptic perceptions.
[0085] According to an embodiment, a self - contained stream format for transmitting MPEG - I haptic data can use a packetization method and can include two levels of packets: an MPEG - I haptic stream (MIHS) unit that covers a duration and includes zero or more MIHS packets; and an MIHS packet that includes metadata or haptic effect data. In an embodiment, the MIHS unit can be referred to as a network abstraction layer unit associated with the haptic data. In an embodiment, the MIHS unit can be referred to as an MIHS sample associated with the haptic data.
[0086] According to an embodiment, an MIHS unit can be a synchronous unit or an asynchronous unit. The synchronous unit resets previous effects and thus provides a haptic experience independent of previous MIHS units. The asynchronous unit is a continuation of the previous MIHS unit and cannot be independently decoded and rendered without decoding the previous MIHS unit.
[0087] According to an embodiment, haptic signals can be encoded on multiple channels. In some embodiments, a haptic channel can define a signal to be rendered at a specific body location using a dedicated actuator / device. Metadata stored at the channel level can include information such as a gain associated with the channel, a mixing weight, a desired body location for haptic feedback, and optionally a reference device and / or orientation. Additional information such as a desired sampling frequency or sample count can also be provided. Finally, the haptic data for a channel is contained in a set of haptic frequency bands defined by its frequency range. A haptic frequency band describes the haptic signal for a channel within a given frequency range. The frequency band is defined by a list of types and orders of haptic effects, each haptic effect containing a set of key frames. For each type of haptic frequency band, a haptic effect can be defined by at least a position and a type. The position can indicate the temporal or spatial position of the effect. In some embodiments, the value 0 is the relative start position of the experience, which depends on the dependent variables of the configured perception modality. The default unit for temporal haptic feedback can be milliseconds, while the default unit for spatial haptic feedback can be millimeters. This embodiment discloses the "start position of the experience" because the binary distribution format does not have any concept of a finite time interval, i.e., frames or samples.
[0088] Depending on the type of the frequency band and the type of the effect, additional attributes can be specified, including a phase, a base signal, a composition and a number of consecutive haptic key frames describing the effect.
[0089] According to an embodiment, a haptic data hierarchy is defined in the present disclosure. ● Haptic channel ○ Haptic frequency band ■ Haptic effect
[0090] Embodiments of the present disclosure describe two anchors for the position of a haptic effect related to an ISOBMFF track.
[0091] Figure 6 A first embodiment 600 is shown. As Figure 6 shown, each MIHS unit (also referred to as a MIHS sample, an ISOBMFF haptic sample, or a sample in an embodiment) includes one or more haptic channel information and one or more haptic frequency band information. As described above, each MIHS unit includes one or more channels, and each channel includes one or more frequency bands. Then, each frequency band can have one or more effects.
[0092] In the first embodiment, the temporal position of an effect can be defined as an offset relative to the start timing of the sample carrying the effect (e.g., the MIHS unit start time). In a second or the same embodiment, the offset is based on the start time and / or the presentation time of the media or the haptic track.
[0093] According to an embodiment, the first embodiment can manipulate the track without affecting the position of the haptic effect because any change in the ISOBMFF sample timing will not affect the relative position of the effect. According to an embodiment, in the case of a basic haptic stream (e.g., an advanced syntax stream), the second embodiment can be used when the basic stream is used without ISOBMFF.
[0094] According to an embodiment, several types of haptic tracks can be used. In an embodiment, samples or MIHS units with a time position of an effect defined as an offset relative to the start timing of the sample can be used in the haptic track. According to another embodiment, for example Figure 7 in Example 700, samples or MIHS units that relate the time position of their effect to the start time of the track are used in the haptic track. In another embodiment, a mixed MIHS unit or sample can be used.
[0095] Embodiments of the present disclosure provide a timing model that can be used to synchronize haptic effects with other media tracks in the same or related ISOBMFF file. Since the timing model of the haptic track is related to the timing model of the related ISOBMFF file, the manipulation and processing of media tracks become more efficient.
[0096] As Figure 8 shown, process 800 illustrates an exemplary process for decoding haptic data.
[0097] At operation 805, a media stream including one or more haptic tracks and one or more video tracks can be received.
[0098] At operation 810, one or more Moving Picture Experts Group (MPEG) Immersive Haptic Stream (MIHS) units can be obtained from the media stream. In some embodiments, the MIHS unit can include one or more haptic effects. The MIHS unit can also include the start time of the MIHS unit.
[0099] In an embodiment, the MIHS unit is associated with at least one haptic channel, the at least one haptic channel includes one or more haptic bands, and each of the one or more haptic bands has at least one haptic effect.
[0100] At operation 815, timing information associated with one or more haptic effects can be obtained. In an embodiment, the timing information can include at least one time position of one or more haptic effects.
[0101] In an embodiment, the temporal position of a haptic effect indicates the effect start time of the haptic effect, where the effect start time of the haptic effect is an offset based on the start time of the corresponding MIHS unit. The effect start time may indicate the start time of the haptic effect corresponding to the start time of the corresponding MIHS unit.
[0102] In an embodiment, the effect start time of a haptic effect is an absolute time based on the start time of at least one haptic track or at least one video track.
[0103] At operation 820, the media stream is rendered based on the obtained timing information.
[0104] According to an embodiment, manipulation of the order of one or more MIHS units does not affect at least one temporal position of one or more haptic effects because the one or more MIHS units correspond to one or more ISO-based media file format (ISOBMFF) samples associated with at least one video track.
[0105] In some embodiments, synchronized MIHS units may be obtained from the media stream. In an embodiment, a synchronized MIHS unit is a special type of MIHS unit configured to provide a reset point in the bitstream. In an embodiment, the synchronized MIHS unit is mapped to a synchronized sample in the video bitstream corresponding to one or more haptic channels.
[0106] As in Figure 9 Example 900, according to embodiments herein, a haptic encoder generates a compact and efficient binary distribution format (.hmpg) for distribution. A haptic decoder may decode such a format and send it to a renderer.
[0107] The ISOBMFF haptic binding working draft defines the following for haptic samples in an ISOBMFF track: As indicated above, the data_packet_type is set to a fixed value.
[0108] An ISOBMFF haptic sample according to embodiments herein may have the following data packets: 1. A silent packet without data, i.e., a zero data packet. 2. One or more temporal data packets. 3. One or more spatial data packets. 4. A combination of 2 and 3 5. A possible combination of 2 and 3, but without any guarantee, i.e., zero or more temporal and / or data packets.
[0109] Embodiments herein can use data_packet_type to signal the above situations such that data_packet_type (i.e., b3b2b1b0) can be signaled as follows: Table 1 - Data Packet / Sample Type Wherein, according to an embodiment, data_packet_type = 0 indicates that the sample is a silent sample, i.e., equivalent to one or more consecutive MIHS silent units.
[0110] Therefore, the following values for data_packet_type in Table 2 have the following meanings: Table 2
[0111] In addition, according to an embodiment, the file format parser can use data_packet_type in the following ways: 1. Identify silent samples and skip further parsing of the silent samples when processing the file, so as to randomly access and fast forward / rewind the file more quickly. 2. Identify which samples only have time data packets to extract the time effect. 3. Identify which samples only have space data packets to extract the space effect. 4. When decomposing an orbit into multiple orbits, it is easier to separate time samples and space samples. 5. When merging multiple orbits into a single orbit, determine the sample structure of the new orbit and whether the samples will merge the time packets and space packets or keep the time packets and space packets in separate samples.
[0112] Therefore, according to an embodiment, there is a method for signaling sample types in ISOBMFF haptic samples, where the samples are identified as silent samples, non - silent samples, non - silent samples with only time effects, non - silent samples with only space effects, non - silent samples with a combination of time and space effects. Among them, bit - based flags are used for signaling for different characteristics. The file format parsing can utilize the information provided by the sample type and browse the file more quickly because silent samples are skipped, or use the sample type information to extract only time information, only space information, or only silent information, or use the information for bitstream manipulation, single - orbit to multiple - orbit conversion, or multiple - orbit to single - orbit conversion.
[0113] As in Figure 9In Example 900, according to the embodiments herein, the haptic encoder generates a compact and efficient binary distribution format (.hmpg) for distribution. The haptic decoder can decode such a format and send it to the renderer. However, in the absence of the embodiments herein, on the other hand, the haptic exchange format (.hjif) does not have binary compression, and thus, the wavelet coefficients are stored in the.hjif instead of a compressed bitstream. In view of this, also according to the embodiments, the embodiments herein extend the band type of the.hjif format to include a new option: binary wavelet.
[0114] Since the haptic committee draft includes a JSON and a binary format. Recently, there has been discussion about adding a semantic description of effects in both formats. However, the semantic dictionary has not been added to the standard, and a flexible method for signaling different semantic dictionaries or extending the semantic dictionary has not been supported. A solution to address these issues is needed.
[0115] In view of this, the embodiments herein add a scheme identifier to the HJIF. For example, see Table 3 below which provides a description of the MPEG_haptics.perception object: Table 3 - Description of the MPEG_haptics.perception object
[0116] As shown in Table 3, the item effect_semantic signals the scheme used to describe the effect scheme. The default scheme is the one specified by MPEG haptics according to the embodiments.
[0117] In addition, an additional solution according to the embodiments herein involves adding a scheme identifier to the MIHS. For example, the embodiments herein add a semantic scheme identifier to the MIHS according to the syntax of readMetadataPerception(), as shown in Table 4 below: Table 4 - Syntax of readMetadataPerception() Table 5 - Syntax of readSemanticScheme()
[0118] As shown in Tables 4 and 5, the parameters effectSemantics, SchemeLength, and schemeCharater provide technical improvements. For example, effectSemantics is a flag for the presence of an effect scheme, rather than a flag defined by a specific document (if the flag is 0, the scheme of that document is used to signal effect semantics), and schemeLength indicates the number of characters in the effect semantics scheme identifier string, and schemeCharacter indicates the characters in the effect semantics scheme identifier in the form of a URN (uniform resource name). According to the embodiments herein, “schemeLength” can also be “schemeLenght”.
[0119] Thus, viewing at least Tables 3 to 5 herein provides a method for signaling the scheme used for effect semantics in the haptic exchange file format and the haptic streaming format, where the URN is used to signal the scheme using the items under the perception level, where in any case, the default scheme is the one defined by the MPEG haptic standard, and in other cases, the signaled scheme defines the semantic code and the corresponding binary code to be used in the haptic exchange file and the haptic streaming format.
[0120] In addition, the embodiments herein also provide single-layer keyword signaling. That is, a flat dictionary with various semantic levels is defined through the embodiments herein. However, according to the embodiments, all semantic keywords receive the same number of bits, such as 8 bits. Then, the MIHS stream signals whether the effect has any semantic keywords, and if the effect has semantic keywords, it signals how many keywords are included in the effect. See Table 6 as an example thereof: Table 6 - Syntax of readEffect()
[0121] According to Table 7, the parameter hasSemantics is a flag for signaling whether the effect has semantics: Table 7 - Values of hasSemantics
[0122] In addition, the parameter semanticNumber indicates the number of semantic keywords used for this effect, and this value should be greater than 2. In addition, the parameter effectSemantic indicates the semantic keyword code according to the semantic code table.
[0123] The embodiments herein also provide a two-layer keyword signal representation, such that there is a defined two-level semantic dictionary here, where layer 1 defines the semantics for a general class, and layer 2 defines more specific semantic keywords within that class. The design of such an embodiment is shown below in Table 8: Table 8 - Syntax of readEffect()
[0124] In Table 8, the parameter hasSemantics represents such a flag that signals whether this effect has semantics according to Table 9 below: Table 9 - Values of hasSemantics
[0125] In addition, according to Table 8, the parameter semanticNumber represents the number of semantic keywords used for this effect, and according to the embodiment, this value should be greater than 1. In addition, according to Table 8, the parameter effectSemanticLayer1 represents the semantic keyword layer 1 code according to the semantic code table. In addition, according to Table 8, each of effectSemanticLayer2, effectSemanticLayer21, and effectSemanticLayer 22 represents the semantic keyword layer 2 code according to the semantic code table. Or in addition, according to Table 8, each of effectSemanticLayer2, effectSemanticLayer21, and effectSemanticLayer21 or effectSemanticLayer 22 represents the semantic keyword layer 1 or semantic keyword layer 2 code according to the semantic code table.
[0126] To this end, semantic keywords for representing haptic effects in a signal in the MIHS binary format are provided herein. In one case, a simple flat signal representation of the keywords is provided, where four different options are explicitly signaled: 1) no semantics, 2) one semantic keyword, 3) two semantic keywords, and 4) three or more semantic keywords. And in a second case, a two-layer semantic keyword signal representation is provided, where four different options are explicitly signaled: signaling 1) no semantics, 2) one layer 1 and one layer 2 semantics, 3) one layer 1 and two layer 3 semantics, and 3) more than one layer 1 and / or two layer 2 semantics.
[0127] Those skilled in the art will understand that the techniques described herein can be implemented on both the encoder side and the decoder side. The techniques described above can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 10 FIG. shows a computer system 1000 suitable for implementing certain embodiments of the present disclosure.
[0128] The computer software can be encoded using any suitable machine code or computer language, which can be subject to mechanisms such as assembly, compilation, linking, etc. to create code including instructions that can be directly executed by a computer central processing unit (CPU), a graphics processing unit (GPU), etc. or executed through interpretation, microcode execution, etc.
[0129] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0130] Figure 10 The components shown in FIG. for the computer system 1000 are examples and are not intended to impose any limitations on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should also not be construed as having any dependencies or requirements related to any one or combination of the components shown in the non-limiting embodiments of the computer system 1000.
[0131] The computer system 1000 may include certain human-machine interface input devices. Such human-machine interface input devices may respond to inputs made by one or more human users through, for example, tactile inputs (e.g., keystrokes, swipes, data glove movements), audio inputs (e.g., voice, taps), visual inputs (e.g., gestures), and olfactory inputs (not depicted). The human-machine interface devices may also be used to capture certain media that are not necessarily directly related to conscious inputs made by humans, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still image cameras), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0132] The input human-machine interface devices may include one or more of the following (only one of each depicted): keyboard 1001, mouse 1002, touchpad 1003, touch screen 1010, data glove, joystick 1005, microphone 1006, scanner 1007, camera device 1008.
[0133] The computer system 1000 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile outputs, sounds, lights, and smells / tastes. Such human-machine interface output devices may include: tactile output devices (e.g., tactile feedback through the touch screen 1010, data glove, or joystick 1005, but there may also be tactile feedback devices that do not function as input devices). For example, such devices may be audio output devices (e.g., speaker 1009, headphones (not depicted)); visual output devices (e.g., screen 1010, including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which may be able to output two-dimensional visual outputs or more than three-dimensional outputs through means such as stereoscopic output; virtual reality glasses (not depicted); holographic displays and fog machines (not depicted)); and printers (not depicted).
[0134] The computer system 1000 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM (Read-Only Memory, ROM) / RW 1020 with media 1021 such as CD / DVDs, thumb drives 1022, removable hard disk drives or solid-state drives 1023, traditional magnetic media such as tapes and floppy disks (not depicted), and dedicated ROM / ASIC-based Devices such as Application Specific Integrated Circuit (ASIC) / PLD (Programmable Logic Device) include, for example, a security dongle (not depicted).
[0135] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not include a transmission medium, a carrier wave, or other transient signals.
[0136] Computer system 1000 may also include an interface to one or more communication networks. The network may be, for example, wireless, wired, optical. The network may also be local, wide area, metropolitan area, vehicular, and industrial, real-time, delay-tolerant, etc. Examples of networks include: local area networks such as Ethernet; wireless LAN (Local Area Network); cellular networks including GSM (Global System for Mobile Communication), 3G (the Third Generation), 4G (the Fourth Generation), 5G (the Fifth Generation), LTE (Long Term Evolution), etc.; TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; vehicular and industrial networks including CANBus (Controller Area Network Bus), etc. Certain networks typically require an external network interface adapter attached to certain common data ports or peripheral buses 1049 (such as, for example, the USB (Universal Serial Bus) port of computer system 1000); other networks are typically integrated into the core of computer system 1000 by attaching to a system bus as described below (e.g., to an Ethernet interface in a PC (Personal Computer) computer system or to a cellular network interface in a smart phone computer system). Using any of these networks, computer system 1000 can communicate with other entities. Such communication can be unidirectional receive-only (e.g., broadcast TV), unidirectional send-only (e.g., CANbus to certain CANbus devices), or bidirectional, e.g., to other computer systems using local or wide area digital networks. Such communication can include communication to a cloud computing environment 1055. As described above, certain protocols and protocol stacks may be used on each of these networks and network interfaces.
[0137] The above-mentioned human-machine interface device, human-accessible storage device, and network interface 1054 can be attached to the core 1040 of the computer system 1000.
[0138] The core 1040 can include one or more central processing units (CPUs) 1041, a graphics processing unit (GPU) 1042, a dedicated programmable processing unit in the form of a field programmable gate area (FPGA) 1043, a hardware accelerator 1044 for certain tasks, etc. These devices, as well as a read-only memory (ROM) 1045, a random access memory 1046, and an internal mass storage device such as an internal non-user-accessible hard disk drive, solid state drive (SSD), etc. 1047, can be connected via a system bus 1048. In some computer systems, the system bus 1048 can be accessed in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices can be attached directly or via a peripheral bus 1049 to the system bus 1048 of the core. The architecture of the peripheral bus includes PCI (Peripheral Component Interconnect), USB, etc. A graphics adapter 1050 can be included in the core 1040.
[0139] The CPU 1041, GPU 1042, FPGA 1043, and accelerator 1044 can execute certain instructions, which, when combined, can constitute the above-mentioned computer code. This computer code can be stored in the ROM 1045 or RAM (Random Access Memory) 1046. Transient data can also be stored in the RAM 1046, while permanent data can be stored, for example, in the internal mass storage device 1047. Fast storage and retrieval of any memory device in the memory devices can be achieved by using a cache memory, which can be closely associated with one or more CPUs 1041, GPUs 1042, mass storage device 1047, ROM 1045, RAM 1046, etc.
[0140] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be media and computer code specifically designed and constructed for the purposes of this disclosure, or the medium and the computer code can be of the type well-known and available to those skilled in the field of computer software.
[0141] By way of example and not limitation, a computer system having the architecture of computer system 1000 and in particular core 1040 can provide functionality due to a processor (including CPU, GPU, FPGA, accelerator, etc.) executing software contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with the user-accessible mass storage device as introduced above and certain storage devices of core 1040 having a non-transitory nature such as core internal mass storage device 1047 or ROM 1045. The software implementing the various embodiments of the present disclosure can be stored in such devices and executed by core 1040. Depending on specific needs, the computer-readable media can include one or more memory devices or chips. The software can cause core 1040 and in particular the processors therein (including CPU, GPU, FPGA, etc.) to perform the specific processes or specific portions of the specific processes described herein, including defining data structures stored in RAM 1046 and modifying such data structures according to the processes defined by the software. Additionally or alternatively, the computer system can provide functionality due to being logically hardwired or otherwise implemented in circuitry (e.g., accelerator 1044) that can operate in place of or in conjunction with the software to perform the specific processes or specific portions of the specific processes described herein. In appropriate cases, references to software can include logic, and references to logic can also include software. In appropriate cases, references to computer-readable media can include circuitry (e.g., integrated circuit (IC)) storing software for execution, circuitry implementing logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0142] Although the present disclosure has described several non-limiting embodiments, there are changes, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it will be recognized that those skilled in the art will be able to envision many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.
Claims
1. A method for decoding video data, the method being executed by at least one processor, the method comprising: Receiving a media stream including data in one of a haptic exchange format and a haptic streaming format; In a case where the data is in the haptic exchange format, obtaining a first syntax from the media stream and determining a scheme for a semantic description of an effect for the media stream according to the first syntax; In a case where the data is in the haptic streaming format, obtaining a second syntax from the media stream and determining a scheme for a semantic description of an effect for the media stream and the number of characters in the scheme according to the second syntax; And Controlling decoding of the media stream based on determining the scheme.
2. The method according to claim 1, wherein, The first syntax indicates "MPEG_haptics.perceptionobject”.
3. The method according to claim 2, wherein In a case where the data is in the haptic exchange format, the scheme is indicated by an "effect_semantic” syntax in "MPEG_haptics.perception object”.
4. The method according to claim 1, wherein The second syntax indicates "readMetadataPerception()”.
5. The method according to claim 4, wherein In a case where the data is in the haptic streaming format, the scheme is indicated by an "effectSemantic” syntax in "readMetadataPerception()”, and the number of characters in the scheme is indicated by a "schemeLength” syntax in "readMetadataPerception()”.
6. The method according to claim 5, wherein "readMetadataPerception()” also indicates that at least one character in an effect semantic scheme identifier is in the form of a Uniform Resource Name (URN).
7. The method according to claim 1, wherein In a case where the data is in the haptic streaming format, the media stream indicates whether an effect has any semantic keywords and indicates how many keywords are included in the effect.
8. The method according to claim 7, wherein In a case where the data is in the haptic streaming format, the media stream indicates whether an effect has any semantic keywords by a "hasSemantics” flag.
9. The method according to claim 8, wherein, In a case where the data is in the haptic streaming format, the media stream indicates how many keywords are included in the effect by a "semanticNumber” flag.
10. The method according to claim 9, wherein, In a case where the data is in the haptic streaming format, the media stream also indicates semantic keyword codes according to a semantic code table by an "effectSemantic” flag.
11. A device for decoding video data, the device comprising: At least one memory configured to store program code; And At least one processor configured to read the program code and operate as indicated by the program code, the program code including: Receiving code that is configured to cause the at least one processor to receive a media stream that includes data in one of a haptic exchange format and a haptic streaming format; Obtaining code that is configured to cause the at least one processor to: In the case where the data is in the haptic exchange format, obtain a first syntax from the media stream and determine a scheme for a semantic description of an effect for the media stream according to the first syntax; and In the case where the data is in the haptic streaming format, obtain a second syntax from the media stream and determine a scheme for a semantic description of an effect for the media stream and the number of characters in the scheme according to the second syntax; and Controlling code that is configured to cause the at least one processor to control decoding of the media stream based on determining the scheme.
12. The apparatus according to claim 11, wherein, The first syntax indicates "MPEG_haptics.perceptionobject".
13. The apparatus according to claim 12, wherein, In the case where the data is in the haptic exchange format, the scheme is indicated by the "effect_semantic" syntax in "MPEG_haptics.perception object".
14. The apparatus according to claim 11, wherein, The second syntax indicates "readMetadataPerception()".
15. The apparatus according to claim 14, wherein, In the case where the data is in the haptic streaming format, the scheme is indicated by the "effectSemantic" syntax in "readMetadataPerception()", and the number of characters in the scheme is indicated by the "schemeLength" syntax in "readMetadataPerception()".
16. The device according to claim 15, wherein, "readMetadataPerception()" also indicates that at least one character in the effect semantic scheme identifier is in the form of a Uniform Resource Name (URN).
17. The apparatus according to claim 11, wherein In the case where the data is in the haptic streaming format, the media stream indicates whether the effect has any semantic keywords and indicates how many keywords are included in the effect.
18. The device according to claim 17, wherein, In the case where the data is in the haptic streaming format, the media stream indicates whether the effect has any semantic keywords by the "hasSemantics" flag.
19. The apparatus according to claim 18, wherein, In the case where the data is in the haptic streaming format, the media stream indicates how many keywords are included in the effect by the "semanticNumber" flag.
20. A non-transitory computer-readable medium storing instructions, the instructions comprising: One or more instructions that, when executed by one or more processors of a device for decoding video data, cause the one or more processors to: Receive a media stream that includes data in one of a haptic exchange format and a haptic streaming format; In the case where the data is in the haptic exchange format, obtain a first syntax from the media stream and determine a scheme for a semantic description of an effect for the media stream; In the case where the data is in the tactile streaming format, a second syntax is obtained from the media stream, and a scheme for the semantic description of the effect for the media stream and the number of characters in the scheme are determined according to the second syntax; and Based on the determined scheme, the decoding of the media stream is controlled.