Storing 3D object model data using iso bmff files

By storing multiple levels of detail (LOD) of 3D object models in ISO BMFF files, the model details are dynamically adjusted to adapt to network and processing capabilities, solving the problem of efficient storage and transmission of 3D object models in XR sessions, and improving resource utilization and user experience.

CN122460083APending Publication Date: 2026-07-24QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QUALCOMM INC
Filing Date
2025-01-03
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently store and transmit 3D object models in extended reality (XR) sessions, especially when distributed and animated across devices with varying network connectivity and processing capabilities, leading to resource waste and a degraded user experience.

Method used

The system uses ISO Basic Media File Format (ISO BMFF) files to store multiple Levels of Detail (LOD) of 3D object models. It dynamically retrieves appropriate LOD models based on distance and network bandwidth, and animates them through an animation stream, prioritizing high LOD models that are closer to the user.

Benefits of technology

It improves the efficiency of 3D object model storage, transmission and rendering, reduces processing and bandwidth requirements, and maintains a good user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122460083A_ABST
    Figure CN122460083A_ABST
Patent Text Reader

Abstract

An example device for retrieving media data includes a memory to store media data and a processing system implemented in circuitry and configured to retrieve data representing a three-dimensional (3D) object model and one or more levels of detail (LODs) of the 3D object model, transmit a request to a server device to access data of one of the LODs of the 3D object model, the data of one of the LODs of the 3D object model including a size, a complexity, and components of the 3D object model for one of the LODs, and receive, in response to the request, the data of one of the LODs of the 3D object model, the data of one of the LODs of the 3D object model having the size, the complexity, and the components of the 3D object model for one of the LODs.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to U.S. Patent Application No. 19 / 008,368, filed January 2, 2025, and U.S. Provisional Application No. 63 / 617,715, filed January 4, 2024, the entire contents of each of which are incorporated herein by reference. U.S. Patent Application No. 19 / 008,368, filed January 2, 2025, claims the benefit of U.S. Provisional Application No. 63 / 617,715, filed January 4, 2024. Technical Field

[0002] This disclosure relates to the storage and transmission of encoded video data. Background Technology

[0003] Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, digital direct broadcasting systems, wireless broadcasting systems, personal digital assistants (PDAs), laptops or desktop computers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite broadcast phones, video conferencing equipment, and more. Digital video devices implement video compression technologies (such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263 or ITU-T H.264 / MPEG-4, Part 10, Advanced Video Decoding (AVC), ITU-T H.265 (also known as Advanced Video Decoding (HEVC)), and extensions to these standards) to send and receive digital video information more efficiently.

[0004] Video compression techniques perform spatial and / or temporal prediction to reduce or remove inherent redundancy in video sequences. For block-based video decoding, video frames or slices can be divided into macroblocks. Each macroblock can be further subdivided. Macroblocks in intra-frame decoding (I) frames or slices can be encoded using spatial prediction of adjacent macroblocks. Macroblocks in inter-frame decoding (P or B) frames or slices can be encoded using spatial prediction of adjacent macroblocks in the same frame or slice, or temporal prediction of other reference frames.

[0005] After video data is encoded, it can be packaged for transmission or storage. Video data can be assembled into video files that conform to any of various standards, such as the International Organization for Standardization (ISO) Basic Media File Format (BMFF) and its extensions, such as AVC. Summary of the Invention

[0006] Generally, this disclosure describes techniques for storing 3D object models (such as user avatars) involved in an extended reality (XR) session using ISO Basic Media File Format (ISO BMFF) files. For example, various levels of detail (LODs) of the 3D object model can be stored, including storing the model's components for each LOD. In this way, the basic model can be distributed to users involved in the XR session and then modified using animation streams. That is, when a user of a device (e.g., a user equipment (UE) device) participates in an XR session with other users, that user's device can share the basic 3D object model with other users and then transmit animation streams to animate that 3D object model.

[0007] This disclosure relates to techniques for efficiently storing base models for rapid distribution and adaptation to the network connectivity and processing capabilities of each participant. Specifically, a device can determine the specific Level of Detail (LOD) of a 3D object model to be retrieved based on factors such as distance to 3D objects in a 3D scene and available network bandwidth, and then retrieve the determined LOD for each 3D object model. Thus, when a 3D object is relatively far from the current user, a relatively low LOD model can be retrieved, while when a 3D object is relatively close to the current user, a relatively high LOD model can be retrieved. Therefore, for objects that do not require a large amount of detail (e.g., due to their distance from the current user), relatively low LOD models can be accessed and processed quickly, and thus processing resources can be prioritized for relatively high LOD models of 3D objects closer to the current user. Therefore, these techniques can reduce the processing operations and bandwidth required to render 3D objects while maintaining a good user experience, as nearby 3D objects can be rendered at a higher quality level.

[0008] In one example, a method for storing media data includes storing a three-dimensional (3D) object model in an ISO Basic Media File Format (ISO BMFF) file, including: storing metadata of the 3D object model in the ISO BMFF file; storing several Levels of Detail (LODs) of the 3D object model in the ISO BMFF file; and for each of the several LODs, storing data in the ISO BMFF file that associates the LOD with data representing the size, complexity, and components of the 3D object model with respect to that LOD.

[0009] In another example, a device for storing media data includes: a memory for storing media data; and a processing system implemented in a circuit and configured to store a three-dimensional (3D) object model in the memory in an ISO Basic Media File Format (ISO BMFF) file of the media data, wherein, in order to store the 3D object model in the ISO BMFF file, the processing system is configured to: store metadata of the 3D object model in the ISO BMFF file; store several levels of detail (LODs) of the 3D object model in the ISO BMFF file; and for each of the several LODs, store data in the ISO BMFF file that associates the LOD with data representing the size, complexity, and components of the 3D object model with respect to the LOD.

[0010] In another example, a method for retrieving media data includes: retrieving, via a client device, data representing a three-dimensional (3D) object model stored in an ISO Basic Media File Format (ISO BMFF) file and one or more Levels of Detail (LODs) of the 3D object model stored in the ISO BMFF file, wherein the ISO BMFF file is stored on a server device; transmitting a request via the client device to the server device to access data of the 3D object model in one LOD of the LOD, the data of the 3D object model in that LOD including the size, complexity, and components of the 3D object model for that LOD; and, in response to the request, receiving, via the client device, the data of the 3D object model in that LOD, the data of the 3D object model in that LOD having the size, complexity, and components of the 3D object model for that LOD.

[0011] In another example, an apparatus for retrieving media data includes: a memory for storing media data; and a processing system implemented in a circuit and configured to: retrieve data representing a three-dimensional (3D) object model and one or more levels of detail (LODs) of the 3D object model; transmit a request to the server device to access data of the 3D object model in one LOD of the LODs, the data of the 3D object model in the one LOD of the LOD including the size, complexity, and components of the 3D object model for the one LOD of the LOD; and, in response to the request, receive the data of the 3D object model in the one LOD of the LOD, the data of the 3D object model in the one LOD of the LOD having the size, complexity, and components of the 3D object model for the one LOD of the LOD.

[0012] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from these descriptions and drawings, and from the claims. Attached Figure Description

[0013] Figure 1 This is a block diagram illustrating an example system for implementing a technology for streaming media data over a network.

[0014] Figure 2 This is a block diagram illustrating the elements of the example video file.

[0015] Figure 3 This is a flowchart illustrating a sample avatar animation workflow that can be used during an XR session.

[0016] Figure 4 This is a flowchart illustrating an example XR session between two user equipment (UE) devices and a shared space server device.

[0017] Figure 5 This is a block diagram illustrating an example 3D object model structure that can be stored in an ISO Basic Media File Format (ISO BMFF) file according to the technology of this disclosure.

[0018] Figure 6 This is an example of the technology that can be included according to this disclosure. Figure 5 A block diagram of an example basic model mapping structure in a 3D object model structure.

[0019] Figure 7 This is an example of the technology that can be included according to this disclosure. Figure 6 A block diagram of an example basic model component structure in the basic model mapping structure.

[0020] Figure 8 This is a flowchart illustrating an example method for storing, transmitting, and retrieving data of a 3D object model structure according to the technology of this disclosure.

[0021] Figure 9 This is a flowchart illustrating an example method for generating an ISO Basic Media File Format (ISO BMFF) file according to the technology of this disclosure, the ISO Basic Media File Format (ISO BMFF) file including a basic model of a 3D object at various levels of detail.

[0022] Figure 10 This is a flowchart illustrating an example method for retrieving 3D object media data according to the technology of this disclosure. Detailed Implementation

[0023] Generally, this disclosure describes techniques for storing and transmitting three-dimensional (3D) object model data, such as avatar model data, which can be used in extended reality (XR) sessions. For example, an XR session can be an extended reality (AR) session, a mixed reality (MR) session, or a virtual reality (VR) session. An XR session can take place between two or more participants, each using a corresponding XR-capable device, such as a user equipment (UE) device.

[0024] Immersive XR experiences are typically based on a shared virtual space where people (represented by their avatars) join and interact with each other and the environment. An avatar can be a real representation of the user or a "cartoonish" representation. Avatars can be animated to mimic the user's body posture and facial expressions.

[0025] This disclosure generally describes the storage and transmission of 3D object models (such as user avatars) involved in an XR session. For example, various levels of detail (LODs) of the 3D object model can be stored, including storing the model's components for each LOD. In this way, a base model can be distributed to users involved in the XR session and then modified using animation streams. That is, when a user of a device (e.g., a user equipment (UE) device) participates in an XR session with other users, that user's device can share a base 3D object model with other users and then transmit animation streams to animate that 3D object model. This disclosure generally relates to techniques for efficiently storing base models for rapid distribution and adaptation to the network connectivity and processing capabilities of each participant. This disclosure also describes techniques for transmitting data representing 3D object models to (and retrieving) the devices of other users involved in the XR session. Although generally described in relation to the storage and transmission of user avatars, these techniques can generally also be applied to other animable 3D objects in an XR scene, such as non-user avatars, animable or movable objects in the XR scene, etc.

[0026] As an example, a user of an XR device (such as an XR headset) can participate in an XR communication session with one or more other users (e.g., during an XR conference call, XR gaming, etc.). Each user can have a corresponding avatar to be depicted to represent that user in the XR communication session. The avatar can be presented at a specific location in 3D space. Each user can control their position in 3D space through physical movement (which can be tracked by the XR device) or by providing movement input using a controller or other input device.

[0027] Generally, when a 3D object is close to the user in 3D space, the user will only be able to appreciate the high-level-of-detail model of the avatar or other 3D objects in that space. Therefore, XR devices can retrieve lower-LOD models instead of high-level-of-detail models of 3D objects that are relatively far from the user in 3D space. This saves bandwidth when retrieving lower-LOD models and reduces the processing requirements for rendering them. Thus, such bandwidth and processing resources can be allocated to higher-LOD models of 3D objects that are relatively close to the user in 3D space. Available network bandwidth can also be used to determine the level of detail of 3D models, for example, to ensure that the data needed for each element in the model can be retrieved in a timely manner.

[0028] After retrieving the Level of Detail (LOD) for each model in the model set, the XR device can receive information indicating the animation stream to be applied to the model. In this way, the animation stream remains independent of the model's LODs. That is, each LOD in the model set can use the same animation stream.

[0029] The techniques disclosed herein can generally improve the efficiency of storing, transmitting, and rendering 3D object models. The techniques disclosed herein can also allow different representations of 3D object models to be delivered to different users, for example, to accommodate varying network connections, available bandwidth, and processing power.

[0030] The technology disclosed herein can be applied to video files that conform to video data encapsulated according to the ISO Basic Media File Format (BMFF) or its extensions. The technology disclosed herein can also be applied to other file formats, such as Scalable Video Codec (SVC), Advanced Video Codec (AVC), 3GPP, and / or Multi-View Video Codec (MVC), or other similar video file formats.

[0031] Figure 1 This is a block diagram illustrating an example system 10 for implementing techniques for streaming media data over a network. In this example, system 10 includes a content preparation device 20, a server device 60, and a client device 40. Client device 40 and server device 60 are communicatively coupled via a network 74, which may include the Internet. In some examples, content preparation device 20 and server device 60 may also be coupled via network 74 or another network, or may be directly communicatively coupled. In some examples, content preparation device 20 and server device 60 may include the same device.

[0032] exist Figure 1In the example, content preparation device 20 includes an audio source 22 and a video source 24. Audio source 22 may include, for example, a microphone that generates electrical signals representing captured audio data to be encoded by audio encoder 26. Alternatively, audio source 22 may include: a storage medium storing previously recorded audio data; an audio data generator, such as a computerized synthesizer; or any other audio data source. Video source 24 may include: a video camera that generates video data to be encoded by video encoder 28; a storage medium encoding previously recorded video data; a video data generation unit, such as a computer graphics source; or any other video data source. Content preparation device 20 is not necessarily communicatively coupled to server device 60 in all examples, but multimedia content may be stored on a separate medium that is read by server device 60.

[0033] The raw audio and video data may include analog or digital data. Analog data may be digitized before being encoded by audio encoder 26 and / or video encoder 28. While a speaker is speaking, audio source 22 may obtain audio data from that speaker, and video source 24 may simultaneously obtain video data of that speaker. In other examples, audio source 22 may include a computer-readable storage medium containing stored audio data, and video source 24 may include a computer-readable storage medium containing stored video data. Thus, the techniques described in this disclosure can be applied to live, streaming, real-time audio and video data, or to archived, pre-recorded audio and video data.

[0034] An audio frame corresponding to a video frame is typically an audio frame containing audio data, which is simultaneously captured (or generated) by audio source 22 and video data captured (or generated) by video source 24 and contained within the video frame. For example, when a speaker typically generates audio data by speaking, audio source 22 captures the audio data, and video source 24 simultaneously (i.e., while audio source 22 is capturing audio data) captures the speaker's video data. Therefore, an audio frame can temporally correspond to one or more specific video frames. Thus, an audio frame corresponding to a video frame typically corresponds to the following situation: in which audio data and video data are captured simultaneously, and in this situation, the audio frame and video frame respectively include the simultaneously captured audio data and video data.

[0035] In some examples, audio encoder 26 may encode a timestamp into each encoded audio frame, where the timestamp indicates the time when the audio data for the encoded audio frame was recorded, and similarly, video encoder 28 may encode a timestamp into each encoded video frame, where the timestamp indicates the time when the video data for the encoded video frame was recorded. In such examples, the audio frame corresponding to the video frame may include: an audio frame including a timestamp, and a video frame including the same timestamp. Content preparation device 20 may include an internal clock, which audio encoder 26 and / or video encoder 28 may use to generate timestamps, or audio source 22 and video source 24 may use the internal clock to associate audio and video data with timestamps, respectively.

[0036] In some examples, audio source 22 may transmit data corresponding to the time when the audio data was recorded to audio encoder 26, and video source 24 may transmit data corresponding to the time when the video data was recorded to video encoder 28. In some examples, audio encoder 26 may encode sequence identifiers into the encoded audio data to indicate the relative time order of the encoded audio data, rather than the absolute time when the audio data was recorded; similarly, video encoder 28 may use sequence identifiers to indicate the relative time order of the encoded video data. Similarly, in some examples, sequence identifiers may be mapped or otherwise associated with timestamps.

[0037] Audio encoder 26 typically produces encoded audio data streams, while video encoder 28 produces encoded video data streams. Each individual data stream (whether audio or video) can be referred to as an elementary stream. An elementary stream is a single, digitally decoded (possibly compressed) component of a media presentation. For example, the decoded video or audio portion of a media presentation can be an elementary stream. Elementary streams can be converted into Packed Elementary Streams (PES) before being encapsulated into a video file. Within the same media presentation, stream IDs can be used to distinguish PES packets belonging to one elementary stream from those belonging to another. The basic unit of data in an elementary stream is the packed elementary stream (PES) packet. Therefore, decoded video data typically corresponds to an elementary video stream. Similarly, audio data corresponds to one or more corresponding elementary streams.

[0038] exist Figure 1In one example, the encapsulation unit 30 of the content preparation device 20 receives a base stream of video data including decoded data from the video encoder 28 and a base stream of audio data including decoded data from the audio encoder 26. In some examples, both the video encoder 28 and the audio encoder 26 may include packers for forming PES packets based on the encoded data. In other examples, both the video encoder 28 and the audio encoder 26 may interface with corresponding packers for forming PES packets based on the encoded data. In other examples, the encapsulation unit 30 may include packers for forming PES packets based on the encoded audio and video data.

[0039] Video encoder 28 is capable of encoding video data of multimedia content in various ways to produce different representations of the multimedia content at various bit rates and utilizing various characteristics such as pixel resolution, frame rate, compliance with various decoding standards, compliance with various profiles and / or profile levels used for various decoding standards, representations with one or more views (e.g., for two-dimensional or three-dimensional playback), or other such characteristics. As used in this disclosure, the representation may include one of the following: audio data, video data, text data (e.g., for closed captions), or other such data. The representation may include a primary stream, such as an audio primary stream or a video primary stream. Each PES packet may include a stream_id, which identifies the primary stream to which the PES packet belongs. Encapsulation unit 30 is responsible for assembling the primary streams into streamable media data.

[0040] Encapsulation unit 30 receives PES packets from audio encoder 26 and video encoder 28 for media presentation of the basic stream, and forms corresponding Network Abstraction Layer (NAL) units based on the PES packets. Decoded video segments can be organized into NAL units that provide a “network-friendly” video representation for addressing applications such as video telephony, storage, broadcasting, or streaming. NAL units can be classified as Video Decoding Layer (VCL) NAL units and non-VCL NAL units. VCL units may contain the core compression engine and may include block, macroblock, and / or slice-level data. Other NAL units may be non-VCL NAL units. In some examples, a picture of the decoded data in a time instance (typically presented as a picture of the main decoded data) may be included in an access unit, which may include one or more NAL units.

[0041] Non-VCL NAL units can include parameter set NAL units and SEI NAL units, etc. Parameter sets can contain sequence-level header information (in the Sequence Parameter Set (SPS)) and picture-level header information that changes very little (in the Picture Parameter Set (PPS)). Using parameter sets (e.g., PPS and SPS), it is not necessary to repeat information that changes very little for each sequence or picture; therefore, decoding efficiency can be improved. Furthermore, the use of parameter sets allows for out-of-band transmission of important header information, thus avoiding redundant transmissions required for error recovery. In an example of out-of-band transmission, parameter set NAL units can be transmitted on a different channel than other NAL units (such as SEI NAL units).

[0042] Supplemental Enhancement Information (SEI) can contain information that is not essential for decoding image samples from VCL NAL units but can assist in processes related to decoding, display, error recovery, and other purposes. SEI messages can be included in non-VCL NAL units. SEI messages are a specification part of some standards and are therefore not always mandatory for specific implementations of standards-compliant decoders. SEI messages can be sequence-level or image-level. Some sequence-level information can be included in SEI messages, such as the scalability information SEI message in the SVC example and the view scalability information SEI message in MVC. These example SEI messages can convey information such as the extraction and characteristics of operation points.

[0043] Server device 60 includes a Real-Time Transport Protocol (RTP) sending unit 70 and a network interface 72. In some examples, server device 60 may include multiple network interfaces. Furthermore, any or all of the features of server device 60 may be implemented on other devices in the content delivery network, such as routers, bridges, proxy devices, switches, or other devices. In some examples, intermediate devices in the content delivery network may cache data for multimedia content 64 and include components that substantially conform to the components of server device 60. Generally, network interface 72 is configured to transmit and receive data via network 74.

[0044] RTP sending unit 70 is configured to deliver media data to client device 40 via network 74 according to RTP, which is standardized in Internet Engineering Task Force (IETF) Request for Comments (RFC) 3550. RTP sending unit 70 may also implement RTP-related protocols such as RTP Control Protocol (RTCP), Real-time Streaming Protocol (RTSP), Session Initiation Protocol (SIP), and / or Session Description Protocol (SDP). RTP sending unit 70 can transmit media data via network interface 72, which implements Uniform Datagram Protocol (UDP) and / or Internet Protocol (IP). Therefore, in some examples, server device 60 may use network 74 to transmit media data via RTP and RTSP over UDP.

[0045] RTP sending unit 70 may receive RTSP description requests from, for example, client device 40. RTSP description requests may include data indicating what types of data the client device 40 supports. RTP sending unit 70 may respond to client device 40 with data indicating a media stream (such as media content 64), which may be transmitted to client device 40 together with a corresponding network location identifier (such as a Uniform Resource Locator (URL) or Uniform Resource Name (URN)).

[0046] Then, RTP sending unit 70 can receive an RTSP establishment request from client device 40. The RTSP establishment request typically indicates how the media stream will be transmitted. The RTSP establishment request may include a network location identifier and transmission specifier for the requested media data (e.g., media content 64), such as the local port used to receive RTP data and control data (e.g., RTCP data) on client device 40. RTP sending unit 70 can respond to the RTSP establishment request with acknowledgments and data representing the port of server device 60, through which the RTP data and control data will be transmitted. RTP sending unit 70 can then receive an RTSP playback request to allow the media stream to be "played," i.e., transmitted to client device 40 via network 74. RTP sending unit 70 can also receive an RTSP teardown request to terminate the streaming session; in response to this RTSP teardown request, RTP sending unit 70 can stop transmitting media data for the corresponding session to client device 40.

[0047] Similarly, RTP receiving unit 52 can initiate a media stream by initially sending an RTSP description request to server device 60. The RTSP description request can indicate the data types supported by client device 40. RTP receiving unit 52 can then receive from server device 60 a reply specifying an available media stream (such as media content 64) that can be transmitted to client device 40 along with a corresponding network location identifier (such as a Uniform Resource Locator (URL) or Uniform Resource Name (URN)).

[0048] Then, RTP receiving unit 52 can generate an RTSP establishment request and transmit it to server device 60. As noted above, the RTSP establishment request may include a network location identifier for the requested media data (e.g., media content 64) and a transport specification, such as the local port used to receive RTP data and control data (e.g., RTCP data) on client device 40. In response, RTP receiving unit 52 can receive an acknowledgment from server device 60, including the port of server device 60 used to transmit media data and control data.

[0049] After a media streaming session is established between server device 60 and client device 40, the RTP sending unit 70 of server device 60 can transmit media data (e.g., packets of media data) to client device 40 according to the media streaming session. Server device 60 and client device 40 can exchange control data (e.g., RTCP data) indicating, for example, the receive statistics of client device 40, so that server device 60 can perform congestion control or otherwise diagnose and resolve transmission failures.

[0050] Network interface 54 can receive selected media presentation and provide it to RTP receiving unit 52, which in turn can provide the media data to decapsulation unit 50. Decapsulation unit 50 can decapsulate the elements of a video file into a PES stream, unpack the PES stream to retrieve encoded data, and transmit the encoded data to audio decoder 46 or video decoder 48, depending on whether the encoded data is part of an audio stream or a video stream (e.g., as indicated by the PES packet header of the stream). Audio decoder 46 decodes the encoded audio data and transmits the decoded audio data to audio output 42, while video decoder 48 decodes the encoded video data and transmits the decoded video data (which may include multiple views of the stream) to video output 44.

[0051] The video encoder 28, video decoder 48, audio encoder 26, audio decoder 46, encapsulation unit 30, RTP receiver unit 52, and decapsulation unit 50 can all be implemented as any of a variety of suitable processing circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic circuits, software, hardware, firmware, or any combination thereof. Each of the video encoder 28 and video decoder 48 can be included in one or more encoders or decoders, and either the video encoder or the video decoder can be integrated as part of a combined video encoder / decoder (CODEC). Similarly, each of the audio encoder 26 and audio decoder 46 can be included in one or more encoders or decoders, and either the audio encoder or the audio decoder can be integrated as part of a combined CODEC. The apparatus including the video encoder 28, video decoder 48, audio encoder 26, audio decoder 46, encapsulation unit 30, RTP receiver unit 52, and / or decapsulation unit 50 can include integrated circuits, microprocessors, and / or wireless communication devices, such as cellular phones.

[0052] Client device 40, server device 60, and / or content preparation device 20 may be configured to operate according to the techniques of this disclosure. For illustrative purposes, these techniques are described with respect to client device 40 and server device 60. However, it should be understood that content preparation device 20 may also be configured to perform these techniques as an alternative to (or other than) server device 60.

[0053] Encapsulation unit 30 can form NAL units, which include a header identifying the program to which the NAL unit belongs and a payload, such as audio data, video data, or data describing the transport or program stream corresponding to the NAL unit. For example, in H.264 / AVC, a NAL unit includes a 1-byte header and a payload of varying size. NAL units whose payloads include video data can include video data at various granularities. For example, a NAL unit can include video data blocks, multiple blocks, video data slices, or entire frames of video data. Encapsulation unit 30 can receive encoded video data in PES packet format with elementary streams from video encoder 28. Encapsulation unit 30 can associate each elementary stream with its corresponding program.

[0054] The encapsulation unit 30 can also assemble access units based on multiple NAL units. Generally, an access unit may include one or more NAL units representing a video data frame, and the corresponding audio data (when the audio data is available). Access units typically include all NAL units for a single output time instance, such as all audio and video data for a single time instance. For example, if each view has a frame rate of 20 frames per second (fps), each time instance may correspond to a time interval of 0.05 seconds. During this time interval, specific frames for all views with the same access unit (same time instance) can be rendered simultaneously. In one example, an access unit may include a decoded image within a time instance, which can be rendered as the primary decoded image.

[0055] Therefore, the access unit can include all audio and video frames of a common time instance, such as those related to time. X All corresponding views. This disclosure also refers to the encoded image of a particular view as a "view component." That is, a view component may include an image (or frame) encoded for a particular view at a particular time. Therefore, an access unit can be defined as all view components including instances of the same time. The decoding order of the access units need not be the same as the output order or display order.

[0056] After the encapsulation unit 30 has assembled the NAL units and / or access units into a video file based on the received data, the encapsulation unit 30 passes the video file to the output interface 32 for output. In some examples, the encapsulation unit 30 may store the video file locally or transmit the video file to a remote server via the output interface 32, instead of transmitting the video file directly to the client device 40. The output interface 32 may include, for example, a transmitter, a transceiver, a device for writing data to a computer-readable medium (such as, for example, an optical drive, a magnetic media drive (e.g., a floppy disk drive)), a universal serial bus (USB) port, a network interface, or other output interface. The output interface 32 outputs the video file to a computer-readable medium, such as, for example, a transmitting signal, a magnetic medium, an optical medium, a memory, a flash drive, or other computer-readable media.

[0057] Network interface 54 can receive NAL units or access units via network 74 and provide NAL units or access units to decapsulation unit 50 via RTP receiving unit 52. Decapsulation unit 50 can decapsulate the elements of the video file into a PES stream, unpack the PES stream to retrieve encoded data, and transmit the encoded data to audio decoder 46 or video decoder 48, depending on whether the encoded data is part of an audio stream or a video stream (e.g., as indicated by the PES packet header of the stream). Audio decoder 46 decodes the encoded audio data and transmits the decoded audio data to audio output 42, while video decoder 48 decodes the encoded video data and transmits the decoded video data (which may include multiple views of the stream) to video output 44.

[0058] Figure 2 This is a block diagram illustrating the elements of example video file 150. As described above, video files based on the ISO Basic Media File Format (BMFF) and its extensions store data in a series of objects called "boxes". Figure 2 In the example, video file 150 includes a file type (FTYP) box 152, a movie (MOOV) box 154, a segment index (sidx) box 162, a movie clip (MOOF) box 164, and a movie clip random access (MFRA) box 166. Although Figure 2 This represents an example of a video file, but it should be understood that other media files may include other types of media data (e.g., audio data, timing text data, etc.) structured according to ISO BMFF and its extensions similar to the data in video file 150.

[0059] The File Type (FTYP) box 152 generally describes the file type of the video file 150. The File Type box 152 may include data identifying the specifications for the best use of the video file 150. The File Type box 152 may optionally be placed before the MOOV box 154, the Movie Clip box 164, and / or the MFRA box 166.

[0060] exist Figure 2 In the example, MOOV box 154 includes a Movie Header (MVHD) box 156, a Track (TRAK) box 158, and one or more Movie Extension (MVEX) boxes 160. Generally, the MVHD box 156 can describe the general characteristics of the video file 150. For example, the MVHD box 156 may include data describing when the video file 150 was initially created, when the video file 150 was last modified, the timestamp of the video file 150, the playback duration of the video file 150, or other data generally describing the video file 150.

[0061] TRAK box 158 may include track data of video file 150. TRAK box 158 may include a Track Header (TKHD) box that describes the characteristics of the track corresponding to TRAK box 158. In some examples, TRAK box 158 may include decoded video images, while in other examples, the decoded video images of the track may be included in movie clip 164, which may be referenced by data from TRAK box 158 and / or sidx box 162.

[0062] In some examples, video file 150 may include more than one track. Therefore, MOOV box 154 may include TRAK boxes in a number equal to the number of tracks in video file 150. TRAK boxes 158 may describe the characteristics of the corresponding tracks in video file 150. For example, TRAK boxes 158 may describe the temporal and / or spatial information of the corresponding tracks. When encapsulation unit 30 ( Figure 1 When a parameter set track is included in a video file (such as video file 150), a TRAK box similar to the TRAK box 158 of the MOOV box 154 can describe the characteristics of the parameter set track. The encapsulation unit 30 can signal the presence of a sequence level SEI message in the parameter set track within the TRAK box describing the parameter set track.

[0063] The MVEX box 160 can describe the characteristics of the corresponding movie clip 164, for example, by signaling to the video file 150 that, in addition to the video data (if any) included in the MOOV box 154, the movie clip 164 is also included. In the context of streaming video data, decoded video images may be included in the movie clip 164 instead of the MOOV box 154. Therefore, all decoded video samples may be included in the movie clip 164 instead of the MOOV box 154.

[0064] MOOV boxes 154 may include an equal number of MVEX boxes 160 as the number of movie segments 164 in the video file 150. Each MVEX box in the MVEX boxes 160 may describe the characteristics of a corresponding movie segment in the movie segment 164. For example, each MVEX box may include a Movie Extended Header Box (MEHD) box, which describes the time duration of a corresponding movie segment in the movie segment 164.

[0065] As noted above, encapsulation unit 30 may store a sequence data set in a video sample that does not include the actual decoded video data. The video sample may typically correspond to an access unit, which is a representation of a decoded picture at a specific time instance. In the context of AVC, the decoded picture includes one or more VCL NAL units (containing information about all pixels used to construct the access unit) and other associated non-VCL NAL units (such as SEI messages). Therefore, encapsulation unit 30 may include a sequence data set in a movie clip of movie clip 164, which may include a sequence-level SEI message. Encapsulation unit 30 may also signal the presence of the sequence data set and / or the sequence-level SEI message as being present in a movie clip of movie clip 164 within an MVEX box of MVEX box 160 corresponding to a movie clip of movie clip 164.

[0066] SIDX box 162 is an optional element of video file 150. That is, video files conforming to 3GPP file formats or other such file formats do not necessarily need to include SIDX box 162. According to the example of the 3GPP file format, SIDX boxes can be used to identify sub-segments of a segment (e.g., a segment contained within video file 150). The 3GPP file format defines a sub-segment as "a self-contained set of one or more consecutive movie clip boxes having a corresponding media data box and a media data box containing data referenced by a movie clip box that must follow that movie clip box and precede the next movie clip box containing information about the same track." The 3GPP file format also instructs that the SIDX box "contains a series of references to sub-segments of the (sub)segment recorded by that box. The referenced sub-segments are consecutive in presentation time. Similarly, bytes referenced by the segment index box are always consecutive within the segment. The size of the reference gives a count of the number of bytes in the referenced material."

[0067] SIDX box 162 typically provides information representing one or more sub-segments of a segment included in video file 150. For example, such information may include the playback time of the start and / or end of the sub-segment, the byte offset of the sub-segment, whether the sub-segment includes (e.g., begins at) a Stream Access Point (SAP), the type of SAP (e.g., whether the SAP is an Instant Decoder Refresh (IDR) picture, a Clean Random Access (CRA) picture, a Broken Link Access (BLA) picture, etc.), the location of the SAP within the sub-segment (in terms of playback time and / or byte offset), etc.

[0068] Movie clip 164 may include one or more decoded video pictures. In some examples, movie clip 164 may include one or more groups of pictures (GOPs), where each GOP may include several decoded video pictures, such as frames or images. Additionally, as described above, in some examples, movie clip 164 may include a sequence of data. Each movie clip in movie clip 164 may include a movie clip header box (MFHD). Figure 2 (Not shown in the image). The MFHD box can describe the characteristics of the corresponding movie clip, such as the movie clip's sequence number. Movie clips 164 can be included in video file 150 in the order of their sequence numbers.

[0069] MFRA box 166 can describe random access points within movie segments 164 of video file 150. This can help perform special effects modes, such as searching for specific time positions (i.e., playback times) within segments encapsulated by video file 150. In some examples, MFRA box 166 is typically optional and does not need to be included in the video file. Similarly, client devices (such as client device 40) do not necessarily need to refer to MFRA box 166 to correctly decode and display the video data of video file 150. MFRA box 166 can include a number equal to the number of tracks in video file 150, or in some examples, a number equal to the number of media tracks (e.g., non-cue tracks) in video file 150 (TFRA) boxes (not shown).

[0070] In some examples, movie clip 164 may include one or more streaming access points (SAPs), such as IDR pictures. Similarly, MFRA box 166 can provide an indication of the location within the SAP video file 150. Therefore, a temporal subsequence of video file 150 can be formed from the SAP of video file 150. The temporal subsequence may also include other pictures, such as SAP-dependent P-frames and / or B-frames. Frames and / or slices of the temporal subsequence can be arranged within segments such that frames / slices of the temporal subsequence that depend on other frames / slices of the subsequence can be correctly decoded. For example, in a hierarchical arrangement of data, data used for prediction of other data may also be included in the temporal subsequence.

[0071] Figure 3This is a flowchart illustrating an example avatar animation workflow that can be used during an XR session. In this example, the received animation stream data 170 includes face blend shape, body blend shape, hand joints, head pose, and audio stream data. The face blend shape, body blend shape, and hand joints may correspond to the animation stream to be applied to the user A avatar base model 172. Specifically, according to the techniques disclosed herein, the data of the user A avatar base model 172 can be stored at various levels of detail. Therefore, the rendering component 174 can retrieve the data of the user A avatar base model 172 at an appropriate level of detail (e.g., based on the distance between the current user and user A in 3D space). The rendering component 174 can then use the received animation stream data 170 to animate the avatar base model. Finally, the animated avatar base model can be presented to the current user via the display 176. Additionally, the current user's movement data can be used by the future pose prediction unit 178 to predict the user's future pose.

[0072] Figure 4 This is a flowchart illustrating an example XR session between two user equipment (UE) devices and a shared space server device. For example... Figure 4 As shown in the example, two or more UEs can participate in an XR session. UEs can send and receive data representing their animation streams and other 3D model data to and from a shared space server. For example, various sensors (such as cameras, trackers, LiDAR, etc.) can track user movements, such as facial movements (e.g., during speech or as an emotional response), hand movements, walking movements, etc. These movements can be converted into animation streams by, for example, UE 182 and transmitted to the shared space server. The shared space server can then transmit the animation streams to UE 184.

[0073] UE 184 can determine the distance between its user and UE 182's user in the 3D space of the XR session. The user's position can be represented by a 3D vector relative to the origin of the 3D space. Therefore, the user of UE 182 can be located at position {x1, y1, z1} in 3D space, while the user of UE 184 can be located at position {x2, y2, z2} in 3D space. The distance between the users of UE 182 and UE 184 can be calculated as follows: Distance =

[0074] UE 184 can use this distance to determine the appropriate Level of Detail (LOD) for the 3D object model of the user's avatar for UE 182, to be requested from the shared space server. For example, UE 184 can be configured with various thresholds corresponding to each available LOD of the avatar's 3D object model. Therefore, UE 184 can determine which threshold the determined distance satisfies, and then request the LOD corresponding to that threshold for the 3D object model of the user's avatar for UE 182. UE 184 can then use the animation stream received from UE 182 via the shared space server to animate the retrieved version of the 3D object model with the requested LOD.

[0075] Figure 5 This is a block diagram illustrating an example 3D object model structure 200 that can be stored in an ISO Basic Media File Format (ISO BMFF) file according to the technology of this disclosure. For example, the 3D object model structure 200 can be stored... Figure 2 The video file 150. Generally speaking, this disclosure describes techniques for storing 3D object models (such as avatar base models) in ISO BMFF files, which can be used during an XR session. This disclosure also describes techniques for retrieving 3D object model data from ISO BMFF files (or from devices that store 3D object model data in ISO BMFF files).

[0076] Storing 3D object models in ISO BMFF files according to the techniques disclosed herein allows for the extraction of relevant representations based on the desired target size and complexity. For example, 3D object models can be stored at various levels of detail (LOD), which can have different complexities, sizes, etc. Thus, for example, if a user is close to the 3D object represented by the 3D object model in a virtual scene, the user's device can retrieve a higher LOD representation of the 3D object, while if the user is relatively far from the 3D object in the virtual scene, the user's device can retrieve a lower LOD representation of the 3D object.

[0077] Storing 3D object models in ISO BMFF files according to the technology disclosed herein also allows for the encryption of all or only a subset of the component sets of the 3D object model. This protects the 3D object model from copying or digital theft by unauthorized users.

[0078] A 3D object model may include metadata describing all components of the 3D object model (e.g., an avatar) and provide that metadata to describe the underlying model.

[0079] Figure 5An example 3D object model 200 is depicted, which includes metadata 202, a number of Levels of Detail (LODs) 204, and a set of basic model maps 206. Metadata 202 may include key items containing master metadata, such as the name, type, gender, age, or other descriptive data of the user represented by the avatar corresponding to the 3D object model 200. The set of basic model maps 206 may include one map for each of the several LODs 204. For example, if the number of LODs 204 has a value of 5, there may be five basic model maps 206. The basic model maps 206 provide the association between LODs, the resulting size and complexity of the corresponding LODs, and the components of the basic model within the corresponding LODs.

[0080] The following pseudocode represents an example data structure corresponding to 3D object model 200: aligned(8) class BaseAvatarModel extends FullBox('bavm', version,flags) { AvatarMetadataEntry(); unsinged int number_lods; BaseModelMapping mappings[number_lods]; }

[0081] Figure 6 This is an example of the technology that can be included according to this disclosure. Figure 5 A block diagram of the example basic model mapping structure 210 in the 3D object model structure 200. That is, Figure 5 Each basic model mapping in the basic model mapping 206 may include Figure 6 The basic model map 210 consists of elements. The basic model map structure 210 includes an LOD identifier (ID) 212, a size 214, a number of mesh faces 216, a number of components 218, and basic model components 220. The basic model map 210 is associated with the level of detail (LOD) indicated by the LOD ID 212. The size 214 provides information about the size and complexity of the corresponding LOD. The number of mesh faces 216 represents the number of faces (e.g., individual planar elements) of the components of the LOD.

[0082] The component number 218 describes the number of basic model component structures 220 included in the basic model map 210. Basic model components 220 represent the component list of the basic model in LOD. The number of basic model components 220 may be equal to the value of component number 218. The map may be available at different resolutions and may be stored as image items. Basic model components 220 may be compressed. For example, mesh compression may be used to compress geometric data, image or video compression standards (such as JPEG, ITU-T H.264, ITU-T H.265, ITU-T H.266, etc.) may be used to compress images, and other data may be entropy-decoded.

[0083] The following pseudocode represents an example data structure corresponding to the basic model mapping 210: aligned(8) class BaseModelMapping extends FullBox("bmma", version,flags) { unsigned int(16) lod_id; unsinged int(32) size_in_bytes; unsigned int(32) num_mesh_faces; unsinged int(16) num_components; BaseModelComponent components[num_components]; }

[0084] Figure 7 This is an example of the technology that can be included according to this disclosure. Figure 6 A block diagram of the example basic model component structure 230 in the basic model mapping structure 210. That is, Figure 6 Each basic model component structure in the basic model component structure 220 may include Figure 7 The basic model component structure 230 consists of the elements of this structure. In this example, the basic model component 230 includes an item identifier (ID) 232, a role 234, and a parent ID 236 (which is an optional element, such as...). Figure 7 (As indicated by the dashed line in the image).

[0085] Components can be assigned roles, describing their role in relation to the underlying model, as indicated by role 234. Examples of roles include joints, blend shapes, meshes, maps, or pose transformations. Each component can be stored separately as a metadata item. Each component may include an ItemInfoEntry describing its content_type and content_encoding, indicating which compression scheme (if any) is applied. Each component can also be individually encrypted. Components can be arranged hierarchically and can use the value of parent ID 236 to indicate their corresponding parent item.

[0086] The following pseudocode represents an example data structure corresponding to basic model component 230: aligned(8) class BaseModelComponent extends FullBox("bmcp", version,flags) { unsigned int(16) item_ID; Role role unsigned int(16) parent_item_ID; optional; }

[0087] Figure 8 This is a flowchart illustrating an example method for storing, transmitting, and retrieving data of a 3D object model structure according to the technology of this disclosure. In this example, a transmitter (e.g., a first UE, such as...) Figure 1 Content preparation equipment 20 or Figure 4 The UE (182) publishes basic model details (250) to the application server (AS), such as metadata of the 3D object model (e.g., avatar) and the available levels of detail for the 3D object model. The transmitter may publish an avatar description including the content type of the main item, such as "application / mpeg.avatar.main". The transmitter may also publish the required decompression engine, information about the required decompression, and the available levels of detail. The transmitter may also publish ISO BMFF files (e.g., conforming to the technology of this disclosure) based on the present disclosure. Figures 5 to 7 (Example) Storing 3D object models to storage devices, such as Figure 1 The server device 60. That is, the transmitter can store various versions of the basic model at each level of detail.

[0088] Receiver (e.g., Figure 1The client device 40 can retrieve a description of the basic model details from the AS. The receiver can examine the description and select one LOD (252). For example, the receiver can determine the LOD based on the proximity (distance) between the receiver user and the 3D object in the virtual scene, available network bandwidth, processing power, etc. The receiver can also ensure the capability of the decoding components. The receiver can then request access to the selected LOD from the AS (254). The AS can direct the request to the transmitter. In this way, the receiver can signal the request to the transmitter for the selected LOD of the 3D object model.

[0089] The transmitter can then authorize the receiver to access the 3D object model at the requested LOD (256). The transmitter can redirect the request to a storage device (e.g., Figure 1 (Server device 60). The storage device can extract the relevant representation from the ISO BMFF file based on the request and redirect the request to a licensed server with the appropriate decryption key.

[0090] The receiver can then request the Level of Detail (LOD) of the 3D object model (e.g., the avatar of the transmitter) from the storage device (258). The storage device can authorize and extract a representation with the requested LOD from the ISO BMFF file (260) and transmit the extracted representation of the base model (in the requested LOD) to the receiver (262). If one or more components of the LOD are encrypted, the receiver can retrieve a decryption key for the encrypted components from a Digital Rights Management (DRM) server device (262). The receiver can then decrypt the encrypted components and use the components of the LOD to render the base model (264).

[0091] During an XR session between the transmitter and receiver, the transmitter can transmit animation data to the receiver. The receiver can use the animation data to transform and / or distort the underlying model and display these animations to the receiver's user via a display device.

[0092] Figure 9 This is a flowchart illustrating an example method for generating an ISO Basic Media File Format (ISO BMFF) file according to the technology of this disclosure, the ISO Basic Media File Format (ISO BMFF) file including a basic model of a 3D object at various levels of detail. Figure 9 The method can be performed by server devices (such as...) Figure 1 Server equipment 60). Figure 1 Client equipment 40 and / or source user equipment (UE) equipment (such as Figure 4 Execution of UE 182). Figure 9 The method can be performed by each UE participating in the multi-user XR communication session, and therefore, Figure 9 The method can be derived from Figure 4 Both UE 182 and UE 184 are executed.

[0093] The source device initially obtains an initial base model of a 3D object (such as a user avatar) (280). For example, the user may create a base model, download a base model, store a base model, modify an existing model, or otherwise load and / or build a base model. The source device may then build multiple Level of Detail (LOD) versions of the base model (282). For example, the source device may build one or more reduced LOD versions from the initial base model. To reduce LOD, the source device may reduce texture resolution, reduce the number of polygons in the model (e.g., the number of vertices / edges / triangles) (which in turn reduces the number of model faces to which textures are applied), reduce the number of components in the model, reduce the size of the model, etc.

[0094] The source device can then generate metadata (284) for the base model. The metadata typically indicates the number of LOD versions of the available base model and the characteristics of each LOD version. For example, for each LOD version, the metadata may indicate the size of the base model for that LOD version, the number of mesh faces for that LOD version, and the number of components for that LOD version. The metadata may also include (e.g., using LOD identifier values) mapping data that maps the metadata for each LOD version to the corresponding LOD version, such as... Figure 6 As shown in the image.

[0095] The source device may then store metadata in an ISO BMFF file (286), and each LOD version in the LOD versions of the base model in an ISO BMFF file (288). The source device may also (e.g., using LOD identifiers) store data that associates the metadata of each LOD with the corresponding LOD (290). The source device may also store the ISO BMFF file in a server device (292).

[0096] so, Figure 9 The method represents an example of a method for storing media data, which includes: storing a three-dimensional (3D) object model in an ISO Basic Media File Format (ISO BMFF) file, including: storing metadata of the 3D object model in the ISO BMFF file; storing several levels of detail (LODs) of the 3D object model in the ISO BMFF file; and for each of the several LODs, storing data in the ISO BMFF file that associates the LOD with data representing the size, complexity, and components of the 3D object model with respect to the LOD.

[0097] Figure 10This is a flowchart illustrating an example method for retrieving 3D object media data according to the technology of this disclosure. Figure 10 The method can be derived from Figure 1 Client device 40 or user equipment (UE) device (such as Figure 4 Execution of UE 184). Figure 10 The method can be performed by each UE participating in the multi-user XR communication session, and therefore, Figure 10 The method can be derived from Figure 4 Both UE 182 and UE 184 are executed.

[0098] The client device can initially retrieve metadata about the base model of a 3D object, which represents the available Level of Detail (LOD) versions (300) of the base model. As discussed above, the metadata can (e.g., using LOD identifiers for LOD versions) indicate, for example, the size of each LOD version, the number of mesh faces per LOD version, the number of components per LOD version, and the mapping between the indicated characteristics of each LOD version and the LOD version.

[0099] The client device can then select one of the LOD versions (302). For example, the client device can determine the distance between the user of the client device and the user who wants to retrieve a 3D object model in 3D space. In some examples, available network bandwidth may also be used when selecting an LOD version. The client device can then request a base model in the selected LOD (304). For example, the client device can request the base model from a storage device such as a central avatar repository. In response to the request, the client device can receive the base model in the selected LOD (306). In some examples, if the base model is encrypted, the client device can further retrieve a decryption key (308) and decrypt the base model using the decryption key (310).

[0100] Furthermore, when participating in an XR communication session, the client device can retrieve the animation stream of the basic model (312), such as... Figure 3 As shown in the diagram. The animation stream can indicate the face blend shape, body blend shape, hand joint movements, and / or head pose to be applied to the base model. The client device can then animate the base model using the animation stream (314). The client device can ultimately present the animated base model to the user of the client device (316), for example, via a display. The animation stream can correspond to the user's actual movements associated with the base model, such as facial movements (e.g., indicating speech), hand movements, body movements, etc.

[0101] so, Figure 10The method represents an example of a method for retrieving media data, the method comprising: retrieving, via a client device, data representing a three-dimensional (3D) object model stored in an ISO Basic Media File Format (ISO BMFF) file and one or more Levels of Detail (LODs) of the 3D object model stored in the ISO BMFF file, wherein the ISO BMFF file is stored on a server device; transmitting, via the client device, a request to the server device to access data of the 3D object model in one LOD of the LOD, the data of the 3D object model in the one LOD of the LOD including the size, complexity, and components of the 3D object model for that one LOD; and, in response to the request, receiving, via the client device, the data of the 3D object model in the one LOD of the LOD, the data of the 3D object model in the one LOD of the LOD having the size, complexity, and components of the 3D object model for that one LOD of the LOD.

[0102] Thus, the techniques disclosed herein allow 3D object models (e.g., avatar base models) to be stored in ISO BMFF files. These techniques can leverage the metadata structure of ISO BMFF to efficiently store and access base models based on the desired level of detail.

[0103] The following clauses summarize various examples of the technology disclosed herein:

[0104] Clause 1: A method for storing media data, the method comprising: storing a three-dimensional (3D) object model in an ISO Basic Media File Format (ISO BMFF) file, including: storing metadata of the 3D object model in the ISO BMFF file; storing a plurality of Levels of Detail (LODs) of the 3D object model in the ISO BMFF file; and for each of the plurality of LODs, storing data in the ISO BMFF file that associates the LOD with data representing the size, complexity, and components of the 3D object model with respect to the LOD.

[0105] Clause 2: The method described in Clause 1, wherein storing the 3D object model includes storing the 3D object model as a Basic Avatar Model (BAVM) box, as an extension of the ISO BMFF FullBox.

[0106] Clause 3: The method according to any one of Clauses 1 and 2, wherein storing the data that associates the LOD with the data representing the size, complexity and components of the 3D object model with respect to the LOD includes storing the basic model mapping structure in the ISO BMFF file.

[0107] Clause 4: The method according to any one of Clauses 1 to 3, wherein the data representing the size, complexity and components of the 3D object model with respect to the LOD includes size values, a plurality of mesh faces, a plurality of components and data representing the corresponding component for each of the plurality of components.

[0108] Clause 5: The method according to Clause 4, wherein storing the data representing the size, complexity and components of the 3D object model with respect to the LOD includes an extension of storing the Basic Model Mapping (BMMA) box as a FullBox of ISO BMFF.

[0109] Clause 6: The method according to any one of Clauses 4 and 5, wherein the data representing the corresponding component includes one or more of the following: data representing the geometry of the component or one or more images of the component.

[0110] Clause 7: The method according to any one of Clauses 4 to 6 further includes, for each of the components, storing data describing the role of the component in relation to the 3D object model.

[0111] Clause 8: The role described in Clause 7 includes one or more of a joint, a blend shape, a mesh, a map, or a pose transformation.

[0112] Clause 9: The method according to any one of Clauses 4 to 8, wherein storing the data representing the components of the 3D object model for the LOD includes storing each component separately as a corresponding metadata item in the ISO BMFF file.

[0113] Clause 10: The method described in Clause 9, wherein storing each component separately as the corresponding metadata item includes storing each component as a Basic Model Component (BMCP) box, as an extension of the FullBox of ISOBMFF.

[0114] Clause 11: The method according to any one of Clauses 9 and 10, wherein storing each component separately as a corresponding metadata item in the ISO BMFF file includes, for each of the components, storing a type value representing the type of the component and an encoding value representing how the component is encoded.

[0115] Clause 12: The method according to any one of Clauses 4 to 11 further includes encrypting data used for one or more of the components.

[0116] Clause 13: The method according to any one of Clauses 4 to 12, wherein storing the data representing the component includes hierarchical storage of the data representing the component, such that each component having a parent component includes data identifying the parent component.

[0117] Clause 14: The method according to any one of Clauses 1 to 13 further includes publishing the identifier of the ISO BMFF file to an application server (AS) in a computer network.

[0118] Clause 15: The method according to any one of Clauses 1 to 14, the method further comprising receiving from the receiving device a request for accessing the 3D object model of one of the LODs.

[0119] Clause 16: The method according to Clause 15 further includes, in response to the request, providing the receiving device with data representing the size, complexity, and components of the 3D object model for the requested LOD in the LOD.

[0120] Clause 17: A method for retrieving media data, the method comprising: retrieving data representing a three-dimensional (3D) object model and one or more levels of detail (LODs) of the 3D object model; transmitting a request to access data of one LOD of the 3D object model in the LODs, the data of the one LOD of the 3D object model in the LODs including the size, complexity, and components of the 3D object model for the LOD; and, in response to the request, receiving data representing the size, complexity, and components of the 3D object model for the LOD.

[0121] Clause 18: The method according to Clause 17, wherein the data that associates the LOD with the data representing the size, complexity and components of the 3D object model with respect to the LOD includes a basic model mapping structure.

[0122] Clause 19: The method according to any one of Clauses 17 and 18, wherein the data representing the size, complexity and components of the 3D object model for the LOD includes size values, a plurality of mesh faces, a plurality of components and data representing the corresponding component for each of the plurality of components.

[0123] Clause 20: The method according to Clause 19, wherein the data representing the size, complexity and components of the 3D object model for the LOD includes a Basic Model Map (BMMA) box, the Basic Model Map (BMMA) box including an extension of the FullBox of ISO BMFF.

[0124] Clause 21: The method according to any one of Clauses 19 and 20, wherein the data representing the corresponding component includes one or more of the following: data representing the geometry of the component or one or more images of the component.

[0125] Clause 22: The method according to any one of Clauses 19 to 21, the method further comprising, for each of the components, retrieving data describing the role of the component in relation to the 3D object model.

[0126] Clause 23: The role described in accordance with Clause 22 includes one or more of a joint, a blend shape, a mesh, a map, or a pose transformation.

[0127] Clause 24: The method according to any one of Clauses 19 to 23, wherein the data representing the 3D object model for the components of the LOD includes separate corresponding metadata items for each component.

[0128] Clause 25: According to the method described in Clause 24, each of the separate corresponding metadata items in the separate corresponding metadata items includes a corresponding Basic Model Component (BMCP) box, which includes an extension of the FullBox of ISOBMFF.

[0129] Clause 26: The method according to any one of Clauses 24 and 25, wherein each of the components includes a type value representing the type of the component and an encoded value representing how the component is encoded.

[0130] Clause 27: The method according to any one of Clauses 19 to 26 further includes decrypting data used for one or more of the components.

[0131] Clause 28: The method described in Clause 27 further includes retrieving a decryption key for each of the cryptographic components.

[0132] Clause 29: The method according to any one of Clauses 19 to 28, wherein the data representing the component includes a hierarchical representation of the component, such that each component having a parent component includes data identifying the parent component.

[0133] Clause 30: An apparatus for transmitting media data, the apparatus comprising one or more components for performing the method according to any one of Clauses 1 to 29.

[0134] Clause 31: The device according to Clause 30, wherein the one or more components include a processing system, the processing system including one or more processors implemented in a circuit.

[0135] Clause 32: The apparatus according to Clause 30, wherein the apparatus includes at least one of the following: an integrated circuit; a microprocessor; or a wireless communication device.

[0136] Clause 33: An apparatus for storing media data, the apparatus comprising: components for storing a three-dimensional (3D) object model in an ISO Basic Media File Format (ISO BMFF) file, the components comprising: components for storing metadata of the 3D object model in the ISO BMFF file; components for storing a plurality of Levels of Detail (LODs) of the 3D object model in the ISO BMFF file; and components for, for each of the plurality of LODs, storing data in the ISO BMFF file that associates the LOD with data representing the size, complexity, and components of the 3D object model with respect to the LOD.

[0137] Clause 34: An apparatus for retrieving media data, the apparatus comprising: components for retrieving data representing a three-dimensional (3D) object model and one or more levels of detail (LODs) of the 3D object model; components for transmitting a request to access data of the 3D object model in one of the LODs, the data of the 3D object model in the one LOD including the size, complexity, and components of the 3D object model with respect to the LOD; and components for receiving, in response to the request, data representing the size, complexity, and components of the 3D object model with respect to the LOD.

[0138] Clause 35: A computer-readable storage medium having instructions stored thereon, which, when executed, cause a processor to perform a method according to any one of Clauses 1 to 29.

[0139] Clause 36: A method for storing media data, the method comprising: storing a three-dimensional (3D) object model in an ISO Basic Media File Format (ISO BMFF) file, including: storing metadata of the 3D object model in the ISO BMFF file; storing a plurality of Levels of Detail (LODs) of the 3D object model in the ISO BMFF file; and for each of the plurality of LODs, storing data in the ISO BMFF file that associates the LOD with data representing the size, complexity, and components of the 3D object model with respect to the LOD.

[0140] Clause 37: The method according to Clause 36, wherein storing the 3D object model includes storing the 3D object model as a Basic Avatar Model (BAVM) box, as an extension of the ISO BMFF FullBox.

[0141] Clause 38: The method according to Clause 36, wherein storing the data that associates the LOD with the data representing the size, complexity and components of the 3D object model with respect to the LOD includes storing the basic model mapping structure in the ISO BMFF file.

[0142] Clause 39: The method according to Clause 36, wherein the data representing the size, complexity and components of the 3D object model for the LOD includes size values, a plurality of mesh faces, a plurality of components and data representing the corresponding component for each of the plurality of components.

[0143] Clause 40: The method according to Clause 39, wherein storing the data representing the size, complexity and components of the 3D object model with respect to the LOD includes an extension of storing the Basic Model Mapping (BMMA) box as a FullBox of ISO BMFF.

[0144] Clause 41: The method according to Clause 39, wherein the data representing the corresponding component includes one or more of the following: data representing the geometry of the component or one or more images of the component.

[0145] Clause 42: The method according to Clause 39 further includes, for each of the components, storing data describing the role of the component in relation to the 3D object model.

[0146] Clause 43: The method described in Clause 42, wherein the character includes one or more of a joint, a blend shape, a mesh, a map, or a pose transformation.

[0147] Clause 44: The method according to Clause 39, wherein storing the data representing the components of the 3D object model for the LOD includes storing each component separately as a corresponding metadata item in the ISO BMFF file.

[0148] Clause 45: The method described in Clause 44, wherein storing each component separately as the corresponding metadata item includes storing each component as a Basic Model Component (BMCP) box, as an extension of the FullBox of ISOBMFF.

[0149] Clause 46: The method according to Clause 44, wherein storing each component separately as a corresponding metadata item in the ISO BMFF file includes, for each of the components, storing a type value indicating the type of the component and an encoding value indicating how the component is encoded.

[0150] Clause 47: The method described in Clause 39 further includes encrypting data used for one or more of the components.

[0151] Clause 48: The method according to Clause 39, wherein storing the data representing the component includes hierarchical storage of the data representing the component, such that each component having a parent component includes data identifying the parent component.

[0152] Clause 49: The method of Clause 36 further includes publishing the identifier of the ISO BMFF file to an application server (AS) in a computer network.

[0153] Clause 50: The method according to Clause 36 further includes receiving from the receiving device a request for accessing the 3D object model of one of the LODs.

[0154] Clause 51: The method according to Clause 50 further includes, in response to the request, providing the receiving device with data representing the size, complexity, and components of the 3D object model for the requested LOD in the LOD.

[0155] Clause 52: A method for retrieving media data, the method comprising: retrieving data representing a three-dimensional (3D) object model and one or more levels of detail (LODs) of the 3D object model; transmitting a request to access data of one LOD of the 3D object model in the LODs, the data of the one LOD of the 3D object model in the LODs including the size, complexity, and components of the 3D object model with respect to the LOD; and, in response to the request, receiving data representing the size, complexity, and components of the 3D object model with respect to the LOD.

[0156] Clause 53: The method according to Clause 52, wherein the data that associates the LOD with the data representing the size, complexity and components of the 3D object model with respect to the LOD includes a basic model mapping structure.

[0157] Clause 54: The method according to Clause 52, wherein the data representing the size, complexity and components of the 3D object model for the LOD includes size values, a plurality of mesh faces, a plurality of components and data representing the corresponding component for each of the plurality of components.

[0158] Clause 55: The method according to Clause 54, wherein the data representing the size, complexity and components of the 3D object model for the LOD includes a Basic Model Map (BMMA) box, the Basic Model Map (BMMA) box including an extension of the FullBox of ISO BMFF.

[0159] Clause 56: The method according to Clause 54, wherein the data representing the corresponding component includes one or more of the following: data representing the geometry of the component or one or more images of the component.

[0160] Clause 57: The method according to Clause 54 further includes, for each of the components, retrieving data describing the role of the component in relation to the 3D object model.

[0161] Clause 58: The method described in Clause 57, wherein the character includes one or more of a joint, a blend shape, a mesh, a map, or a pose transformation.

[0162] Clause 59: The method according to Clause 54, wherein the data representing the 3D object model for the components of the LOD includes separate corresponding metadata items for each component.

[0163] Clause 60: According to the method described in Clause 59, each of the separate corresponding metadata items in the separate corresponding metadata items includes a corresponding Basic Model Component (BMCP) box, which includes an extension of the FullBox of ISOBMFF.

[0164] Clause 61: The method described in Clause 59, wherein each of the components includes a type value representing the type of the component and an encoded value representing how the component is encoded.

[0165] Clause 62: The method described in Clause 54 further includes decrypting data used for one or more of the components.

[0166] Clause 63: The method according to Clause 62 further includes retrieving a decryption key for each of the cryptographic components.

[0167] Clause 64: The method according to Clause 54, wherein the data representing the component includes a hierarchical representation of the component, such that each component having a parent component includes data identifying the parent component.

[0168] Clause 65: A method for storing media data, the method comprising: storing a three-dimensional (3D) object model in an ISO Basic Media File Format (ISO BMFF) file, including: storing metadata of the 3D object model in the ISO BMFF file; storing a plurality of Levels of Detail (LODs) of the 3D object model in the ISO BMFF file; and for each of the plurality of LODs, storing data in the ISO BMFF file that associates the LOD with data representing the size, complexity, and components of the 3D object model with respect to the LOD.

[0169] Clause 66: The method of claim 65, wherein storing the 3D object model comprises storing the 3D object model as a Basic Avatar Model (BAVM) box, as an extension of the ISO BMFF FullBox.

[0170] Clause 67: The method of claim 65, wherein storing the data that associates the LOD with the data representing the size, complexity and components of the 3D object model with respect to the LOD comprises storing a basic model mapping structure in the ISO BMFF file.

[0171] Clause 68: The method of claim 65, wherein the data representing the size, complexity, and components of the 3D object model for the LOD includes size values, a plurality of mesh faces, a plurality of components, and data representing the corresponding component for each of the plurality of components.

[0172] Clause 69: The method according to Clause 68, wherein storing the data representing the size, complexity and components of the 3D object model with respect to the LOD includes an extension of storing the Basic Model Mapping (BMMA) box as a FullBox of ISO BMFF.

[0173] Clause 70: The method according to Clause 68, wherein the data representing the corresponding component includes one or more of the following: data representing the geometry of the component or one or more images of the component.

[0174] Clause 71: The method according to Clause 68 further includes, for each of the components, storing data describing the role of the component in relation to the 3D object model.

[0175] Clause 72: The role described in accordance with Clause 71 includes one or more of a joint, a blend shape, a mesh, a map, or a pose transformation.

[0176] Clause 73: The method according to Clause 68, wherein storing the data representing the 3D object model for the components of the LOD includes storing each component separately as a corresponding metadata item in the ISO BMFF file.

[0177] Clause 74: The method described in Clause 73, wherein storing each component separately as the corresponding metadata item includes storing each component as a Basic Model Component (BMCP) box, as an extension of the FullBox of ISOBMFF.

[0178] Clause 75: The method according to Clause 73, wherein storing each component separately as a corresponding metadata item in the ISO BMFF file includes, for each of the components, storing a type value indicating the type of the component and an encoding value indicating how the component is encoded.

[0179] Clause 76: The method described in Clause 68 further includes encrypting data used for one or more of the components.

[0180] Clause 77: The method according to Clause 68, wherein storing the data representing the component includes hierarchical storage of the data representing the component, such that each component having a parent component includes data identifying the parent component.

[0181] Clause 78: The method of claim 65 further comprises publishing the identifier of the ISO BMFF file to an application server (AS) in a computer network.

[0182] Clause 79: The method of claim 65, further comprising receiving from a receiving device a request for accessing the 3D object model of one of the LODs.

[0183] Clause 80: The method according to Clause 79 further includes, in response to the request, providing the receiving device with data representing the size, complexity, and components of the 3D object model for the requested LOD in the LOD.

[0184] Clause 81: An apparatus for storing media data, the apparatus comprising: a memory for storing media data; and a processing system implemented in circuitry and configured to store a three-dimensional (3D) object model in an ISO Basic Media File Format (ISO BMFF) file of the media data in the memory, wherein, in order to store the 3D object model in the ISO BMFF file, the processing system is configured to: store metadata of the 3D object model in the ISO BMFF file; store a plurality of Levels of Detail (LODs) of the 3D object model in the ISO BMFF file; and for each of the plurality of LODs, store data in the ISO BMFF file that associates the LOD with data representing the size, complexity, and components of the 3D object model with respect to the LOD.

[0185] Clause 82: A method for retrieving media data, the method comprising: retrieving, via a client device, data representing a three-dimensional (3D) object model stored in an ISO Basic Media File Format (ISO BMFF) file and one or more Levels of Detail (LODs) of the 3D object model stored in the ISO BMFF file, wherein the ISO BMFF file is stored on a server device; transmitting, via the client device, a request to the server device to access data of the 3D object model in one of the LODs, the data of the 3D object model in the one LOD including the size, complexity, and components of the 3D object model for the one LOD; and, in response to the request, receiving, via the client device, the data of the 3D object model in the one LOD having the size, complexity, and components of the 3D object model for the one LOD.

[0186] Clause 83: The method according to Clause 82, wherein the data that associates the LOD with the data representing the size, complexity and components of the 3D object model with respect to the LOD includes a basic model mapping structure.

[0187] Clause 84: The method according to Clause 82, wherein the data representing the size, complexity and components of the 3D object model for the LOD includes size values, a plurality of mesh faces, a plurality of components and data representing the corresponding component for each of the plurality of components.

[0188] Clause 85: The method according to Clause 84, wherein the data representing the size, complexity and components of the 3D object model with respect to the LOD includes a Basic Model Map (BMMA) box, the Basic Model Map (BMMA) box including an extension of the FullBox of ISO BMFF.

[0189] Clause 86: The method according to Clause 84, wherein the data representing the corresponding component includes one or more of the following: data representing the geometry of the component or one or more images of the component.

[0190] Clause 87: The method according to Clause 84 further includes, for each of the components, retrieving data describing the role of the component in relation to the 3D object model.

[0191] Clause 88: The method described in Clause 87, wherein the character includes one or more of a joint, a blend shape, a mesh, a map, or a pose transformation.

[0192] Clause 89: The method according to Clause 84, wherein the data representing the 3D object model for the components of the LOD includes separate corresponding metadata items for each component.

[0193] Clause 90: The method described in Clause 89, wherein each of the separate corresponding metadata items in the separate corresponding metadata items includes a corresponding Basic Model Component (BMCP) box, the corresponding Basic Model Component (BMCP) box including an extension of the FullBox of ISOBMFF.

[0194] Clause 91: The method described in Clause 89, wherein each of the components includes a type value representing the type of the component and an encoded value representing how the component is encoded.

[0195] Clause 92: The method according to Clause 84 further includes: retrieving a decryption key for each of the cryptographic components; and using the decryption key to decrypt data for one or more of the components.

[0196] Clause 93: The method according to Clause 84, wherein the data representing the component includes a hierarchical representation of the component, such that each component having a parent component includes data identifying the parent component.

[0197] Clause 94: A client device for retrieving media data, the device comprising: a memory for storing media data; and a processing system implemented in circuitry and configured to: retrieve data representing a three-dimensional (3D) object model stored in an ISO Basic Media File Format (ISO BMFF) file and one or more Levels of Detail (LODs) of the 3D object model stored in the ISO BMFF file, wherein the ISO BMFF file is stored on a server device; transmit a request to the server device to access data of one LOD of the 3D object model, the data of the one LOD of the 3D object model including the size, complexity, and components of the 3D object model for the one LOD of the LOD; and, in response to the request, receive the data of the one LOD of the 3D object model having the size, complexity, and components of the 3D object model for the one LOD of the LOD.

[0198] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored as one or more instructions or code on a computer-readable medium or transmitted via a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium (which corresponds to a tangible medium such as a data storage medium) or a communication medium, including, for example, any medium that facilitates the transfer of a computer program from one place to another according to a communication protocol. Thus, a computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to extract instructions, code, and / or data structures for implementing the techniques described in this disclosure. Computer program products may include computer-readable media.

[0199] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other medium capable of storing desired program code in the form of instructions or data structures and accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies (such as infrared, radio, and microwave), then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies (such as infrared, radio, and microwave) are included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but instead refer to non-transient tangible storage media. As used herein, disks and optical discs include: compact optical discs (CDs), laser optical discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. The combinations described above should also be included within the scope of computer-readable media.

[0200] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the term "processor" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Furthermore, these techniques can be fully implemented in one or more circuit or logic elements.

[0201] The techniques disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless mobile phones, integrated circuits (ICs), or IC sets (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but implementation by different hardware units is not necessarily required. Specifically, as described above, various units may be combined in a codec hardware unit, or various units may be provided by a collection of interoperable hardware units (including one or more processors as described above) combined with appropriate software and / or firmware.

[0202] Various examples have been described. These and other examples are within the scope of the following claims.

Claims

1. A method for storing media data, the method comprising: Storing three-dimensional (3D) object models in ISO Basic Media File Format (ISO BMFF) files includes: The metadata of the 3D object model is stored in the ISO BMFF file; The 3D object model is stored at several levels of detail (LOD) in the ISO BMFF file; and For each of the plurality of LODs, data is stored in the ISO BMFF file that associates the LOD with data representing the size, complexity, and components of the 3D object model for the LOD.

2. The method of claim 1, wherein storing the 3D object model comprises storing the 3D object model as a Basic Avatar Model (BAVM) box, as an extension of the ISO BMFF FullBox.

3. The method of claim 1, wherein storing the data that associates the LOD with the data representing the size, complexity, and components of the 3D object model with respect to the LOD comprises storing a basic model mapping structure in the ISO BMFF file.

4. The method of claim 1, wherein the data representing the size, complexity, and components of the 3D object model for the LOD includes size values, a plurality of mesh faces, a plurality of components, and data representing the corresponding component for each of the plurality of components.

5. The method of claim 4, wherein storing the data representing the size, complexity, and components of the 3D object model with respect to the LOD comprises an extension of storing the Basic Model Mapping (BMMA) box as a FullBox of ISO BMFF.

6. The method of claim 4, wherein the data representing the corresponding component includes one or more of the following: data representing the geometry of the component or one or more images of the component.

7. The method of claim 4, further comprising, for each of the components, storing data describing the role of the component in relation to the 3D object model.

8. The method of claim 7, wherein the role comprises one or more of a joint, a blend shape, a mesh, a map, or a pose transformation.

9. The method of claim 4, wherein storing the data representing the components of the 3D object model for the LOD comprises storing each component separately as a corresponding metadata item in the ISO BMFF file.

10. The method of claim 9, wherein storing each component separately as the corresponding metadata item comprises storing each component as a Basic Model Component (BMCP) box, as an extension of the FullBox of ISOBMFF.

11. The method of claim 9, wherein storing each component separately as a corresponding metadata item in the ISO BMFF file includes, for each of the components, storing a type value indicating the type of the component and an encoding value indicating how the component is encoded.

12. The method of claim 4, further comprising encrypting data used for one or more of the components.

13. The method of claim 4, wherein storing the data representing the component comprises hierarchically storing the data representing the component, such that each component having a parent component includes data identifying the parent component.

14. The method of claim 1, further comprising publishing the identifier of the ISO BMFF file to an application server (AS) in a computer network.

15. The method of claim 1, further comprising receiving from a receiving device a request for accessing the 3D object model of one of the LODs.

16. The method of claim 15, further comprising, in response to the request, providing the receiving device with data representing the size, complexity, and components of the 3D object model for the requested LOD in the LOD.

17. An apparatus for storing media data, the apparatus comprising: The memory is used to store media data; and A processing system, implemented in a circuit and configured to store a three-dimensional (3D) object model in an ISO Basic Media File Format (ISO BMFF) file of the media data in the memory, wherein, in order to store the 3D object model in the ISO BMFF file, the processing system is configured to: The metadata of the 3D object model is stored in the ISO BMFF file; The 3D object model is stored in several levels of detail (LOD) in the ISO BMFF file; as well as For each of the plurality of LODs, data is stored in the ISO BMFF file that associates the LOD with data representing the size, complexity, and components of the 3D object model for the LOD.

18. A method for retrieving media data, the method comprising: Data representing a three-dimensional (3D) object model stored in an ISO Basic Media File Format (ISO BMFF) file and one or more levels of detail (LOD) of the 3D object model stored in the ISO BMFF file are retrieved via a client device, wherein the ISO BMFF file is stored on a server device. The client device transmits a request to the server device to access the data of the 3D object model in one of the LODs. The data of the 3D object model in the LOD includes the size, complexity, and components of the 3D object model in the LOD. as well as In response to the request, the client device receives the data of the 3D object model in the LOD of the LOD, wherein the data of the 3D object model in the LOD of the LOD includes the size, complexity and components of the 3D object model in the LOD of the LOD.

19. The method of claim 18, wherein the data that associates the LOD with the data representing the size, complexity, and components of the 3D object model with respect to the LOD includes a basic model mapping structure.

20. The method of claim 18, wherein the data representing the size, complexity, and components of the 3D object model for the LOD includes size values, a plurality of mesh faces, a plurality of components, and data representing the corresponding component for each of the plurality of components.

21. The method of claim 20, wherein the data representing the size, complexity, and components of the 3D object model with respect to the LOD includes a Basic Model Mapping (BMMA) box, the Basic Model Mapping (BMMA) box including an extension of the FullBox of ISO BMFF.

22. The method of claim 20, wherein the data representing the corresponding component includes one or more of the following: data representing the geometry of the component or one or more images of the component.

23. The method of claim 20, further comprising, for each of the components, retrieving data describing the role of the component in relation to the 3D object model.

24. The method of claim 23, wherein the role comprises one or more of a joint, a blend shape, a mesh, a map, or a pose transformation.

25. The method of claim 20, wherein the data representing the 3D object model for the components of the LOD includes separate corresponding metadata items for each of the components.

26. The method of claim 25, wherein each of the separate corresponding metadata items in the separate corresponding metadata items includes a corresponding Basic Model Component (BMCP) box, the corresponding Basic Model Component (BMCP) box including an extension of the FullBox of ISOBMFF.

27. The method of claim 25, wherein each of the components includes a type value representing the type of the component and an encoded value representing how the component is encoded.

28. The method of claim 20, wherein one or more of the components include an encryption component, and the method further comprises: Retrieve the decryption key for each of the encryption components; as well as Use the decryption key to decrypt the data of the encrypted component.

29. The method of claim 20, wherein the data representing the component includes a hierarchical representation of the component, such that each component having a parent component includes data identifying the parent component.

30. A client device for retrieving media data, the device comprising: The memory is used to store media data; and A processing system, implemented in a circuit and configured to: Retrieve data representing a three-dimensional (3D) object model stored in an ISO Basic Media File Format (ISO BMFF) file and one or more levels of detail (LOD) of the 3D object model stored in the ISO BMFF file, wherein the ISO BMFF file is stored on a server device. A request is sent to the server device to access data of the 3D object model in one LOD (Level of Detail), wherein the data of the 3D object model in that LOD includes the size, complexity, and components of the 3D object model for that LOD; and In response to the request, the data of the 3D object model in the LOD of the LOD is received, the data of the 3D object model in the LOD of the LOD having the size, complexity and components of the 3D object model in the LOD of the LOD.