Using ISO BMFF files to save 3D object model data
Patent Information
- Application Number
- KR1020267020258
- Authority / Receiving Office
- KR · KR
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-01-02
- Filing Date
- 2025-01-03
- Publication Date
- 2026-09-04
Smart Images

Figure PCT00010_ABST
Abstract
Description
Technology Field
[0001] The present application claims priority to U.S. Patent Application No. 19 / 008,368 filed January 2, 2025 and U.S. Provisional Application No. 63 / 617,715 filed January 4, 2024, the entire contents of each of these applications incorporated by reference. U.S. Patent Application No. 19 / 008,368 filed January 2, 2025 claims the benefit of U.S. Provisional Application No. 63 / 617,715 filed January 4, 2024.
[0002] Technology field
[0003] The present disclosure relates to the storage and transmission of encoded video data. Background Technology
[0004] Digital video capabilities can be integrated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptops or desktop computers, digital cameras, digital recording devices, digital media players, video gaming devices, video game consoles, cellular or satellite radio phones, video teleconferencing devices, etc. Digital video devices implement video compression techniques, such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263 or ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), and ITU-T H.265 (also referred to as High Efficiency Video Coding (HEVC)), and extensions of such standards, in order to transmit and receive digital video information more efficiently.
[0005] Video compression techniques perform spatial prediction and / or temporal prediction to reduce or eliminate redundancy inherent to video sequences. For block-based video coding, a video frame or slice can be partitioned into macroblocks. Each macroblock can be further partitioned. Macroblocks in an intra-coded (I) frame or slice are encoded using spatial prediction for neighboring macroblocks. Macroblocks in an inter-coded (P or B) frame or slice may use spatial prediction for neighboring macroblocks in the same frame or slice, or temporal prediction for other reference frames.
[0006] After video data is encoded, the video data can be packetized for transmission or storage. The video data can be assembled into a video file conforming to any of the various standards, such as ISO (International Organization for Standardization) BMFF (base media file format) and its extensions, such as AVC.
[0007] Generally, the present disclosure describes techniques for using ISO Base Media File Format (ISO BMFF) files to store 3D object models, such as user avatars associated with an extended reality (XR) session. For example, various levels of detail (LODs) of a 3D object model may be stored, each containing components of the model within that LOD. In this way, the base model can be modified using animation streams distributed to users associated with the XR session. That is, when a user of a device (e.g., a user equipment (UE) device) joins an XR session with other users, the user's device may share the base 3D object model with the other users and then transmit animation streams to animate the 3D object model.
[0008] The present disclosure relates to techniques for efficiently storing base models for rapid deployment and adaptation to the network connectivity and processing capabilities of each participant. In particular, a device may determine a specific LOD for a 3D object model to retrieve based, for example, the distance from a 3D scene to a 3D object, available network bandwidth, etc., and then retrieve the determined LOD for the 3D object model. In this way, relatively lower LOD models may be retrieved when the 3D object is relatively far from the current user, whereas relatively higher LOD models may be retrieved when the 3D object is relatively close to the current user. Thus, relatively lower LOD models can be quickly accessed and processed for objects that do not require a large amount of detail (e.g., because they are far from the current user), and thus, processing resources can be prioritized for relatively higher LOD models for 3D objects close to the current user. Therefore, these techniques can reduce the processing operations and bandwidth required for 3D objects that need to be rendered, while maintaining a good experience for the user because nearby 3D objects can be rendered at a higher quality level.
[0009] In one example, a method for storing media data includes the step of storing a three-dimensional object model in an ISO base media file format (ISO BMFF) file, wherein the storing step includes: storing metadata for the three-dimensional object model in the ISO BMFF file; storing a number of levels of detail (LODs) for the three-dimensional object model in the ISO BMFF file; and for each number of LODs, storing data in the ISO BMFF file that associates the LOD with data representing the size, complexity, and components of the three-dimensional object model for the LOD.
[0010] In another example, a device for storing media data comprises: a memory for storing media data; and a processing system implemented in circuitry, wherein the processing system is configured to store a three-dimensional object model in an ISO base media file format (ISO BMFF) file of media data in memory, and to store the three-dimensional object model in the ISO BMFF file, the processing system is configured to: store metadata for the three-dimensional object model in the ISO BMFF file; store a number of levels of detail (LOD) for the three-dimensional object model in the ISO BMFF file; and for each number of LODs, store data in the ISO BMFF file that associates the LOD with data representing the size, complexity, and components of the three-dimensional object model for the LOD.
[0011] In another example, the method for retrieving media data is:
[0012] The method comprises the steps of: retrieving, by a client device, data representing a three-dimensional object model stored in an ISO BMFF (ISO base media file format) file and one or more levels of detail (LODs) for the three-dimensional object model stored in the ISO BMFF file, wherein the ISO BMFF file is stored on a server device; transmitting to the server device a request to access data for the three-dimensional object model in one of the LODs, wherein the data for the three-dimensional object model in one of the LODs includes the size, complexity, and components of the three-dimensional object model in one of the LODs; and receiving, by the client device in response to the request, data for the three-dimensional object model in one of the LODs, wherein the data for the three-dimensional object model in one of the LODs includes the size, complexity, and components of the three-dimensional object model in one of the LODs.
[0013] In another example, a device for retrieving media data comprises: memory for storing media data; and a processing system implemented in circuitry, wherein the processing system retrieves data representing a three-dimensional object model and one or more levels of detail (LODs) for the three-dimensional object model; transmits a request to a server device for accessing data for the three-dimensional object model in one of the LODs, wherein the data for the three-dimensional object model in one of the LODs includes the size, complexity, and components of the three-dimensional object model in one of the LODs; and is configured to receive, in response to the request, data for the three-dimensional object model in one of the LODs, wherein the data for the three-dimensional object model in one of the LODs includes the size, complexity, and components of the three-dimensional object model in one of the LODs.
[0014] Details of one or more examples are described in the attached drawings and the description below. Other features, purposes, and advantages will be apparent from the description, drawings, and claims. Brief explanation of the drawing
[0015] Figure 1 is a block diagram illustrating an exemplary system for implementing techniques for streaming media data over a network. Figure 2 is a block diagram illustrating the elements of an exemplary video file. Figure 3 is a flowchart illustrating an exemplary avatar animation workflow that can be used during an XR session. Figure 4 is a flowchart illustrating an exemplary XR session between two UE (user equipment) devices and a shared space server device. FIG. 5 is a block diagram illustrating an exemplary 3D object model structure that can be stored in an ISO BMFF (ISO Base Media File Format) file according to the techniques of the present disclosure. FIG. 6 is a block diagram illustrating an exemplary basic model mapping structure that may be included in the 3D object model of FIG. 5 according to the techniques of the present disclosure. FIG. 7 is a block diagram illustrating an exemplary basic model component structure that may be included in the basic model mapping structure of FIG. 6 according to the techniques of the present disclosure. FIG. 8 is a flowchart illustrating an exemplary method that can be used to store, transmit, and retrieve data of a 3D object model structure according to the techniques of the present disclosure. FIG. 9 is a flowchart illustrating an exemplary method for generating an ISO BMFF (ISO base media file format) file containing a base model of a 3D object at various levels of detail (LOD) according to the techniques of the present disclosure. FIG. 10 is a flowchart illustrating an exemplary method for retrieving 3D object media data according to the techniques of the present disclosure. Specific details for implementing the invention
[0016] Generally, the present disclosure describes techniques for storing and transmitting three-dimensional object model data, such as avatar model data, that can be used in an extended reality (XR) session. The XR may be an augmented reality (AR), mixed reality (MR), or virtual reality (VR) session. An XR session may take place between two or more participants, each using individual XR-enabled devices, such as user equipment (UE) devices.
[0017] Immersive XR experiences are typically based on shared virtual spaces where people (represented by their avatars) join in and interact with each other and the environment. Avatars can be realistic or "cartoon-like" representations of the user. Avatars can be animated to mimic the user's body poses and facial expressions.
[0018] The present disclosure generally describes the storage and transmission of 3D object models, such as user avatars associated with an XR session. For example, various levels of detail (LODs) of a 3D object model may be stored, each containing, for each LOD, components of the model within that LOD. In this way, the base model can be distributed to users associated with the XR session and then modified using animation streams. That is, when a user of a device (e.g., a user equipment (UE) device) joins an XR session with other users, the user's device may share the base 3D object model with the other users and then transmit animation streams to animate the 3D object model. The present disclosure generally relates to techniques for efficiently storing the base model for rapid distribution and adaptation to the network connectivity and processing capabilities of each participant. The present disclosure also describes techniques that enable data representing the 3D object model to be transmitted (and retrieved by) to the devices of other users associated with the XR session. Although generally described in relation to the storage and transmission of user avatars, these techniques can also generally be applied to other animable 3D objects, such as non-user avatars, animable or movable objects within an XR scene.
[0019] For example, a user of an XR device, such as an XR headset, may participate in an XR communication session with one or more other users during, for example, an XR conference call or an XR game. Each user may have an individual avatar to be depicted to represent the user in the XR communication session. The avatar may be presented at a specific position within 3D space. Each user may control their position within 3D space by physically moving (which can be tracked by the XR device) or by providing movement input using controllers or other input devices.
[0020] Generally, users will only be able to perceive high-level-of-detail (LOD) models of avatars or other 3D objects within 3D space when, for example, the 3D objects are close to the user. Therefore, instead of retrieving high-level-of-detail (LOD) models of 3D objects relatively far from the user in 3D space, the XR device can retrieve lower-level-of-detail (LOD) models. In this way, bandwidth can be saved when retrieving lower-level-of-detail (LOD) models, and processing requirements for rendering lower-level-of-detail (LOD) models can be reduced. Consequently, such bandwidth and processing resources can be allocated to higher-level-of-detail (LOD) models of 3D objects relatively close to the user in 3D space. Available network bandwidth can also be used to determine the LODs of the 3D models, ensuring that, for example, the data required for each model can be retrieved in a timely manner.
[0021] After retrieving LODs for each model, the XR device can receive information indicating the animation stream to be applied to the model. In this way, the animation stream can remain agnostic to the LODs for the model. That is, the same animation stream can be used for each of the model's LODs.
[0022] The techniques of the present disclosure can generally improve the efficiency of storing, transmitting, and rendering 3D object models. The techniques of the present disclosure can also allow different representations of 3D object models to be transmitted to different users to accommodate, for example, various network connections, available bandwidth, and processing capabilities.
[0023] The techniques of the present disclosure may be applied to video files corresponding to video data encapsulated according to ISO BMFF (base media file format) or its extensions. The techniques of the present disclosure may also be applied to other file formats such as SVC (Scalable Video Coding) file format, AVC (Advanced Video Coding) file format, 3GPP (Third Generation Partnership Project) file format, and / or MVC (Multiview Video Coding) file format, or other similar video file formats.
[0024] FIG. 1 is a block diagram illustrating an exemplary system (10) for implementing techniques for streaming media data over a network. In this example, the system (10) includes a content preparation device (20), a server device (60), and a client device (40). The client device (40) and the server device (60) are coupled to communicate with each other by a network (74) which may include the Internet. In some examples, the content preparation device (20) and the server device (60) may also be coupled to the network (74) or another network, or may be coupled to communicate directly. In some examples, the content preparation device (20) and the server device (60) may include the same device.
[0025] The content preparation device (20) includes, in the example of FIG. 1, an audio source (22) and a video source (24). The audio source (22) may include, for example, a microphone that generates electrical signals representing captured audio data to be encoded by an audio encoder (26). Alternatively, the audio source (22) may include a storage medium that stores previously recorded audio data, an audio data generator such as a computerized synthesizer, or any other source of audio data. The video source (24) may include a video camera that generates video data to be encoded by a video encoder (28), a storage medium encoded with previously recorded video data, a video data generation unit such as a computer graphics source, or any other source of video data. In all examples, the content preparation device (20) does not necessarily need to be communicateably coupled to the server device (60), but may store multimedia content on a separate medium read by the server device (60).
[0026] Raw audio and video data may include analog or digital data. Analog data may be digitized before being encoded by an audio encoder (26) and / or a video encoder (28). An audio source (22) may acquire audio data from a speaking participant while the speaking participant is speaking, and a video source (24) may acquire video data of the speaking participant simultaneously. In other examples, the audio source (22) may include a computer-readable storage medium containing stored audio data, and the video source (24) may include a computer-readable storage medium containing stored video data. In this way, the techniques described in this disclosure may be applied to live, streaming, real-time audio and video data or to archived, pre-recorded audio and video data.
[0027] Audio frames corresponding to video frames are generally audio frames that contain audio data captured (or generated) by an audio source (22) simultaneously with video data captured (or generated) by a video source (24) that is included within the video frames. For example, while a speaking participant generates audio data by generally speaking, the audio source (22) captures the audio data, and the video source (24) simultaneously captures the video data of the speaking participant, that is, while the audio source (22) is capturing the audio data. Thus, an audio frame may correspond temporally to one or more specific video frames. Thus, an audio frame corresponding to a video frame generally corresponds to a situation where audio data and video data are captured simultaneously, and the audio frame and video frame each contain the audio data and video data that were captured simultaneously.
[0028] In some examples, the audio encoder (26) may encode a timestamp in each encoded audio frame representing the time at which audio data for the encoded audio frame was recorded, and similarly, the video encoder (28) may encode a timestamp in each encoded video frame representing the time at which video data for the encoded video frame was recorded. In these examples, the audio frame corresponding to the video frame may include an audio frame containing a timestamp and a video frame containing the same timestamp. The content preparation device (20) may include an internal clock that enables the audio encoder (26) and / or the video encoder (28) to generate timestamps or that the audio source (22) and the video source (24) can use to associate audio and video data with timestamps, respectively.
[0029] In some examples, the audio source (22) may transmit data corresponding to the time at which the audio data was recorded to the audio encoder (26), and the video source (24) may transmit data corresponding to the time at which the video data was recorded to the video encoder (28). In some examples, the audio encoder (26) may encode a sequence identifier into the encoded audio data, which indicates the relative time ordering of the encoded audio data but does not necessarily indicate the absolute time at which the audio data was recorded, and similarly, the video encoder (28) may also use the sequence identifiers to indicate the relative time ordering of the encoded video data. Similarly, in some examples, the sequence identifier may be mapped to a timestamp or otherwise correlated with a timestamp.
[0030] The audio encoder (26) generally generates a stream of encoded audio data, while the video encoder (28) generates a stream of encoded video data. Each individual stream of data (whether audio or video) may be referred to as an elementary stream. An elementary stream is a single digitally coded (possibly compressed) component of a media presentation. For example, a coded video or audio part of a media presentation may be an elementary stream. An elementary stream may be converted into a packetized elementary stream (PES) before being encapsulated within a video file. Within the same media presentation, a stream ID may be used to distinguish PES packets belonging to one elementary stream from others. The basic unit of data in an elementary stream is a packet of the packetized elementary stream (PES). Thus, coded video data generally corresponds to elementary video streams. Similarly, audio data corresponds to one or more individual elementary streams.
[0031] In the example of FIG. 1, the encapsulation unit (30) of the content preparation device (20) receives base streams containing coded video data from the video encoder (28) and base streams containing coded audio data from the audio encoder (26). In some examples, the video encoder (28) and the audio encoder (26) may each include packetizers for forming PES packets from the encoded data. In other examples, the video encoder (28) and the audio encoder (26) may each interface with individual packetizers for forming PES packets from the encoded data. In yet other examples, the encapsulation unit (30) may include packetizers for forming PES packets from the encoded audio and video data.
[0032] A video encoder (28) can encode video data of multimedia content in various ways to generate different representations of multimedia content having various bit rates and various characteristics, such as pixel resolutions, frame rates, compliance with various coding standards, compliance with various profiles and / or levels of profiles for various coding standards, representations having one or more views (e.g., for 2D or 3D playback), or other such characteristics. A representation may include audio data, video data, text data (e.g., for closed captions), or any of other such data, as used in the present disclosure. A representation may include a base stream, such as an audio base stream or a video base stream. Each PES packet may include a stream_id that identifies the base stream to which the PES packet belongs. An encapsulation unit (30) is responsible for assembling the base streams into streamable media data.
[0033] The encapsulation unit (30) receives PES packets for base streams of media presentation from the audio encoder (26) and the video encoder (28) and forms corresponding network abstraction layer (NAL) units from the PES packets. Coded video segments can be organized into NAL units, which provide "network-friendly" video representations for applications such as video telephony, storage, broadcasting, or streaming. NAL units can be classified into Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL units may include a core compression engine and may include block, macroblock, and / or slice-level data. Other NAL units may be non-VCL NAL units. In some examples, a coded picture at a time instance, usually presented as a primary coded picture, may be contained in an access unit that may include one or more NAL units.
[0034] Non-VCL NAL units may include, in particular, parameter set NAL units and SEI NAL units. Parameter sets may include sequence-level header information (in sequence parameter sets (SPS)) and rarely changing picture-level header information (in picture parameter sets (PPS)). In parameter sets (e.g., PPS and SPS), rarely changing information does not need to be repeated for each sequence or picture; thus, coding efficiency can be improved. Furthermore, the use of parameter sets enables out-of-band transmission of important header information, thereby avoiding the need for redundant transmissions for error resilience. In out-of-band transmission examples, parameter set NAL units may be transmitted on a different channel from other NAL units, such as SEI NAL units.
[0035] Supplemental Enhancement Information (SEI) may contain information that assists processes related to decoding, display, error tolerance, and other purposes, even though it is not necessary to decode picture samples coded from VCL NAL units. SEI messages may be included in non-VCL NAL units. SEI messages are a normative part of some standard specifications and are therefore not always mandatory for standard-compliant decoder implementations. SEI messages may be sequence-level SEI messages or picture-level SEI messages. Some sequence-level information may be included in SEI messages, such as scalability information SEI messages in the example of SVC and view scalability information SEI messages in MVC. These exemplary SEI messages may convey information, for example, regarding the extraction of action points and the characteristics of action points.
[0036] The server device (60) includes a Real-time Transport Protocol (RTP) transmission unit (70) and a network interface (72). In some examples, the server device (60) may include a plurality of network interfaces. Furthermore, any or all of the features of the server device (60) may be implemented on other devices of the content delivery network, such as routers, bridges, proxy devices, switches, or other devices. In some examples, intermediate devices of the content delivery network may cache data of multimedia content (64) and may include components substantially corresponding to the components of the server device (60). Generally, the network interface (72) is configured to transmit and receive data through the network (74).
[0037] The RTP transmission unit (70) is configured to transmit media data to a client device (40) over a network (74) in accordance with RTP, which is standardized in Request for Comment (RFC) 3550 by the Internet Engineering Task Force (IETF). The RTP transmission unit (70) may also implement RTP-related protocols such as the RTP Control Protocol (RTCP), Real-time Streaming Protocol (RTSP), Session Initiation Protocol (SIP), and / or Session Description Protocol (SDP). The RTP transmission unit (70) may transmit media data through a network interface (72), which may implement the Uniform Datagram Protocol (UDP) and / or the Internet Protocol (IP). Thus, in some examples, a server device (60) may transmit media data over RTP and RTSP using the network (74).
[0038] The RTP transmission unit (70) may receive an RTSP description request from, for example, a client device (40). The RTSP description request may include data indicating which types of data are supported by the client device (40). The RTP transmission unit (70) may respond to the client device (40) with data indicating media streams, such as media content (64), that can be transmitted to the client device (40), along with a corresponding network location identifier such as a URL (uniform resource locator) or a URN (uniform resource name).
[0039] Next, the RTP transmission unit (70) may receive an RTSP setup request from the client device (40). The RTSP setup request may generally indicate how the media stream should be transmitted. The RTSP setup request may include a network location identifier for the requested media data (e.g., media content (64)) and a transport identifier, such as a local port for receiving RTP data and control data (e.g., RTCP data) on the client device (40). The RTP transmission unit (70) may respond to the RTSP setup request with data and acknowledgments representing the ports of the server device (60) to which the RTP data and control data will be transmitted. Next, the RTP transmission unit (70) may receive an RTSP playback request to cause the media stream to be "played," that is, transmitted to the client device (40) over the network (74). The RTP transmission unit (70) may also receive an RTSP teardown request to terminate a streaming session, and in response to this, the RTP transmission unit (70) may stop transmitting media data to the client device (40) for the corresponding session.
[0040] The RTP receiving unit (52) can likewise initiate a media stream by initially sending an RTSP description request to the server device (60). The RTSP description request may indicate the types of data supported by the client device (40). Then, the RTP receiving unit (52) can receive from the server device (60) a reply specifying available media streams, such as media content (64), that can be sent to the client device (40), along with a corresponding network location identifier such as a URL (uniform resource locator) or URN (uniform resource name).
[0041] Next, the RTP receiving unit (52) may generate an RTSP setup request and transmit the RTSP setup request to the server device (60). As mentioned above, the RTSP setup request may include a transport specifier, such as a network location identifier for the requested media data (e.g., media content (64)) and local ports for receiving RTP data and control data (e.g., RTCP data) on the client device (40). In response, the RTP receiving unit (52) may receive an acknowledgment from the server device (60) including ports of the server device (60) that the server device (60) will use to transmit the media data and control data.
[0042] After establishing a media streaming session between the server device (60) and the client device (40), the RTP transmission unit (70) of the server device (60) can transmit media data (e.g., packets of media data) to the client device (40) according to the media streaming session. The server device (60) and the client device (40) can exchange control data (e.g., RTCP data) that displays reception statistics by the client device (40), for example, so that the server device (60) can perform congestion control or otherwise diagnose and handle transmission faults.
[0043] A network interface (54) may receive media of a selected media presentation and provide it to an RTP receiving unit (52), which in turn may provide the media data to a decapsulation unit (50). The decapsulation unit (50) may decapsulate the elements of a video file into constituent PES streams, depacketize the PES streams to retrieve the encoded data, and transmit the encoded data to either an audio decoder (46) or a video decoder (48), depending on whether the encoded data is part of an audio stream or part of a video stream, for example, as indicated by the PES packet headers of the stream. The audio decoder (46) decodes the encoded audio data and transmits the decoded audio data to an audio output (42), whereas the video decoder (48) decodes the encoded video data and transmits the decoded video data, which may include multiple views of the stream, to a video output (44).
[0044] The video encoder (28), video decoder (48), audio encoder (26), audio decoder (46), encapsulation unit (30), RTP receiving unit (52), and decapsulation unit (50) may each be implemented as various suitable processing circuits where applicable, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuits, software, hardware, firmware, or any combination thereof. Each of the video encoder (28) and video decoder (48) may be included in one or more encoders or decoders, any one of which may be integrated as part of a combined video encoder / decoder (CODEC). Likewise, each of the audio encoder (26) and audio decoder (46) may be included in one or more encoders or decoders, any one of which may be integrated as part of a combined CODEC. A device comprising a video encoder (28), a video decoder (48), an audio encoder (26), an audio decoder (46), an encapsulation unit (30), an RTP receiving unit (52), and / or a decapsulation unit (50) may include an integrated circuit, a microprocessor, and / or a wireless communication device, such as a cellular telephone.
[0045] The client device (40), the server device (60), and / or the content preparation device (20) may be configured to operate according to the techniques of the present disclosure. For the purposes of example, the present disclosure describes these techniques with respect to the client device (40) and the server device (60). However, it should be understood that the content preparation device (20) may be configured to perform these techniques instead of (or in addition to) the server device (60).
[0046] The encapsulation unit (30) may form NAL units that include a header identifying the program to which the NAL unit belongs, as well as a payload, such as audio data, video data, or data describing the transmission or program stream corresponding to the NAL unit. For example, in H.264 / AVC, the NAL unit includes a 1-byte header and a variable-size payload. A NAL unit that includes video data in its payload may include video data of various granularity levels. For example, the NAL unit may include a block of video data, multiple blocks, a slice of video data, or an entire picture of video data. The encapsulation unit (30) may receive encoded video data from the video encoder (28) in the form of PES packets of base streams. The encapsulation unit (30) may associate each base stream with a corresponding program.
[0047] The encapsulation unit (30) may also assemble access units from a plurality of NAL units. Generally, the access unit may include one or more NAL units for representing audio data corresponding to a frame when audio data is available as well as frames of video data. The access unit generally includes all NAL units for one output time instance, e.g., all audio and video data for one time instance. For example, if each view has a frame rate of 20 fps (frames per second), each time instance may correspond to a time interval of 0.05 seconds. During this time interval, specific frames for all views of the same access unit (same time instance) may be rendered simultaneously. In one example, the access unit may include a picture coded in one time instance, which may be presented as a primary coded picture.
[0048] Therefore, the access unit includes all audio and video frames of a common time instance, for example, time X It may include all views corresponding to it. The present disclosure also refers to an encoded picture of a specific view as a "view component." That is, a view component may include an encoded picture (or frame) of a specific view at a specific time. Accordingly, an access unit may be defined as including all view components of a common time instance. The decoding order of the access units does not necessarily have to be the same as the output or display order.
[0049] After the encapsulation unit (30) assembles the NAL units and / or access units into a video file based on the received data, the encapsulation unit (30) transmits the video file to the output interface (32) for output. In some examples, the encapsulation unit (30) may store the video file locally, or transmit the video file to a remote server via the output interface (32) rather than transmitting the video file directly to the client device (40). The output interface (32) may include, for example, a transmitter, a transceiver, a device for writing data to a computer-readable medium such as an optical drive, a magnetic media drive (e.g., a floppy drive), a USB (universal serial bus) port, a network interface, or other output interfaces. The output interface (32) outputs the video file to a computer-readable medium such as a transmission signal, a magnetic medium, an optical medium, a memory, a flash drive, or other computer-readable medium.
[0050] The network interface (54) may receive a NAL unit or access unit through the network (74) and provide the NAL unit or access unit to the decapsulation unit (50) through the RTP receiving unit (52). The decapsulation unit (50) may decapsulate the elements of the video file into constituent PES streams, depacketize the PES streams to retrieve the encoded data, and transmit the encoded data to either an audio decoder (46) or a video decoder (48), depending on whether the encoded data is part of an audio stream or part of a video stream, for example, as indicated by the PES packet headers of the stream. The audio decoder (46) decodes the encoded audio data and transmits the decoded audio data to the audio output (42), while the video decoder (48) decodes the encoded video data and transmits the decoded video data, which may include multiple views of the stream, to the video output (44).
[0051] FIG. 2 is a block diagram illustrating the elements of an exemplary video file (150). As described above, video files according to ISO BMFF (base media file format) and its extensions store data in a series of objects referred to as "boxes." In the example of FIG. 2, the video file (150) includes a file type (FTYP) box (152), a movie (MOOV) box (154), segment index (sidx) boxes (162), movie fragment (MOOF) boxes (164), and a movie fragment random access (MFRA) box (166). Although FIG. 2 illustrates an example of a video file, it should be understood that other media files may include other types of media data (e.g., audio data, timing data, etc.) that are structured similarly to the data of the video file (150) according to ISO BMFF and its extensions.
[0052] The file type (FTYP) box (152) generally describes the file type for the video file (150). The file type box (152) may contain data identifying a specification describing the best use for the video file (150). Alternatively, the file type box (152) may be placed before the MOOV box (154), movie fragment boxes (164), and / or MFRA box (166).
[0053] In the example of FIG. 2, the MOOV box (154) includes a Movie Header (MVHD) box (156), a Track (TRAK) box (158), and one or more Movie Extension (MVEX) boxes (160). Generally, the MVHD box (156) can describe general characteristics of the video file (150). For example, the MVHD box (156) may include data describing the time scale of the video file (150), the duration of playback of the video file (150), or other data describing the video file (150) in general, when the video file (150) was originally created, when the video file (150) was last modified.
[0054] The TRAK box (158) may contain data for a track of the video file (150). The TRAK box (158) may contain a track header (TKHD) box describing the characteristics of the track corresponding to the TRAK box (158). In some examples, the TRAK box (158) may contain coded video pictures, whereas in other examples, the coded video pictures of the track may be contained in movie fragments (164) that can be referenced by data in the TRAK box (158) and / or sidx boxes (162).
[0055] In some examples, the video file (150) may contain more than one track. Accordingly, the MOOV box (154) may contain a number of TRAK boxes equal to the number of tracks in the video file (150). The TRAK box (158) may describe the characteristics of the corresponding track in the video file (150). For example, the TRAK box (158) may describe temporal and / or spatial information for the corresponding track. When the encapsulation unit (30) (FI. 1) includes a parameter set track in the video file, e.g., the video file (150), a TRAK box similar to the TRAK box (158) of the MOOV box (154) may describe the characteristics of the parameter set track. The encapsulation unit (30) may signal the presence of sequence-level SEI messages in the parameter set track within the TRAK box describing the parameter set track.
[0056] MVEX boxes (160) may describe the characteristics of corresponding movie fragments (164) to signal that, for example, if a video file (150) exists, in addition to the video data contained within the MOOV box (154). In the context of streaming video data, coded video pictures may be contained in the movie fragments (164) rather than in the MOOV box (154). Thus, all coded video samples may be contained in the movie fragments (164) rather than in the MOOV box (154).
[0057] The MOOV box (154) may contain the same number of MVEX boxes (160) as the number of movie fragments (164) in the video file (150). Each MVEX box (160) may describe the characteristics of the corresponding movie fragment among the movie fragments (164). For example, each MVEX box may include a MEHD (movie extends header box) box describing the time duration of the corresponding movie fragment among the movie fragments (164).
[0058] As mentioned above, the encapsulation unit (30) may store a sequence data set that does not contain actual coded video data in a video sample. The video sample may generally correspond to an access unit, which is a representation of the coded picture at a specific time instance. In the context of AVC, the coded picture includes one or more VCL NAL units containing information for configuring all pixels of the access unit, and other associated non-VCL NAL units, such as SEI messages. Thus, the encapsulation unit (30) may include a sequence data set, which may contain sequence-level SEI messages in one of the movie fragments (164). The encapsulation unit (30) may further signal the presence of the sequence data set and / or sequence-level SEI messages as existing in one of the movie fragments (164) within one of the MVEX boxes (160) corresponding to one of the movie fragments (164).
[0059] SIDX boxes (162) are optional elements of a video file (150). That is, video files conforming to the 3GPP file format, or other such file formats, do not necessarily contain SIDX boxes (162). According to the example of the 3GPP file format, SIDX boxes can be used to identify sub-segments of a segment (e.g., a segment contained within the video file (150)). The 3GPP file format defines a sub-segment as “a standalone set of one or more consecutive movie fragment boxes having corresponding media data box(s) and a media data box containing data referenced by a movie fragment box must follow that movie fragment box and precede the next movie fragment box containing information about the same track.” The 3GPP file format also indicates that the SIDX box "contains a sequence of references to the subsegments of the (sub)segment documented by the box. The referenced subsegments are contiguous at presentation time. Similarly, the bytes referenced by the segment index box are always contiguous within the segment. The referenced size provides a count of the number of bytes in the referenced material."
[0060] SIDX boxes (162) generally provide information representing one or more sub-segments of a segment included in a video file (150). For example, such information may include playback times when the sub-segments start and / or end, byte offsets for the sub-segments, whether the sub-segments contain a stream access point (SAP) (e.g., whether they start with a SAP), the type of the SAP (e.g., whether the SAP is an instantaneous decoder refresh (IDR) picture, a clean random access (CRA) picture, a broken link access (BLA) picture, etc.), the position of the SAP in the sub-segment (in terms of playback time and / or byte offset), etc.
[0061] Movie fragments (164) may include one or more coded video pictures. In some examples, movie fragments (164) may include one or more group of pictures (GOPs), each of which may include multiple coded video pictures, such as frames or pictures. Additionally, as described above, movie fragments (164) may include sequence data sets in some examples. Each movie fragment (164) may include a movie fragment header box (MFHD) (not shown in FIG. 2). The MFHD box may describe the characteristics of the corresponding movie fragment, such as the sequence number for the movie fragment. Movie fragments (164) may be included in the video file (150) in the order of sequence numbers.
[0062] An MFRA box (166) can describe random access points within movie fragments (164) of a video file (150). This can assist in performing trick modes, such as performing searches for specific time locations (i.e., playback times) within a segment encapsulated by the video file (150). An MFRA box (166) is generally optional and, in some examples, does not need to be included in video files. Likewise, a client device, such as a client device (40), does not necessarily need to refer to the MFRA box (166) to accurately decode and display video data of the video file (150). An MFRA box (166) may include a number of TFRA (track fragment random access) boxes (not shown) equal to the number of tracks of the video file (150), or, in some examples, equal to the number of media tracks (e.g., non-hint tracks) of the video file (150).
[0063] In some examples, movie fragments (164) may include one or more stream access points (SAPs), such as IDR pictures. Likewise, an MFRA box (166) may provide indications of the locations of the SAPs within the video file (150). Thus, a time sub-sequence of the video file (150) may be formed from the SAPs of the video file (150). The time sub-sequence may also include other pictures, such as P-frames and / or B-frames, that depend on the SAPs. Frames and / or slices of the time sub-sequence may be arranged within segments so that frames / slices of the time sub-sequence that depend on other frames / slices of the sub-sequence can be properly decoded. For example, in a hierarchical arrangement of data, data used for prediction regarding other data may also be included in the time sub-sequence.
[0064] FIG. 3 is a flowchart illustrating an exemplary avatar animation workflow that can be used during an XR session. In this example, the received animation stream data (170) includes face blend shapes, body blend shapes, hand joints, head poses, and audio stream data. The face blend shapes, body blend shapes, and hand joints may correspond to an animation stream to be applied to the User A avatar base model (172). In particular, data for the User A avatar base model (172) may be stored at various levels of detail (LODs) according to the techniques of the present disclosure. Thus, rendering components (174) may retrieve data of the User A avatar base model (172) at an appropriate level of detail (LOD), for example, based on the distance between the current user and User A in 3D space. Then, the rendering components (174) may animate the avatar base model using the received animation stream data (170). Ultimately, the animated avatar base model can be presented to the current user through the display (176). Additionally, the current user's movement data can be used by the future pose prediction unit (178) to predict the user's future pose.
[0065] FIG. 4 is a flowchart illustrating an exemplary XR session between two UE (user equipment) devices and a shared space server device. As illustrated in the example of FIG. 4, two or more UEs may participate in the XR session. The UEs may receive data representing their animation streams and other 3D model data from the shared space server and transmit it to the shared space server. For example, various sensors such as cameras, trackers, LIDAR, etc., may track user movements such as facial movements (e.g., while speaking or as emotional responses), hand movements, walking movements, etc. These movements may be translated into animation streams by, for example, the UE (182) and transmitted to the shared space server. Then, the shared space server may transmit the animation streams to the UE (184).
[0066] The UE (184) can determine the distance between the user of the UE (184) in 3D space for the XR session and the user of the UE (182) in 3D space for the XR session. User positions can be represented by 3D vectors relative to the origin in 3D space. Thus, the user of the UE (182) may be at position {x1, y1, z1} in 3D space, and the user of the UE (184) may be at position {x2, y2, z2} in 3D space. The distance between the user of the UE (182) and the user of the UE (184) can be calculated as follows:
[0067] Distance =
[0068] The UE (184) can use this distance to determine the appropriate LOD for the 3D object model of the UE (182)'s user's avatar in response to a request from the shared space server. For example, the UE (184) may be composed of various thresholds corresponding to each available LOD for the 3D object model of the avatar. Thus, the UE (184) can determine which threshold is satisfied by the determined distance and then request the LOD for the 3D object model of the UE (182)'s user's avatar corresponding to this threshold. Then, the UE (184) can animate the retrieved version of the 3D object model having the requested LOD using the animation streams received from the UE (182) through the shared space server.
[0069] FIG. 5 is a block diagram illustrating an exemplary 3D object model structure (200) that can be stored in an ISO Base Media File Format (ISO BMFF) file according to the techniques of the present disclosure. For example, the 3D object model structure (200) can be stored in the video file (150) of FIG. 2. Generally, the present disclosure describes techniques for storing a 3D object model, such as an avatar base model that can be used during an XR session, in an ISO BMFF file. The present disclosure also describes techniques for retrieving 3D object model data from an ISO BMFF file (or from a device that stores 3D object model data in an ISO BMFF file).
[0070] Storage of a 3D object model within an ISO BMFF file according to the techniques of the present disclosure may allow for the extraction of relevant representations based on desired target sizes and complexities. For example, the 3D object model may be stored at various different levels of detail (LODs), which may have different complexities, sizes, etc. In this way, for example, if a user is close to a 3D object represented by a 3D object model in a virtual scene, the user's device may retrieve a higher LOD representation of the 3D object, whereas if the user is relatively far from the 3D object in the virtual scene, the user's device may retrieve a lower LOD representation of the 3D object.
[0071] According to the techniques of the present disclosure, the storage of a 3D object model within an ISO BMFF file may also allow for the encryption of all or only a subset of the components of the 3D object model. In this way, the 3D object model can be protected from copying or digital theft by users who are not authorized to access the 3D object model.
[0072] A 3D object model may include metadata containing descriptions of all components of the 3D object model (e.g., an avatar), and may provide metadata to describe the base model.
[0073] FIG. 5 depicts an exemplary 3D object model (200) comprising metadata (202), a number of LODs (levels of detail) (204), and a set of base model mappings (206). The metadata (202) may include key items containing key metadata such as name, type, gender, age, or other descriptive data for a user represented by an avatar corresponding to the 3D object model (200). The set of base model mappings (206) may include one mapping for each LOD number (204). For example, if the LOD number (204) has a value of 5, there may be 5 base model mappings (206). The base model mappings (206) may provide LODs, the resulting size and complexity for the corresponding LODs, and associations between the components of the base model in the corresponding LODs.
[0074] The following pseudocode represents an exemplary data structure corresponding to a 3D object model (200):
[0075] aligned(8) class BaseAvatarModel extends FullBox('bavm', version, flags) {
[0076] AvatarMetadataEntry();
[0077] unsigned int number_lods;
[0078] BaseModelMapping mappings[number_lods];
[0079] }
[0080] FIG. 6 is a block diagram illustrating an exemplary basic model mapping structure (210) that may be included in the 3D object model structure (200) of FIG. 5 according to the techniques of the present disclosure. That is, each of the basic model mappings (206) of FIG. 5 may include elements of the basic model mapping (210) of FIG. 6. The basic model mapping structure (210) includes an LOD identifier (ID) (212), a size (214), a number of mesh faces (216), a number of components (218), and basic model components (220). The basic model mapping (210) is associated with a level of detail (LOD) indicated by the LOD ID (212). The size (214) provides information about the size and complexity of the corresponding LOD. The number of mesh faces (216) represents the number of faces (e.g., individual planar elements) of the components of the LOD.
[0081] The number of components (218) describes the number (220) of basic model component structures included in the basic model mapping (210). The basic model components (220) represent a list of components of the basic model in the LOD. The number of basic model components (220) may be equal to the value of the number of components (218). Maps may be available at different resolutions and may be stored as image items. The basic model components (220) may be compressed. For example, geometric structure data may be compressed using mesh compression, images may be compressed using image or video compression standards such as JPEG, ITU-T H.264, ITU-T H.265, ITU-T H.266, etc., and other data may be entropy coded.
[0082] The following pseudocode represents an exemplary data structure corresponding to the basic model mapping (210):
[0083] aligned(8) class BaseModelMapping extends FullBox("bmma", version, flags) {
[0084] unsigned int(16) lod_id;
[0085] unsinged int(32) size_in_bytes;
[0086] unsigned int(32) num_mesh_faces;
[0087] unsinged int(16) num_components;
[0088] BaseModelComponent components[num_components];
[0089] }
[0090] FIG. 7 is a block diagram illustrating an exemplary basic model component structure (230) that may be included in the basic model mapping structure (210) of FIG. 6 according to the techniques of the present disclosure. That is, each of the basic model component structures (220) of FIG. 6 may include elements of the basic model component structure (230) of FIG. 7. In this example, the basic model component (230) includes an item identifier (ID) (232), a role (234), and a parent ID (236) (which is an optional element as indicated by the dotted lines in FIG. 7).
[0091] Components may be assigned roles that describe their roles to the base model, as indicated by the role (234). Examples of roles include joints, blend shapes, meshes, maps, or pose transformations. Each component may be stored individually as a metadata item. Each component may include an ItemInfoEntry describing its content_type and content_encoding indicating which compression method was applied (if applied). Each component may also be individually encrypted. Components may be arranged hierarchically and their individual parent items may be indicated using the value of the parent ID (236).
[0092] The following pseudocode represents an exemplary data structure corresponding to the basic model component (230):
[0093] aligned(8) class BaseModelComponent extends FullBox("bmcp", version, flags) {
[0094] unsigned int(16) item_ID;
[0095] Role role;
[0096] unsigned int(16) parent_item_ID; optional;
[0097] }
[0098] FIG. 8 is a flowchart illustrating an exemplary method that can be used to store, transmit, and retrieve data of a 3D object model structure according to the techniques of the present disclosure. In this example, a transmitter (e.g., a content preparation device (20) of FIG. 1 or a first UE, such as the UE (182) of FIG. 4) discloses basic model details (250), such as metadata for a 3D object model (e.g., an avatar), and available levels of detail (LODs) for the 3D object model to an application server (AS). The transmitter may disclose an avatar description including the content type of a main item, such as "application / mpeg.avatar.main". The transmitter may also disclose necessary decompression engines, information on necessary decoding, and available levels of detail (LODs). The transmitter may also store a 3D object model in a storage device, such as the server device (60) of FIG. 1, within an ISO BMFF file according to the techniques of the present disclosure (e.g., conforming to the examples of FIG. 5 through 7). That is, the transmitter may store various versions of the base model at each of the levels of detail (LODs).
[0099] A receiver (e.g., client device (40) of FIG. 1) can retrieve a description of the basic model details from the AS. The receiver can review the description and select one of the LODs (252). For example, the receiver can determine the LOD based on the receiver's user proximity (distance) to the 3D object within the virtual scene, available network bandwidth, processing capabilities, etc. The receiver can also ensure the ability to decode the components. Then, the receiver can request access to the selected LOD from the AS (254). The AS can forward the request to the transmitter. In this way, the receiver can signal the request for the selected LOD of the 3D object model to the transmitter.
[0100] Next, the transmitter can authorize the receiver to access the 3D object model in the requested LOD (256). The transmitter can redirect the request to a storage device (e.g., the server device (60) of FIG. 1). The storage device can extract the relevant representation from the ISO BMFF file based on the request and redirect the request to appropriate license servers for decryption keys.
[0101] The receiver can then request an LOD for a 3D object model (e.g., an avatar of a transmitter) from the storage device (258). The storage device can authenticate and extract a representation having the requested LOD from an ISO BMFF file (260), and transmit the extracted representation of the base model (from the requested LOD) to the receiver (262). If one or more of the components of the LOD are encrypted, the receiver can retrieve decryption keys for the encrypted components from a digital rights management (DRM) server device (262). The receiver can then decrypt the encrypted components and present the base model using the components of the LOD (264).
[0102] During an XR session between a transmitter and a receiver, the transmitter can transmit animation data to the receiver. The receiver can use the animation data to transform and / or warp a base model and display these animations to the receiver's user through a display device.
[0103] FIG. 9 is a flowchart illustrating an exemplary method for generating an ISO base media file format (ISO BMFF) file containing a base model of a 3D object at various levels of detail (LOD) according to the techniques of the present disclosure. The method of FIG. 9 may be performed by a server device such as the server device (60) of FIG. 1, a client device (40) of FIG. 1, and / or a source UE (user equipment) device such as the UE (182) of FIG. 4. The method of FIG. 9 may be performed by each UE participating in a multi-user XR communication session, and thus the method of FIG. 9 may be performed by both the UE (182) and the UE (184) of FIG. 4.
[0104] Initially, the source device obtains an initial base model of a 3D object, such as a user avatar (280). For example, the user may create a base model, download a base model, save a base model, modify an existing model, or otherwise load and / or configure a base model. The source device may then configure multiple levels of detail (LOD) versions of the base model (282). For example, the source device may configure one or more reduced LOD versions from the initial base model. To reduce the LOD, the source device may reduce texture resolutions, reduce the number of polygons in the model (e.g., the number of vertices / edges / triangles) (which may thereby reduce the number of model faces to which textures are applied), reduce the number of components in the model, or reduce the size of the model.
[0105] The source device can then generate metadata for the base model (284). The metadata can generally display information expressing the number of available LOD versions of the base model and characteristics for each LOD version. For example, the metadata can display, for each LOD version, the size of the base model for that LOD version, the number of mesh faces for that LOD version, and the number of components for that LOD version. The metadata can also include mapping data that maps the metadata for each LOD version to the corresponding LOD version, for example, using LOD identifier values, as shown in FIG. 6.
[0106] The source device can then store metadata in an ISO BMFF file (286), as well as each LOD version for the base model in an ISO BMFF file (288). The source device can also store data that associates metadata for each LOD with the corresponding LOD, for example, using LOD identifiers (290). The source device can also store the ISO BMFF file in a server device (292).
[0107] In this way, the method of FIG. 9 illustrates an example of a method for storing media data, the method comprising the step of storing a three-dimensional object model in an ISO BMFF (ISO base media file format) file, the storing step comprising: storing metadata for the three-dimensional object model in the ISO BMFF file; storing a number of levels of detail (LODs) for the three-dimensional object model in the ISO BMFF file; and for each number of LODs, storing data in the ISO BMFF file that associates the LOD with data representing the size, complexity, and components of the three-dimensional object model for the LOD.
[0108] FIG. 10 is a flowchart illustrating an exemplary method for retrieving 3D object media data according to the techniques of the present disclosure. The method of FIG. 10 may be performed by a user equipment (UE) device such as the client device (40) of FIG. 1 or the UE (184) of FIG. 4. The method of FIG. 10 may be performed by each UE participating in a multi-user XR communication session, and thus the method of FIG. 10 may be performed by both the UE (182) and the UE (184) of FIG. 4.
[0109] The client device can initially retrieve metadata for a base model of a 3D object that represents the available levels of detail (LOD) versions of the base model (300). As discussed above, the metadata may represent, for example, the size for each LOD version, the number of mesh faces for each LOD version, the number of components for each LOD version, and the mapping between the LOD versions and the attributes indicated for each LOD version, for example, using LOD identifiers for the LOD versions.
[0110] The client device can then select one of the LOD versions (302). For example, the client device can determine the distance between the user of the client device and the user to whom the 3D object model will be retrieved in 3D space. In some examples, when selecting the LOD version, available network bandwidth may also be used. Then, the client device can request a base model in the selected LOD (304). For example, the client device can request a base model from a storage device such as a central avatar repository. In response to the request, the client device can receive a base model in the selected LOD (306). In some examples, if the base model is encrypted, the client device can additionally retrieve decryption keys (308) and use the decryption keys to decrypt the base model (310).
[0111] Furthermore, as illustrated in FIG. 3, for example, while participating in an XR communication session, the client device can retrieve an animation stream for the base model (312). The animation stream may display face blend shapes, body blend shapes, hand joint movements, and / or head poses to be applied to the base model. Then, the client device can animate the base model using the animation stream (314). Ultimately, the client device can present the animated base model to the user of the client device, for example, through a display (316). The animation stream may correspond to the user's actual movements associated with the base model, such as face movements (e.g., for expressing speech), hand movements, body movements, etc.
[0112] In this way, the method of FIG. 10 illustrates an example of a method for retrieving media data, the method comprising: a step of retrieving, by a client device, data representing a three-dimensional object model stored in an ISO BMFF (ISO base media file format) file and one or more levels of detail (LODs) for the three-dimensional object model stored in the ISO BMFF file - the ISO BMFF file is stored on a server device -; a step of, by the client device, transmitting a request to the server device for accessing data for the three-dimensional object model in one of the LODs - the data for the three-dimensional object model in one of the LODs includes the size, complexity, and components of the three-dimensional object model in one of the LODs -; and a step of, by the client device, receiving data for the three-dimensional object model in one of the LODs in response to the request, wherein the data for the three-dimensional object model in one of the LODs includes the size, complexity, and components of the three-dimensional object model in one of the LODs.
[0113] In this way, the techniques of the present disclosure can allow the storage of 3D object models (e.g., avatar base models) within ISO BMFF files. These techniques can utilize the metadata structures of ISO BMFF to allow efficient storage and access to base models based on a desired level of detail (LOD).
[0114] The following provisions summarize various examples of the techniques of the present disclosure:
[0115] Clause 1: A method for storing media data comprising the step of storing a three-dimensional object model in an ISO base media file format (ISO BMFF) file, wherein the storing step comprises: storing metadata for the three-dimensional object model in the ISO BMFF file; storing a number of levels of detail (LODs) for the three-dimensional object model in the ISO BMFF file; and, for each number of LODs, storing data in the ISO BMFF file that associates the LOD with data representing the size, complexity, and components of the three-dimensional object model for the LOD.
[0116] Clause 2: The method of Clause 1, wherein the step of saving a 3D object model includes the step of saving the 3D object model as a BAVM (base avatar model) box as an extension of the ISO BMFF FullBox.
[0117] Clause 3: In Clause 1 or Clause 2, the step of storing data that associates the LOD with data representing the size, complexity, and components of a 3D object model for the LOD comprises the step of storing a basic model mapping structure in an ISO BMFF file.
[0118] Clause 4: A method in which, in any one of Clauses 1 to 3, the data representing the size, complexity, and components of a 3D object model for an LOD includes a size value, the number of mesh faces, the number of components, and for each of the number of components, data representing a corresponding component.
[0119] Clause 5: The method of Clause 4, wherein the step of storing data representing the size, complexity, and components of a 3D object model for LOD includes the step of storing a BMMA (base model mapping) box as an extension of the ISO BMFF FullBox.
[0120] Clause 6: The method of Clause 4 or Clause 5, wherein the data representing the corresponding component comprises one or more images of the component or one or more data representing the geometric structure of the component.
[0121] Clause 7: A method comprising, for each of Clauses 4 through 6, a step of storing data describing the role of the component with respect to the 3D object model.
[0122] Clause 8: In Clause 7, the role is a method comprising one or more of a joint, blend shape, mesh, map, or pose transformation.
[0123] Clause 9: In any one of Clauses 4 to 8, the step of storing data representing components of a 3D object model for an LOD comprises the step of individually storing each component as an individual metadata item in an ISO BMFF file.
[0124] Clause 10: The method of Clause 9, wherein the step of storing each component individually as an individual metadata item includes the step of storing each component as a BMCP (base model component) box as an extension of the ISO BMFF FullBox.
[0125] Clause 11: In Clause 9 or Clause 10, the step of storing each component individually as a separate metadata item in an ISO BMFF file comprises, for each component, storing a type value representing the type of the component and an encoding value representing how the component is encoded.
[0126] Clause 12: A method comprising, in any one of Clauses 4 through 11, further a step of encrypting data for one or more of the components.
[0127] Clause 13: A method according to any one of Clauses 4 through 12, wherein the step of storing data representing components comprises the step of hierarchically storing data representing components such that each component having a parent component includes data identifying the parent component.
[0128] Clause 14: A method comprising, in any one of Clauses 1 through 13, further the step of disclosing an identifier for an ISO BMFF file to an application server (AS) of a computer network.
[0129] Clause 15: A method comprising, in any one of Clauses 1 through 14, further receiving a request from a receiving device for access to a 3D object model in one of the LODs.
[0130] Clause 16: The method of Clause 15, further comprising the step of providing to a receiving device, in response to a request, data representing the size, complexity, and components of a 3D object model for one requested LOD among the LODs.
[0131] Clause 17: A method for retrieving media data comprising the steps of: retrieving data representing a three-dimensional object model and one or more levels of detail (LODs) for the three-dimensional object model; transmitting a request to access data for the three-dimensional object model in one of the LODs, wherein the data for the three-dimensional object model in one of the LODs includes the size, complexity, and components of the three-dimensional object model for the LOD; and receiving, in response to the request, data representing the size, complexity, and components of the three-dimensional object model for the LOD.
[0132] Clause 18: In Clause 17, the data representing the size, complexity, and components of a 3D object model for the LOD comprises a basic model mapping structure.
[0133] Clause 19: In Clause 17 or Clause 18, the data representing the size, complexity, and components of a 3D object model for an LOD comprises a size value, the number of mesh faces, the number of components, and, for each of the number of components, data representing a corresponding component.
[0134] Clause 20: In Clause 19, the data representing the size, complexity, and components of a 3D object model for an LOD comprises a BMMA (base model mapping) box that includes an extension of the ISO BMFF FullBox.
[0135] Clause 21: A method according to Clause 19 or Clause 20, wherein the data representing the corresponding component comprises one or more images of the component or one or more data representing the geometric structure of the component.
[0136] Clause 22: A method comprising, for each of Clauses 19 through 21, a step of retrieving data describing the role of the component with respect to the 3D object model.
[0137] Clause 23: In Clause 22, the role is a method comprising one or more of a joint, blend shape, mesh, map, or pose transformation.
[0138] Clause 24: A method in which, in any one of Clauses 19 through 23, the data representing the components of a 3D object model for an LOD comprises separate individual metadata items for each of the components.
[0139] Clause 25: In Clause 24, a method wherein each of the distinct individual metadata items comprises an individual BMCP (base model component) box containing an extension of the ISO BMFF FullBox.
[0140] Clause 26: A method in which, in Clause 24 or Clause 25, each of the components comprises a type value representing a type for the component and an encoding value representing how the component is encoded.
[0141] Clause 27: A method comprising, in any one of Clauses 19 through 26, further a step of decoding data for one or more of the components.
[0142] Clause 28: A method according to Clause 27, further comprising the step of retrieving decryption keys for each of the encrypted components.
[0143] Clause 29: A method comprising, in any one of Clauses 19 through 28, a hierarchical representation of components such that the data representing the components includes data in which each component having a parent component identifies the parent component.
[0144] Clause 30: A device for transmitting media data, comprising one or more means for performing the method of any one of Clauses 1 through 29.
[0145] Clause 31: In Clause 30, one or more means comprises a processing system comprising one or more processors implemented as circuits, a device.
[0146] Clause 32: In Clause 30, the device comprises at least one of an integrated circuit; a microprocessor; or a wireless communication device.
[0147] Clause 33: A device for storing media data, comprising means for storing a three-dimensional object model in an ISO base media file format (ISO BMFF) file, wherein the means for storing comprises: means for storing metadata for the three-dimensional object model in an ISO BMFF file; means for storing a number of levels of detail (LODs) for the three-dimensional object model in an ISO BMFF file; and for each number of LODs, means for storing data in the ISO BMFF file that associates the LOD with data representing the size, complexity, and components of the three-dimensional object model for the LOD.
[0148] Clause 34: A device for retrieving media data, comprising: means for retrieving data representing a three-dimensional object model and one or more levels of detail (LODs) for the three-dimensional object model; means for transmitting a request for accessing data for the three-dimensional object model in one of the LODs, wherein the data for the three-dimensional object model in one of the LODs comprises the size, complexity, and components of the three-dimensional object model for the LOD; and means for receiving, in response to the request, data representing the size, complexity, and components of the three-dimensional object model for the LOD.
[0149] Clause 35: A computer-readable storage medium in which instructions are stored, wherein the instructions, when executed, cause a processor to perform the method of any one of Clauses 1 through 29.
[0150] Clause 36: A method for storing media data comprising the step of storing a three-dimensional object model in an ISO base media file format (ISO BMFF) file, wherein the storing step comprises: storing metadata for the three-dimensional object model in the ISO BMFF file; storing a number of levels of detail (LODs) for the three-dimensional object model in the ISO BMFF file; and for each number of LODs, storing data in the ISO BMFF file that associates the LOD with data representing the size, complexity, and components of the three-dimensional object model for the LOD.
[0151] Clause 37: The method of Clause 36, wherein the step of saving a 3D object model includes the step of saving the 3D object model as a BAVM (base avatar model) box as an extension of the ISO BMFF FullBox.
[0152] Clause 38: In Clause 36, the step of storing data that associates the LOD with data representing the size, complexity, and components of the 3D object model for the LOD comprises the step of storing a base model mapping structure in an ISO BMFF file.
[0153] Clause 39: In Clause 36, the data representing the size, complexity, and components of a 3D object model for an LOD comprises a size value, the number of mesh faces, the number of components, and, for each of the number of components, data representing a corresponding component.
[0154] Clause 40: The method of Clause 39, wherein the step of storing data representing the size, complexity, and components of a 3D object model for an LOD includes the step of storing a BMMA (base model mapping) box as an extension of the ISO BMFF FullBox.
[0155] Clause 41: The method of Clause 39, wherein the data representing the corresponding component comprises one or more images of the component or one or more data representing the geometric structure of the component.
[0156] Clause 42: The method of Clause 39, further comprising the step of storing data describing the role of the component to the 3D object model for each of the components.
[0157] Clause 43: In Clause 42, the role is a method comprising one or more of a joint, blend shape, mesh, map, or pose transformation.
[0158] Clause 44: The method of Clause 39, wherein the step of storing data representing the components of a 3D object model for an LOD includes the step of storing each component individually as a separate metadata item in an ISO BMFF file.
[0159] Clause 45: The method of Clause 44, wherein the step of storing each component individually as an individual metadata item includes the step of storing each component as a BMCP (base model component) box as an extension of the ISO BMFF FullBox.
[0160] Clause 46: The method of Clause 44, wherein the step of storing each component individually as a separate metadata item in an ISO BMFF file comprises, for each component, storing a type value representing the type of the component and an encoding value representing how the component is encoded.
[0161] Clause 47: A method according to Clause 39, further comprising the step of encrypting data for one or more of the components.
[0162] Clause 48: The method of Clause 39, wherein the step of storing data representing components comprises the step of hierarchically storing data representing components such that each component having a parent component includes data identifying the parent component.
[0163] Clause 49: A method according to Clause 36, further comprising the step of disclosing an identifier for an ISO BMFF file to an application server (AS) of a computer network.
[0164] Clause 50: The method of Clause 36, further comprising the step of receiving a request from a receiving device for access to a 3D object model in one of the LODs.
[0165] Clause 51: The method of Clause 50, further comprising the step of providing to a receiving device, in response to a request, data representing the size, complexity, and components of a 3D object model for one requested LOD among the LODs.
[0166] Clause 52: A method for retrieving media data comprising the steps of: retrieving data representing a three-dimensional object model and one or more levels of detail (LODs) for the three-dimensional object model; transmitting a request to access data for the three-dimensional object model in one of the LODs, wherein the data for the three-dimensional object model in one of the LODs includes the size, complexity, and components of the three-dimensional object model for the LOD; and receiving, in response to the request, data representing the size, complexity, and components of the three-dimensional object model for the LOD.
[0167] Clause 53: In Clause 52, the data representing the size, complexity, and components of a 3D object model for the LOD comprises a basic model mapping structure.
[0168] Clause 54: In Clause 52, the data representing the size, complexity, and components of a 3D object model for an LOD comprises a size value, the number of mesh faces, the number of components, and, for each of the number of components, data representing a corresponding component.
[0169] Clause 55: In Clause 54, the data representing the size, complexity, and components of a 3D object model for an LOD comprises a BMMA (base model mapping) box that includes an extension of the ISO BMFF FullBox.
[0170] Clause 56: The method of Clause 54, wherein the data representing the corresponding component comprises one or more images of the component or one or more data representing the geometric structure of the component.
[0171] Clause 57: A method according to Clause 54, further comprising, for each of the components, a step of retrieving data describing the role of the component in the 3D object model.
[0172] Clause 58: In Clause 57, the role is a method comprising one or more of a joint, blend shape, mesh, map, or pose transformation.
[0173] Clause 59: In Clause 54, the data representing the components of a 3D object model for an LOD comprises separate individual metadata items for each of the components.
[0174] Clause 60: In Clause 59, a method wherein each of the distinct individual metadata items comprises an individual BMCP (base model component) box containing an extension of the ISO BMFF FullBox.
[0175] Clause 61: In Clause 59, a method wherein each of the components includes a type value representing the type of the component and an encoding value representing how the component is encoded.
[0176] Clause 62: A method according to Clause 54, further comprising the step of decoding data for one or more of the components.
[0177] Clause 63: A method according to Clause 62, further comprising the step of retrieving decryption keys for each of the encrypted components.
[0178] Clause 64: The method of Clause 54, wherein the data representing the components includes a hierarchical representation of the components such that each component having a parent component includes data identifying the parent component.
[0179] Clause 65: A method for storing media data comprising the step of storing a three-dimensional object model in an ISO base media file format (ISO BMFF) file, wherein the storing step comprises: storing metadata for the three-dimensional object model in the ISO BMFF file; storing a number of levels of detail (LODs) for the three-dimensional object model in the ISO BMFF file; and for each number of LODs, storing data in the ISO BMFF file that associates the LOD with data representing the size, complexity, and components of the three-dimensional object model for the LOD.
[0180] Clause 66: The method of Clause 65, wherein the step of saving a 3D object model includes the step of saving the 3D object model as a BAVM (base avatar model) box as an extension of the ISO BMFF FullBox.
[0181] Clause 67: In Clause 65, the step of storing data that associates the LOD with data representing the size, complexity, and components of the 3D object model for the LOD comprises the step of storing a base model mapping structure in an ISO BMFF file.
[0182] Clause 68: In Clause 65, the data representing the size, complexity, and components of a 3D object model for an LOD comprises a size value, the number of mesh faces, the number of components, and, for each of the number of components, data representing a corresponding component.
[0183] Clause 69: In Clause 68, the step of storing data representing the size, complexity, and components of a 3D object model for an LOD comprises the step of storing a BMMA (base model mapping) box as an extension of the ISO BMFF FullBox.
[0184] Clause 70: The method of Clause 68, wherein the data representing the corresponding component comprises one or more images of the component or one or more data representing the geometric structure of the component.
[0185] Clause 71: A method according to Clause 68, further comprising, for each component, a step of storing data describing the role of the component to the 3D object model.
[0186] Clause 72: In Clause 71, the role is a method comprising one or more of a joint, blend shape, mesh, map, or pose transformation.
[0187] Clause 73: In Clause 68, the step of storing data representing components of a 3D object model for an LOD comprises the step of storing each component individually as a separate metadata item in an ISO BMFF file.
[0188] Clause 74: The method of Clause 73, wherein the step of storing each component individually as an individual metadata item includes the step of storing each component as a BMCP (base model component) box as an extension of the ISO BMFF FullBox.
[0189] Clause 75: The method of Clause 73, wherein the step of storing each component individually as a separate metadata item in an ISO BMFF file comprises, for each component, storing a type value representing the type of the component and an encoding value representing how the component is encoded.
[0190] Clause 76: A method of Clause 68, further comprising the step of encrypting data for one or more of the components.
[0191] Clause 77: The method of Clause 68, wherein the step of storing data representing components comprises the step of hierarchically storing data representing components such that each component having a parent component includes data identifying the parent component.
[0192] Clause 78: A method according to Clause 65, further comprising the step of disclosing an identifier for an ISO BMFF file to an application server (AS) of a computer network.
[0193] Clause 79: The method of Clause 65, further comprising the step of receiving a request from a receiving device for access to a 3D object model in one of the LODs.
[0194] Clause 80: The method of Clause 79, further comprising the step of providing to a receiving device, in response to a request, data representing the size, complexity, and components of a 3D object model for one requested LOD among the LODs.
[0195] Clause 81: A device for storing media data comprising: a memory for storing media data; and a processing system implemented in a circuit, wherein the processing system is configured to store a three-dimensional object model in an ISO BMFF (ISO base media file format) file of media data in memory, and to store the three-dimensional object model in the ISO BMFF file, the processing system comprises: storing metadata for the three-dimensional object model in the ISO BMFF file; storing a number of LODs (levels of detail) for the three-dimensional object model in the ISO BMFF file; and
[0196] A device configured to store, for each LOD, data associating the LOD with data representing the size, complexity, and components of the 3D object model for the LOD in an ISO BMFF file.
[0197] Clause 82: A method for retrieving media data comprising the steps of: retrieving, by a client device, data representing a three-dimensional object model stored in an ISO base media file format (ISO BMFF) file and one or more levels of detail (LODs) for the three-dimensional object model stored in the ISO BMFF file, wherein the ISO BMFF file is stored on a server device; transmitting to the server device a request to access data for the three-dimensional object model in one of the LODs, wherein the data for the three-dimensional object model in one of the LODs includes the size, complexity, and components of the three-dimensional object model in one of the LODs; and receiving, by the client device in response to the request, data for the three-dimensional object model in one of the LODs, wherein the data for the three-dimensional object model in one of the LODs includes the size, complexity, and components of the three-dimensional object model in one of the LODs.
[0198] Clause 83: In Clause 82, the data representing the size, complexity, and components of a 3D object model for the LOD comprises a basic model mapping structure.
[0199] Clause 84: In Clause 82, the data representing the size, complexity, and components of a 3D object model for an LOD comprises a size value, the number of mesh faces, the number of components, and, for each of the number of components, data representing a corresponding component.
[0200] Clause 85: In Clause 84, the data representing the size, complexity, and components of a 3D object model for an LOD comprises a BMMA (base model mapping) box that includes an extension of the ISO BMFF FullBox.
[0201] Clause 86: The method of Clause 84, wherein the data representing the corresponding component comprises one or more images of the component or one or more data representing the geometric structure of the component.
[0202] Clause 87: The method of Clause 84, further comprising, for each of the components, a step of retrieving data describing the role of the component in the 3D object model.
[0203] Clause 88: In Clause 87, the role is a method comprising one or more of a joint, blend shape, mesh, map, or pose transformation.
[0204] Clause 89: In Clause 84, the data representing the components of a 3D object model for an LOD comprises separate individual metadata items for each of the components.
[0205] Clause 90: In Clause 89, a method wherein each of the distinct individual metadata items comprises an individual BMCP (base model component) box containing an extension of the ISO BMFF FullBox.
[0206] Clause 91: In Clause 89, a method wherein each of the components includes a type value representing the type of the component and an encoding value representing how the component is encoded.
[0207] Clause 92: The method of Clause 84, further comprising the step of retrieving decryption keys for each of the encrypted components; and the step of decrypting data for one or more of the components using the decryption keys.
[0208] Clause 93: A method in Clause 84 wherein the data representing the components comprises a hierarchical representation of the components such that each component having a parent component includes data identifying the parent component.
[0209] Clause 94: A client device for retrieving media data comprising: memory for storing media data; and a processing system implemented in circuitry, wherein the processing system retrieves data representing a three-dimensional object model stored in an ISO base media file format (ISO BMFF) file and one or more levels of detail (LODs) for the three-dimensional object model stored in the ISO BMFF file—the ISO BMFF file is stored on a server device—; transmits a request to the server device for accessing data for the three-dimensional object model in one of the LODs—the data for the three-dimensional object model in one of the LODs includes the size, complexity, and components of the three-dimensional object model in one of the LODs—; and configured to receive, in response to the request, data for the three-dimensional object model in one of the LODs, wherein the data for the three-dimensional object model in one of the LODs includes the size, complexity, and components of the three-dimensional object model in one of the LODs.
[0210] In one or more examples, the described functions may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functions may be stored as one or more instructions or code on a computer-readable medium or transmitted therethrough and may be executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media corresponding to tangible media such as data storage media, or communication media including any medium that facilitates the transfer of a computer program from one place to another according to a communication protocol, for example. In this way, computer-readable media may generally correspond to (1) non-transitory tangible computer-readable storage media or (2) communication media such as signals or carrier waves. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code and / or data structures for the implementation of the techniques described in this disclosure. Computer program products may include computer-readable media.
[0211] As an example, not a limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be accessed by a computer and used to store desired program code in the form of instructions or data structures. Additionally, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL (digital subscriber line), or wireless technologies (such as infrared, radio, and microwave), then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies (such as infrared, radio, and microwave) are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carriers, signals, or other transient media, but instead refer to non-transient types of storage media. As used herein, disks and discs include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy discs, and Blu-ray discs, wherein disks typically reproduce data magnetically, while discs reproduce data optically using a laser. Combinations of the above should also be included within the scope of computer-readable media.
[0212] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Accordingly, as used herein, the term "processor" may refer to any of the aforementioned structures or any other structures suitable for implementing the techniques described herein. Additionally, in some embodiments, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or integrated into a combined codec. Furthermore, the techniques may be fully implemented by one or more circuits or logic elements.
[0213] The techniques of the present disclosure may be implemented in a wide variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chipsets). Various components, modules, or units are described in the present disclosure to highlight functional modes of devices configured to perform the disclosed techniques, but are not necessarily required to be realized by different hardware units. Rather, as described above, various units may be combined into a codec hardware unit or provided by a set of interacting hardware units including one or more processors as described above, together with suitable software and / or firmware.
[0214] Various examples have been described. These and other examples exist within the scope of the following claims.
Claims
Claim 1 A method for storing media data includes the step of storing a 3D (three-dimensional) object model in an ISO BMFF (ISO base media file format) file, wherein the storing step is: A step of storing metadata for the 3D object model in the ISO BMFF file; A step of storing the number of LODs (levels of detail) for the 3D object model in the ISO BMFF file; and A method comprising the step of, for each of the above LODs, storing data in the ISO BMFF file that associates the LOD with data representing the size, complexity, and components of the 3D object model for the LOD. Claim 2 A method according to claim 1, wherein the step of saving the 3D object model includes the step of saving the 3D object model as a BAVM (base avatar model) box as an extension of the ISO BMFF FullBox. Claim 3 A method according to claim 1, wherein the step of storing the data that associates the LOD with the data representing the size, complexity, and components of the 3D object model for the LOD comprises the step of storing a basic model mapping structure in the ISO BMFF file. Claim 4 A method according to claim 1, wherein the data representing the size, complexity, and components of the 3D object model for the LOD includes a size value, the number of mesh faces, the number of components, and, for each of the number of components, data representing a corresponding component. Claim 5 In claim 4, the step of storing the data representing the size, complexity, and components of the 3D object model for the LOD comprises the step of storing a BMMA (base model mapping) box as an extension of the ISO BMFF FullBox. Claim 6 A method according to claim 4, wherein the data representing the corresponding component comprises one or more images of the component or one or more data representing the geometric structure of the component. Claim 7 A method according to claim 4, further comprising the step of storing data describing the role of the component with respect to the 3D object model for each of the above components. Claim 8 In claim 7, the above role comprises one or more of a joint, blend shape, mesh, map, or pose transformation. Claim 9 In claim 4, the step of storing the data representing the components of the 3D object model for the LOD comprises the step of individually storing each component as an individual metadata item in the ISO BMFF file. Claim 10 In claim 9, the step of individually storing each component as the individual metadata item includes the step of storing each component as a BMCP (base model component) box as an extension of the ISO BMFF FullBox. Claim 11 In claim 9, the step of individually storing each component as an individual metadata item in the ISO BMFF file comprises, for each component, storing a type value representing the type of said component and an encoding value representing how said component is encoded. Claim 12 A method according to claim 4, further comprising the step of encrypting data for one or more of the above components. Claim 13 In claim 4, the step of storing the data representing the components comprises the step of hierarchically storing the data representing the components such that each component having a parent component includes data identifying the parent component. Claim 14 A method according to claim 1, further comprising the step of disclosing an identifier for the ISO BMFF file to an AS (application server) of a computer network. Claim 15 A method according to claim 1, further comprising the step of receiving a request from a receiving device for access to the 3D object model in one of the LODs. Claim 16 A method according to claim 15, further comprising the step of providing to the receiving device, in response to the above request, the data representing the size, complexity, and components of the 3D object model for one of the requested LODs. Claim 17 A device for storing media data comprises: a memory for storing media data; and a processing system implemented as a circuit part, wherein the processing system is configured to store a 3D (three-dimensional) object model in an ISO BMFF (ISO base media file format) file of the media data within the memory, and to store the 3D object model in the ISO BMFF file, the processing system comprises: Metadata for the 3D object model is stored in the ISO BMFF file; Store the number of LODs (levels of detail) for the 3D object model in the ISO BMFF file; and A device configured to store, for each of the above LODs, data relating the LOD to data representing the size, complexity, and components of the 3D object model for the LOD in the ISO BMFF file. Claim 18 A method for retrieving media data, comprising the steps of: retrieving, by a client device, data representing a three-dimensional object model stored in an ISO BMFF (ISO base media file format) file and one or more Levels of Detail (LODs) for said 3D object model stored in said ISO BMFF file - said ISO BMFF file is stored on a server device -; and by the client device transmitting a request to the server device for accessing data for said 3D object model in one of said LODs - said data for said 3D object model in said LOD among said LODs includes the size, complexity, and components of said 3D object model for said LOD among said LODs -; A method comprising the step of receiving data for a 3D object model in one of the LODs in response to a request by the client device, wherein the data for the 3D object model in one of the LODs has the size, the complexity, and the components of the 3D object model in one of the LODs. Claim 19 In claim 18, the data that associates the LOD with the size, complexity, and components of the 3D object model for the LOD comprises a basic model mapping structure. Claim 20 In claim 18, the data representing the size, complexity, and components of the 3D object model for the LOD comprises a size value, the number of mesh faces, the number of components, and data representing a corresponding component for each of the number of components. Claim 21 In claim 20, the data representing the size, complexity, and components of the 3D object model for the LOD comprises a BMMA (base model mapping) box including an extension of the ISO BMFF FullBox. Claim 22 A method according to claim 20, wherein the data representing the corresponding component comprises one or more images of the component or one or more data representing the geometric structure of the component. Claim 23 A method according to claim 20, further comprising the step of searching for data describing the role of the component with respect to the 3D object model for each of the above components. Claim 24 In paragraph 23, the above role comprises one or more of a joint, blend shape, mesh, map, or pose transformation. Claim 25 In claim 20, the method wherein the data representing the components of the 3D object model for the LOD comprises separate individual metadata items for each of the components. Claim 26 In paragraph 25, a method wherein each of the above separate individual metadata items includes an individual BMCP (base model component) box that includes an extension of the ISO BMFF FullBox. Claim 27 A method according to claim 25, wherein each of the above components includes a type value representing a type for said component and an encoding value representing how said component is encoded. Claim 28 In claim 20, one or more of the above components include encrypted components, and the method further comprises: a step of retrieving decryption keys for each of the above encrypted components; and a step of decrypting data for the above encrypted components using the decryption keys. Claim 29 In claim 20, the method comprises a hierarchical representation of the components such that the data representing the components includes data in which each component having a parent component identifies the parent component. Claim 30 A client device for retrieving media data, comprising: a memory for storing media data; and a processing system implemented in a circuit portion, wherein the processing system is Retrieves data representing a 3D (three-dimensional) object model stored in an ISO BMFF (ISO base media file format) file and one or more LODs (levels of detail) for the 3D object model stored in the ISO BMFF file - the ISO BMFF file is stored on a server device -; Sending a request to the server device to access data for the 3D object model in one of the LODs, wherein the data for the 3D object model in one of the LODs includes the size, complexity, and components of the 3D object model for one of the LODs; and A client device configured to receive the data for the 3D object model in one of the LODs in response to the above request, wherein the data for the 3D object model in one of the LODs has the size, the complexity, and the components of the 3D object model for one of the LODs.