Augmented Reality Call Content Protection

A DRM system encrypts 3D assets in AR calls with session keys, addressing theft and misuse by controlling access, thus safeguarding digital assets and preventing impersonation.

JP2026502480APending Publication Date: 2026-01-23QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025539956
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-09
Filing Date
2024-01-10
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Digital assets in augmented reality (AR) calls are vulnerable to theft and misuse, posing risks of intellectual property infringement and impersonation by malicious users.

Method used

Implementing a digital rights management (DRM) system that encrypts 3D assets using session keys, managed by a DRM server, and controls access through encryption key distribution during the AR call.

Benefits of technology

Protects digital assets from theft and misuse by ensuring only authorized participants can access and use them, thereby safeguarding intellectual property and preventing impersonation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026502480000001_ABST
    Figure 2026502480000001_ABST
Patent Text Reader

Abstract

An exemplary device for participating in an augmented reality (AR) call includes a memory configured to store AR data and a processing system having one or more processors implemented in circuitry, wherein the processing system is configured to receive a scene description for the AR call, the scene description including data representing one or more encrypted digital assets for the AR call, request permission to access the encrypted one or more digital assets for the AR call, receive key data used to decrypt the one or more digital assets in response to requesting permission, decrypt the one or more digital assets using the key data to form decrypted digital assets, and render the decrypted digital assets during the AR call.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to U.S. Patent Application No. 18 / 407,996, filed January 9, 2024, and U.S. Provisional Patent Application No. 63 / 479,520, filed January 11, 2023, the entire contents of each of which are incorporated herein by reference. U.S. Patent Application No. 18 / 407,996, filed January 9, 2024, claims the benefit of U.S. Provisional Patent Application No. 63 / 479,520, filed January 11, 2023.

[0002] FIELD This disclosure relates to the storage and transport of encoded video data. [Background technology]

[0003] Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, digital cameras, digital recording devices, digital media players, video gaming devices, video game consoles, cellular or satellite radiotelephones, video teleconferencing devices, etc. Digital video devices implement video compression techniques, such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263 or ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 (also referred to as High Efficiency Video Coding, HEVC), and extensions to such standards, to more efficiently transmit and receive digital video information.

[0004] Video compression techniques perform spatial and / or temporal prediction to reduce or remove redundancy inherent in video sequences. In block-based video coding, video frames or slices may be partitioned into macroblocks. Each macroblock may be further partitioned. Macroblocks in intra-coded (I) frames or slices are coded using spatial prediction with respect to neighboring macroblocks. Macroblocks in inter-coded (P or B) frames or slices may use spatial prediction with respect to neighboring macroblocks in the same frame or slice, or temporal prediction with respect to other reference frames.

[0005] After the video data is encoded, it can be packetized for transmission or storage and assembled into a video file that conforms to any of a variety of standards, such as the International Organization for Standardization (ISO) Base Media File Format and its extensions, e.g., AVC. Summary of the Invention

[0006] This disclosure generally describes techniques for protecting digital assets exchanged during an augmented reality (AR) call. Participants in an AR call may have digital assets that they want to protect, such as digital avatars, clothing for the digital avatars, and items held by or used as decorations for the digital avatars. Participants may want to present these digital assets within a virtual scene for the AR call, but they may also want to prevent others from stealing the digital assets. Theft of digital assets could infringe intellectual property or could be used by malicious users to impersonate users whose digital assets were stolen. Techniques of this disclosure can be used to protect digital assets used in an AR call from theft.

[0007] In one embodiment, a method for participating in an augmented reality (AR) call includes receiving a scene description for the AR call, the scene description including data representing one or more encrypted digital assets for the AR call; requesting permission to access the encrypted one or more digital assets for the AR call; receiving key data used to decrypt the one or more digital assets in response to requesting permission; decrypting the one or more digital assets using the key data to form decrypted digital assets; and rendering the decrypted digital assets during the AR call.

[0008] In another example, a device for participating in an augmented reality (AR) call includes a memory configured to store AR data and a processing system having one or more processors implemented in circuitry, wherein the processing system is configured to receive a scene description for the AR call, the scene description including data representing one or more encrypted digital assets for the AR call, request permission to access the encrypted one or more digital assets for the AR call, receive key data used to decrypt the one or more digital assets in response to requesting permission, decrypt the one or more digital assets using the key data to form decrypted digital assets, and render the decrypted digital assets during the AR call.

[0009] In another example, a device for participating in an augmented reality (AR) call includes means for receiving a scene description for the AR call, the scene description including data representing one or more encrypted digital assets for the AR call; means for requesting permission to access the encrypted one or more digital assets for the AR call; means for receiving key data used to decrypt the one or more digital assets in response to requesting permission; means for decrypting the one or more digital assets using the key data to form decrypted digital assets; and means for rendering the decrypted digital assets during the AR call.

[0010] In another example, a computer-readable storage medium stores instructions that cause a processor of a device for participating in an augmented reality (AR) call to receive a scene description for the AR call, the scene description including data representing one or more encrypted digital assets for the AR call; request permission to access the encrypted one or more digital assets for the AR call; in response to requesting permission, receive key data used to decrypt the one or more digital assets; decrypt the one or more digital assets using the key data to form decrypted digital assets; and render the decrypted digital assets during the AR call.

[0011] In another example, a method for participating in an augmented reality (AR) call includes receiving a request from a first client device participating in the AR call to access one or more protected digital assets of a second client device participating in the AR call; receiving authorization from the second client device to provide the first client device with access to the one or more protected digital assets; and, in response to the authorization from the second client device, providing a decryption key associated with the one or more protected digital assets to the first client device.

[0012] In another example, a device for participating in an augmented reality (AR) call includes a memory configured to store decryption keys and a processing system having one or more processors implemented in circuitry, wherein the processing system is configured to receive, from a first client device participating in the AR call, a request to access one or more protected digital assets of a second client device participating in the AR call; receive, from the second client device, permission to provide the first client device with access to the one or more protected digital assets; and, in response to the permission from the second client device, provide, to the first client device, one of the decryption keys, the decryption key being associated with the one or more protected digital assets.

[0013] In another example, a device for participating in an augmented reality (AR) call includes means for receiving a request from a first client device participating in the AR call to access one or more protected digital assets of a second client device participating in the AR call; means for receiving permission from the second client device to provide the first client device with access to the one or more protected digital assets; and means for providing a decryption key associated with the one or more protected digital assets to the first client device in response to the permission from the second client device.

[0014] In another example, a computer-readable storage medium stores instructions that cause a processor of a device for participating in an augmented reality (AR) call to receive a request from a first client device participating in the AR call to access one or more protected digital assets of a second client device participating in the AR call, receive permission from the second client device to provide the first client device with access to the one or more protected digital assets, and, in response to the permission from the second client device, provide to the first client device one of the decryption keys, the decryption key being associated with the one or more protected digital assets.

[0015] In another example, a method for retrieving digital assets for an augmented reality (AR) call includes receiving a scene description for the AR call, the scene description including data representing one or more digital assets for the AR call; requesting permission to access the one or more digital assets for the AR call; receiving data for the one or more digital assets in response to requesting permission; and rendering the one or more digital assets during the AR call.

[0016] In another example, a device for retrieving digital assets for an augmented reality (AR) call includes a memory configured to store AR data and a processing system having one or more processors implemented in circuitry, the processing system configured to receive a scene description for the AR call, the scene description including data representing one or more digital assets for the AR call, request permission to access the one or more digital assets for the AR call, receive data for the one or more digital assets in response to requesting permission, and render the one or more digital assets during the AR call.

[0017] In another example, a device for transmitting digital assets for an augmented reality (AR) call includes a memory configured to store one or more assets of AR data from a first device participating in the AR call, and a processing system including one or more processors implemented in circuitry, wherein the processing system is configured to receive the one or more assets of AR data from the first device participating in the AR call, receive a request to provide the one or more assets of AR data to a second device participating in the AR call, and in response to the request, transmit the one or more assets of AR data to the second device participating in the AR call.

[0018] In another example, a device for participating in an augmented reality (AR) call includes a memory configured to store decryption keys and a processing system having one or more processors implemented in circuitry, wherein the processing system is configured to receive, from a first client device participating in the AR call, a request to access one or more protected digital assets of a second client device participating in the AR call; receive, from the second client device, permission to provide the first client device with access to the one or more protected digital assets; and, in response to the permission from the second client device, provide, to the first client device, one of the decryption keys, the decryption key being associated with the one or more protected digital assets.

[0019] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will become apparent from the description, drawings, and claims. [Brief explanation of the drawings]

[0020] [Figure 1] 1 is a block diagram illustrating an example system implementing techniques for streaming media data over a network. [Figure 2]FIG. 2 is a block diagram illustrating an example video. [Figure 3] FIG. 1 is a conceptual diagram illustrating an exemplary extension of a glTF scene description to primitive elements. [Figure 4] FIG. 1 is a conceptual diagram illustrating an exemplary extension of a glTF scene description to buffer elements. [Figure 5] 1 is a flow diagram illustrating an example method for encrypting and decrypting 3D assets for an augmented reality (AR) call, according to techniques of this disclosure. [Figure 6] 1 is a flowchart illustrating an example method for exchanging protected digital assets for an augmented reality (AR) call, consistent with techniques of this disclosure. [Figure 7] FIG. 1 is a conceptual diagram illustrating an example method for exchanging protected digital assets for an augmented reality (AR) call in accordance with techniques of this disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0021] Generally, this disclosure describes techniques for protecting content (e.g., images, virtual object data, audio data, or other content) exchanged during an augmented reality (AR) or other extended reality (XR) call, such as a mixed reality (MR) or virtual reality (VR) call.

[0022] GL Transmission Format 2.0 (glTF2) can be used as a scene description format to address the needs of MPEG-I (Moving Pictures Experts Group - Immersive) and 6DoF (Six Degrees of Freedom) applications. Specifying extensions to glTF2 is described, for example, in Khronos Group, The GL Transmission Format (glTF), version 2.0, github.com / KhronosGroup / glTF / tree / master / specification / 2.0#specifying-extensions.

[0023] In general, glTF2 may include data describing static or dynamic scenes. In the context of the techniques of this disclosure, glTF2 can be used to describe scenes that include audio, video, and dynamic media data, such as XR / AR / MR / VR data. For example, a 3D rendered scene may include objects such as a display screen presenting video data or other objects. Similarly, a 3D rendered scene may include audio objects positioned on speakers within the 3D rendered scene.

[0024] In an XR call, a user may present themselves to others in the call using their own 3D assets, such as three-dimensional (3D) avatars, clothing, and AR effects. During an AR call or AR experience in a shared space, a user may need to share their assets with other participants on the call. An AR call / experience may be described by a 3D scene that includes all participants. A gLTF 2.0 scene or scene update may represent the 3D scene. Assets are represented as 3D objects, such as meshes and / or point clouds.

[0025] Participants in an AR call / experience may receive 3D representations of other call member assets. If not protected, other users may make copies of these assets and use them for other purposes after the AR call / experience. In some cases, malicious users may even misuse the 3D assets to impersonate participants in future AR calls / experiences or otherwise misappropriate user-created digital assets that may be protected as intellectual property, for example, under copyright.

[0026] This disclosure describes techniques that can be used to encrypt 3D assets using a digital rights management (DRM) system, which allows for retrieval of encryption keys during an AR call / experience. The use of DRM protection can be signaled in the scene description document through a glTF 2.0 extension.

[0027] In some examples, the DRM license includes a session key that is used by all participants to encrypt the cryptographic keys that they use for their assets. A signaling server, such as a WebRTC signaling server or an IP Multimedia Core Network Subsystem (IMS) proxy, interrogator, or Serving Call Session Control Function (P / I / S-CSCF), may perform tasks attributed to a DRM server.

[0028] In some examples, a user (or user client software) may send data to a DRM server at the beginning of an AR call / experience that authorizes participants to receive their 3D assets. At the end of the call / experience, the user (or user client software) may notify the DRM server that the user's license to access the 3D assets should be revoked.

[0029] In this manner, the techniques of this disclosure may be used to protect user digital assets exchanged during an AR (or other XR, e.g., MR or VR) call with one or more other users. Without such protection, these digital assets may be vulnerable to misuse by malicious users participating in such calls. By implementing these techniques, the problem of digital asset theft or other misuse that arises particularly in the realm of computer-based technologies such as XR / AR / VR / MR calls may be overcome through a solution rooted in computer-based technologies.

[0030] 1 is a block diagram illustrating an example system 10 that implements techniques for streaming media data over a network. In this example, system 10 includes a content preparation device 20, a server device 60, and a client device 40. Client device 40 and server device 60 are communicatively coupled by a network 74, which may include the Internet. In some examples, content preparation device 20 and server device 60 may also be coupled by network 74 or another network, or may be communicatively coupled directly. In some examples, content preparation device 20 and server device 60 may comprise the same device.

[0031] 1 includes an audio source 22 and a video source 24. Audio source 22 may include, for example, a microphone that generates electrical signals representing captured audio data to be encoded by audio encoder 26. Alternatively, audio source 22 may include a storage medium storing previously recorded audio data, an audio data generator such as a computerized synthesizer, or any other source of audio data. Video source 24 may include a video camera, a storage medium encoded with previously recorded video data, a video data generation unit such as a computer graphics source, or any other source of video data that generates video data to be encoded by video encoder 28. Content preparation device 20 is not necessarily communicatively coupled to server device 60 in all embodiments; multimedia content may also be stored on a separate medium that is read by server device 60.

[0032] The raw audio and video data may include analog or digital data. Analog data may be digitized before being encoded by audio encoder 26 and / or video encoder 28. Audio source 22 may acquire audio data from a speaking participant while the speaking participant is speaking, and video source 24 may simultaneously acquire video data of the speaking participant. In other embodiments, audio source 22 may include a computer-readable storage medium containing stored audio data, and video source 24 may include a computer-readable storage medium containing stored video data. In this manner, the techniques described in this disclosure may be applied to live, streaming, real-time audio and video data, or to archived, pre-recorded audio and video data.

[0033] An audio frame corresponding to a video frame is generally an audio frame that includes audio data captured (or generated) by audio source 22 contemporaneously with the video data captured (or generated) by video source 24 contained within that video frame. For example, a speaking participant typically generates audio data by speaking while audio source 22 captures that audio data, and video source 24 captures video data of the speaking participant contemporaneously, i.e., while audio source 22 is capturing the audio data. Thus, an audio frame may correspond in time to one or more particular video frames. Thus, an audio frame corresponding to a video frame generally corresponds to a situation in which the audio and video data are captured contemporaneously, and the audio and video frames contain those contemporaneously captured audio and video data, respectively.

[0034] In some embodiments, audio encoder 26 may encode, in each encoded audio frame, a timestamp representing the time the audio data for that encoded audio frame was recorded, and similarly, video encoder 28 may encode, in each encoded video frame, a timestamp representing the time the video data for that encoded video frame was recorded. In such embodiments, the correspondence of an audio frame to a video frame may include the audio frame including a timestamp and the video frame including the same timestamp. Content preparation device 20 may include internal clocks that enable audio encoder 26 and / or video encoder 28 to generate timestamps, or audio source 22 and video source 24 may include internal clocks that can be used to associate audio data and video data, respectively, with timestamps.

[0035] In some embodiments, audio source 22 may send data to audio encoder 26 corresponding to the time the audio data was recorded, and video source 24 may send data to video encoder 28 corresponding to the time the video data was recorded. In some embodiments, audio encoder 26 may encode a sequence identifier in the encoded audio data that indicates the relative temporal order of the encoded audio data, but not necessarily the absolute time the audio data was recorded; similarly, video encoder 28 may also use a sequence identifier to indicate the relative temporal order of the encoded video data. Similarly, in some embodiments, the sequence identifier may be mapped to or otherwise correlated to a timestamp.

[0036] The audio encoder 26 typically generates a stream of coded audio data, while the video encoder 28 generates a stream of coded video data. Each individual stream of data (whether audio or video) may be referred to as an elementary stream. An elementary stream is a single, digitally coded (and possibly compressed) component of a media presentation. For example, a coded video or audio portion of a media presentation may be an elementary stream. An elementary stream may be converted into a packetized elementary stream (PES) before being encapsulated within a video file. Within the same media presentation, a stream ID may be used to distinguish PES packets belonging to one elementary stream from others. The basic unit of data for an elementary stream is the packetized elementary stream (PES) packet. Therefore, coded video data typically corresponds to an elementary video stream. Similarly, audio data corresponds to one or more respective elementary streams.

[0037] 1, encapsulation unit 30 of content preparation device 20 receives an elementary stream including coded video data from video encoder 28 and an elementary stream including coded audio data from audio encoder 26. In some embodiments, video encoder 28 and audio encoder 26 may each include a packetizer for forming PES packets from the coded data. In other embodiments, video encoder 28 and audio encoder 26 may each interface with a corresponding packetizer for forming PES packets from the coded data. In still other embodiments, encapsulation unit 30 may include packetizers for forming PES packets from the coded audio and video data.

[0038] The video encoder 28 may encode the video data of the multimedia content in various manners to generate different representations of the multimedia content at various bit rates and with different characteristics, such as pixel resolution, frame rate, compliance with various coding standards, compliance with various profiles and / or levels of profiles for various coding standards, representations with one or more views (e.g., for two-dimensional or three-dimensional playback), or other such characteristics. As used in this disclosure, a representation may include one of audio data, video data, text data (e.g., for closed captioning), or other such data. A representation may include an elementary stream, such as an audio elementary stream or a video elementary stream. Each PES packet may include a stream_id, which identifies the elementary stream to which the PES packet belongs. The encapsulation unit 30 is responsible for assembling the elementary streams into streamable media data.

[0039] Encapsulation unit 30 receives PES packets for the elementary streams of the media presentation from audio encoder 26 and video encoder 28 and forms corresponding network abstraction layer (NAL) units from the PES packets. Coded video segments can be organized into NAL units, which provide a "network-friendly" video representation that addresses applications such as video telephony, storage, broadcast, or streaming. NAL units can be classified into Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL units may contain the core compression engine and may contain block-, macroblock-, and / or slice-level data. Other NAL units may be non-VCL NAL units. In some embodiments, a coded picture at a time instance, typically presented as a primary coded picture, can be contained within an access unit, which may contain one or more NAL units.

[0040] Non-VCL NAL units may include, among other things, parameter set NAL units and SEI NAL units. Parameter sets may contain sequence-level header information (in sequence parameter sets (SPS)) and picture-level header information that does not change frequently (in picture parameter sets (PPS)). Parameter sets (e.g., PPS and SPS) allow information that does not change frequently to not need to be repeated for each sequence or picture, thus improving coding efficiency. Furthermore, the use of parameter sets may avoid the need for redundant transmission for error resilience by enabling out-of-band transmission of important header information. In an embodiment of out-of-band transmission, parameter set NAL units may be transmitted on a different channel from other NAL units, such as SEI NAL units.

[0041] Supplemental Enhancement Information (SEI) may contain information that is not necessary for decoding coded picture samples from VCL NAL units, but that can assist processes related to decoding, display, error resilience, and other purposes. SEI messages can be included in non-VCL NAL units. SEI messages are normative parts of some standard specifications and are therefore not necessarily mandatory for the implementation of standard-compliant decoders. SEI messages can be sequence-level SEI messages or picture-level SEI messages. Some sequence-level information can be included in SEI messages, such as the scalability information SEI message in SVC embodiments and the view scalability information SEI message in MVC. These exemplary SEI messages can convey information about, for example, operation point extraction and operation point characteristics.

[0042] Content preparation device 20 may prepare a GL Transmission Format 2.0 (glTF2) bitstream that includes one or more timed media objects, such as audio and video objects. In particular, content preparation device 20 (e.g., its encapsulation unit 30) may prepare a glTF2 scene description for the glTF2 bitstream that indicates the presence of the timed media objects and the location of the timed media objects in a presentation environment (e.g., a three-dimensional space navigable by a user in a virtual reality, video game, or other rendered virtual environment). The timed media objects may also be associated with a presentation time such that client device 40 can present the timed media objects at the timed media objects' current time. For example, content preparation device 20 may capture audio and video data live and stream the live-captured audio and video data to client device 40 in real time. The scene description data may indicate that the timed media objects are stored on server device 60 or another device remote from client device 40, or that the timed media objects are included in the glTF2 bitstream or are otherwise already present on client device 40.

[0043] In accordance with the techniques of this disclosure, client device 40 may also include elements of content preparation device 20 to prepare content to be distributed to other client devices (not shown) via network 74. Server device 60 may include one or more components for a digital rights management (DRM) server, as described in more detail below. In some examples, client device 40 may send content to server device 60 to be distributed to other client devices participating in the AR call (e.g., user movement information, user viewport orientation information, user interaction information (e.g., button presses or interactions with 3D objects in the virtual scene, etc., and audio, video, and / or 3D object content). In other examples, client device 40 may send content directly to other client devices participating in the AR call.

[0044] A client device 40 may receive a glTF scene description extended according to the techniques of this disclosure at a glTF node, mesh, or primitive. The extension may also be associated with a texture, map (e.g., normal map, height map, bump map, etc.), shader, light, or other 3D asset. A glTF extension may indicate that all attribute data of the associated primitive, mesh, or node is encrypted. If added at the node level, all attributes of a mesh primitive may be encrypted. Alternatively, only a subset of the attribute data may be encrypted. The encrypted attributes may then be explicitly signaled. In some examples, the signaling is associated with a buffer element. By being associated with a buffer element, the signaling is generic to all types of media data.

[0045] Client device 40 may retrieve a glTF2 bitstream that includes a scene description that includes data describing the timed media object, such as the location from which the timed media object can be retrieved, the position of the timed media object in the presentation environment, and the presentation time of the timed media object. In this manner, client device 40 may retrieve current timed media data for the timed media object for the current presentation time and present the timed media data at the appropriate location in the presentation environment at the current playback time.

[0046] According to the techniques of this disclosure, client device 40 may receive a glTF scene description to determine one or more 3D assets of a scene that are encrypted or otherwise protected. Client device 40 may further determine how to decrypt the 3D assets, for example, from the glTF scene description. For example, client device 40 may receive an encryption key from server device 60 or another server device according to WebRTC signaling or IMS P / I / S-CSCF. Client device 40 may then use the encryption key to decrypt the 3D assets and present the decrypted 3D assets during a call to a user of client device 40.

[0047] Client device 40 may be configured to instantiate a circular buffer in its memory (not shown in FIG. 1). Audio decoder 46 and video decoder 48 may store frames of audio or video data in the circular buffer, and audio output 42 and video output 44 may retrieve frames from the circular buffer. For example, audio output 42, video output 44, audio decoder 46, and video decoder 48 may maintain read and write pointers in the circular buffer, and audio decoder 46 and video decoder 48 may store decoded frames at the write pointer and then advance the write pointer, while audio output 42 and video output 44 may retrieve decoded frames at the read pointer and then advance the read pointer. Furthermore, client device 40 may prevent the read pointer from exceeding the write pointer and the write pointer from overtaking the read pointer to prevent buffer overflow and buffer underflow.

[0048] Server device 60 includes a Real-time Transport Protocol (RTP) transmission unit 70 and a network interface 72. In some embodiments, server device 60 may include multiple network interfaces. Furthermore, any or all of the features of server device 60 may be implemented on other devices in the content delivery network, such as routers, bridges, proxy devices, switches, or other devices. In some embodiments, intermediate devices in the content delivery network may cache data for multimedia content 64 and may include components that substantially conform to those of server device 60. Generally, network interface 72 is configured to send and receive data over network 74.

[0049] The RTP transmission unit 70 is configured to deliver media data to the client device 40 over the network 74 in accordance with RTP, which is standardized in Request for Comment (RFC) 3550 by the Internet Engineering Task Force (IETF). The RTP transmission unit 70 may also implement protocols related to RTP, such as the RTP Control Protocol (RTCP), the Real-time Streaming Protocol (RTSP), the Session Initiation Protocol (SIP), and / or the Session Description Protocol (SDP). The RTP transmission unit 70 may transmit the media data over a network interface 72, which may implement the Uniform Datagram Protocol (UDP) and / or the Internet Protocol (IP). Thus, in some embodiments, the server device 60 may transmit media data over RTP and RTSP over UDP using the network 74.

[0050] The RTP sending unit 70 may receive an RTSP description request, for example, from the client device 40. The RTSP description request may include data indicating what types of data are supported by the client device 40. The RTP sending unit 70 may respond to the client device 40 with data indicating media streams, such as media content 64, that may be sent to the client device 40 along with corresponding network location identifiers, such as uniform resource locators (URLs) or uniform resource names (URNs).

[0051] The RTP sending unit 70 may then receive an RTSP setup request from the client device 40. The RTSP setup request may generally indicate how the media stream should be transported. The RTSP setup request may include a network location identifier for the requested media data (e.g., media content 64) and a transport specifier, such as a local port for receiving RTP data and control data (e.g., RTCP data) on the client device 40. The RTP sending unit 70 may reply to the RTSP setup request with a confirmation and data indicating the port on the server device 60 to which the RTP data and control data will be sent. The RTP sending unit 70 may then receive an RTSP play request to “play” the media stream, i.e., to send the media stream to the client device 40 over the network 74. The RTP sending unit 70 may also receive an RTSP teardown request to terminate the streaming session, and in response to that request, the RTP sending unit 70 may stop sending media data to the client device 40 for the corresponding session.

[0052] RTP receiving unit 52 may similarly initiate a media stream by first sending an RTSP description request to server device 60. The RTSP description request may indicate the type of data supported by client device 40. RTP receiving unit 52 may then receive a reply from server device 60 specifying available media streams, such as media content 64, that may be sent to client device 40, along with corresponding network location identifiers, such as uniform resource locators (URLs) or uniform resource names (URNs).

[0053] RTP receiving unit 52 may then generate an RTSP setup request and send the RTSP setup request to server device 60. As described above, the RTSP setup request may include a network location identifier for the requested media data (e.g., media content 64) and a transport specifier, such as a local port for receiving RTP data and control data (e.g., RTCP data) on client device 40. In response, RTP receiving unit 52 may receive a confirmation from server device 60 that includes the port of server device 60 that server device 60 will use to transmit the media data and control data.

[0054] After establishing a media streaming session between server device 60 and client device 40, an RTP sending unit 70 of server device 60 may transmit media data (e.g., packets of media data) to client device 40 according to the media streaming session. Server device 60 and client device 40 may exchange control data (e.g., RTCP data), for example, indicating reception statistics by client device 40, thereby enabling server device 60 to perform congestion control or otherwise diagnose and address transmission failures.

[0055] Network interface 54 may receive and provide media of a selected media presentation to RTP receiving unit 52, which in turn may provide the media data to de-encapsulation unit 50. De-encapsulation unit 50 may de-encapsulate elements of a video file into constituent PES streams, de-packetize the PES streams to remove encoded data, and send the encoded data to either audio decoder 46 or video decoder 48, depending, for example, on whether the encoded data is part of an audio stream or a video stream, as indicated by the PES packet headers for that stream. Audio decoder 46 decodes the encoded audio data and sends the decoded audio data to audio output 42, while video decoder 48 decodes the encoded video data and sends the decoded video data, which may include multiple views of a stream, to video output 44.

[0056] Each of the video encoder 28, video decoder 48, audio encoder 26, audio decoder 46, encapsulation unit 30, RTP receiving unit 52, and decapsulation unit 50, as applicable, may be implemented as any of a variety of suitable processing circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuits, software, hardware, firmware, or any combination thereof. Each of the video encoder 28 and video decoder 48 may be included within one or more encoders or decoders, any of which may be integrated as part of a composite video encoder / decoder (CODEC). Similarly, each of the audio encoder 26 and audio decoder 46 may be included within one or more encoders or decoders, any of which may be integrated as part of a composite CODEC. An apparatus including the video encoder 28, the video decoder 48, the audio encoder 26, the audio decoder 46, the encapsulation unit 30, the RTP receiving unit 52, and / or the decapsulation unit 50 may include an integrated circuit, a microprocessor, and / or a wireless communication device such as a cellular telephone.

[0057] Client device 40, server device 60, and / or content preparation device 20 may be configured to operate in accordance with the techniques of this disclosure. For illustrative purposes, this disclosure describes these techniques with respect to client device 40 and server device 60. However, it should be understood that content preparation device 20 may be configured to perform these techniques instead of (or in addition to) server device 60.

[0058] Encapsulation unit 30 can form NAL units, which include a header that identifies the program to which the NAL unit belongs and a payload, e.g., audio data, video data, or data describing the transport stream or program stream to which the NAL unit corresponds. For example, in H.264 / AVC, an NAL unit includes a one-byte header and a variable-sized payload. NAL units that include video data within their payloads can contain video data at various levels of granularity. For example, an NAL unit can include a block of video data, multiple blocks, a slice of video data, or an entire picture of video data. Encapsulation unit 30 can receive encoded video data from video encoder 28 in the form of PES packets of elementary streams. Encapsulation unit 30 can associate each elementary stream with a corresponding program.

[0059] Encapsulation unit 30 can also assemble access units from multiple NAL units. Generally, an access unit may include one or more NAL units for representing a frame of video data, as well as audio data corresponding to that frame, if such audio data is available. An access unit generally includes all NAL units for one output time instance, e.g., all audio and video data for one time instance. For example, if each view has a frame rate of 20 frames per second (fps), each time instance may correspond to a time interval of 0.05 seconds. During this time interval, a particular frame for all views of the same access unit (same time instance) can be rendered simultaneously. In one embodiment, an access unit may include a coded picture at one time instance, which may be presented as a primary coded picture.

[0060] Thus, an access unit may include all audio and video frames of a common time instance, e.g., all views corresponding to time X. This disclosure also refers to the coded pictures of a particular view as a "view component." That is, a view component may include coded pictures (or frames) for a particular view at a particular time. Thus, an access unit may be defined as including all view components of a common time instance. The decoding order of access units is not necessarily the same as the output order or display order.

[0061] After encapsulation unit 30 assembles NAL units and / or access units into a video file based on the received data, encapsulation unit 30 passes the video file to output interface 32 for output. In some embodiments, encapsulation unit 30 may store the video file locally or transmit the video file to a remote server via output interface 32 rather than sending the video file directly to client device 40. Output interface 32 may include, for example, a transmitter, a transceiver, a device for writing data to a computer-readable medium such as an optical drive, a magnetic media drive (e.g., a floppy drive), a universal serial bus (USB) port, a network interface, or other output interface. Output interface 32 outputs the video file to a computer-readable medium such as, for example, a transmission signal, a magnetic medium, an optical medium, a memory, a flash drive, or other computer-readable medium.

[0062] Network interface 54 may receive NAL units or access units via network 74 and provide the NAL units or access units to de-encapsulation unit 50 via RTP receiving unit 52. De-encapsulation unit 50 may de-encapsulate elements of the video file into constituent PES streams, de-packetize the PES streams to extract encoded data, and send the encoded data to either audio decoder 46 or video decoder 48, depending, for example, on whether the encoded data is part of an audio stream or a video stream, as indicated by the stream's PES packet headers. Audio decoder 46 decodes the encoded audio data and sends the decoded audio data to audio output 42, while video decoder 48 decodes the encoded video data and sends the decoded video data, which may include multiple views of a stream, to video output 44.

[0063] FIG. 2 is a block diagram illustrating elements of an exemplary video file 150. As described above, video files according to the ISO Base Media File Format and its extensions store data in a series of objects called "boxes." In the example of FIG. 2, video file 150 includes a file type (FTYP) box 152, a movie (MOOV) box 154, a segment index (sidx) box 162, a movie fragment (MOOF) box 164, and a movie fragment random access (MFRA) box 166. While FIG. 2 represents one example of a video file, it should be understood that other media files may contain other types of media data (e.g., audio data, timed text data, etc.) that are structured similarly to the data in video file 150 according to the ISO Base Media File Format and its extensions.

[0064] The file type (FTYP) box 152 generally represents the file type for the video file 150. The file type box 152 may contain data identifying specifications that describe best use for the video file 150. The file type box 152 may alternatively be placed before the MOOV box 154, the movie fragment box 164, and / or the MFRA box 166.

[0065] 2, MOOV box 154 includes a movie header (MVHD) box 156, a track (TRAK) box 158, and one or more movie extends (MVEX) boxes 160. In general, MVHD box 156 may describe general characteristics of video file 150. For example, MVHD box 156 may include data describing when video file 150 was originally created, data describing when video file 150 was last modified, data describing a timescale for video file 150, data describing a playback duration for video file 150, or other data that generally describes video file 150.

[0066] TRAK box 158 may contain data about a track of video file 150. TRAK box 158 may include a track header (TKHD) box that describes characteristics of the track corresponding to TRAK box 158. In some embodiments, TRAK box 158 may contain coded video pictures, while in other embodiments, the coded video pictures for that track may be contained within a movie fragment 164 that may be referenced by data in TRAK box 158 and / or sidx box 162.

[0067] In some embodiments, video file 150 may include two or more tracks. Thus, MOOV box 154 may include a number of TRAK boxes equal to the number of tracks in video file 150. TRAK box 158 may describe characteristics of the corresponding track in video file 150. For example, TRAK box 158 may describe temporal and / or spatial information about the corresponding track. When encapsulation unit 30 (FIG. 1) includes a parameter set track in a video file, such as video file 150, a TRAK box similar to TRAK box 158 in MOOV box 154 may describe characteristics of the parameter set track. Encapsulation unit 30 may signal, within a TRAK box describing a parameter set track, the presence of a sequence-level SEI message in that parameter set track.

[0068] MVEX box 160 may describe the characteristics of the corresponding movie fragment 164, for example, to signal that video file 150 includes movie fragment 164 in addition to the video data, if any, contained in MOOV box 154. In the context of streaming video data, coded video pictures may be contained in movie fragment 164 rather than in MOOV box 154. Thus, all coded video samples may be contained in movie fragment 164 rather than in MOOV box 154.

[0069] The MOOV box 154 may contain a number of MVEX boxes 160 equal to the number of movie fragments 164 in the video file 150. Each of the MVEX boxes 160 may describe the characteristics of a corresponding one of the movie fragments 164. For example, each MVEX box may contain a movie extends header box (MEHD) box that describes the duration for the corresponding one of the movie fragments 164.

[0070] As described above, encapsulation unit 30 may store a sequence data set within a video sample that does not contain the actual coded video data. A video sample may generally correspond to an access unit, which is a representation of a coded picture at a particular time instance. In the context of AVC, a coded picture includes one or more VCL NAL units that contain information for constructing all pixels of the access unit and other associated non-VCL NAL units, such as SEI messages. Accordingly, encapsulation unit 30 may include a sequence data set, which may include a sequence-level SEI message, within one of the movie fragments 164. Encapsulation unit 30 may further signal the presence of the sequence data set and / or the sequence-level SEI message within one of the movie fragments 164 in one of the MVEX boxes 160 corresponding to that movie fragment 164.

[0071] The SIDX box 162 is an optional element of video file 150; that is, a video file that conforms to the 3GPP® file format, or other such file formats, does not necessarily include a SIDX box 162. According to an embodiment of the 3GPP file format, the SIDX box can be used to identify a sub-segment of a segment (e.g., a segment contained within video file 150). The 3GPP file format defines a sub-segment as "a self-contained set of one or more consecutive Movie Fragment Boxes with corresponding Media Data Boxes, where the Media Data Box(es) containing data referenced by a Movie Fragment Box must follow that Movie Fragment Box and precede the next Movie Fragment Box containing information for the same track." The 3GPP file format also indicates that a SIDX box "contains a sequence of references to subsegments of the (sub)segment documented by that box. The referenced subsegments are contiguous in presentation time. Similarly, the bytes referenced by a Segment Index box are always contiguous within the segment. The referenced size gives a count of the number of bytes in the referenced material."

[0072] SIDX box 162 generally provides information describing one or more subsegments of a segment contained within video file 150. For example, such information may include the playback time at which the subsegment begins and / or ends, a byte offset for the subsegment, whether the subsegment includes (e.g., starts with) a stream access point (SAP), the type for the SAP (e.g., whether the SAP is an instantaneous decoder refresh (IDR) picture, a clean random access (CRA) picture, a broken link access (BLA) picture, etc.), the location of the SAP (in terms of playback time and / or byte offset) within the subsegment, etc.

[0073] A movie fragment 164 may include one or more coded video pictures. In some embodiments, a movie fragment 164 may include one or more groups of pictures (GOPs), each of which may include several coded video pictures, e.g., frames or pictures. Furthermore, as noted above, a movie fragment 164 may include a sequence data set in some embodiments. Each movie fragment 164 may include a movie fragment header box (MFHD, not shown in FIG. 2). The MFHD box may describe characteristics of the corresponding movie fragment, such as a sequence number for that movie fragment. Movie fragments 164 may be included in video file 150 in sequence number order.

[0074] MFRA box 166 can describe random access points within movie fragments 164 of video file 150. This can assist in performing trick modes, such as performing a seek to a particular temporal location (i.e., playback time) within a segment encapsulated by video file 150. MFRA box 166 is generally optional in some embodiments and need not be included within a video file. Similarly, a client device, such as client device 40, does not necessarily need to reference MFRA box 166 to correctly decode and display the video data of video file 150. MFRA box 166 may include a number of track fragment random access (TFRA) boxes (not shown) equal to the number of tracks in video file 150, or, in some embodiments, equal to the number of media tracks (e.g., non-hint tracks) in video file 150.

[0075] In some embodiments, movie fragment 164 may include one or more stream access points (SAPs), such as an IDR picture. Similarly, MFRA box 166 may provide an indication of the location of those SAPs within video file 150. Thus, a temporal sub-sequence of video file 150 may be formed from the SAPs of video file 150. This temporal sub-sequence may also include other pictures, such as P-frames and / or B-frames, that depend on the SAPs. Frames and / or slices of a temporal sub-sequence may be arranged within segments such that frames / slices of the temporal sub-sequence that depend on other frames / slices of the sub-sequence can be properly decoded. For example, in a hierarchical arrangement of data, data used for prediction with respect to other data may also be included within the temporal sub-sequence.

[0076] 3 is a conceptual diagram illustrating an exemplary extension of a glTF scene description to primitive elements. In this example, a node element 180 includes a primitive element 182. A node element 180 may include one or more such primitive elements. Additionally or alternatively, a protected asset may include multiple nodes 180. A primitive element 182 may represent a graphical primitive, such as a triangle or other geometric shape bounded by vertices and edges in a 3D mesh. A primitive element 182 includes attributes 184, such as a position (which describes the 3D position of the corresponding primitive), a normal (which describes the surface normal direction of the corresponding primitive), texture coordinates (which describe the texture of the corresponding primitive), various indices, and a material (e.g., the color and surface shading information of the corresponding primitive).

[0077] The data of Figure 3 may be stored as metadata items, and not necessarily in an ISO Base Media File Format file, such as the file of Figure 2. In some examples, the data of Figure 3 may be stored as metadata in an ISO Base Media File Format file, as metadata in a scene description, as metadata in an independent file, or elsewhere.

[0078] Additionally, in this example, primitive element 182 includes content protection information 186. In accordance with the techniques of this disclosure, content protection information 186 describes the content protection scheme for the corresponding primitive, such as a scheme identifier (schemeId), an address for a DRM server, the network location of a key used to decrypt the protected data, and what data is protected (e.g., what attributes are encrypted). The following pseudocode represents an example set of data that may be used to represent primitive element 182 in accordance with the techniques of this disclosure:

[0079] [Table 1]

[0080] 4 is a conceptual diagram illustrating an example extension of a glTF scene description to a buffer element. In this example, buffer element 190 includes buffer element 192. Buffer element 192 includes a uniform resource identifier (URI), a byteLength value, and a name value, as well as content protection information 194. The following pseudocode represents an example set of data that may be used to represent buffer element 192 in accordance with the techniques of this disclosure.

[0081] [Table 2]

[0082] 5 is a flow diagram illustrating an example method for encrypting and decrypting 3D assets for an AR call in accordance with the techniques of this disclosure. In this example, participants in the method include two user equipment (UE) devices labeled "UE1" and "UE2." Each of the UE devices may include components similar to those of client device 40 and content preparation device 20 of FIG. 1. In the example of FIG. 5, additional participants include a DRM server, a scene manager, and an AR data server. Each of these elements may be included on a different server device, on the same server device, or on any combination of common or separate servers. The servers may be physical servers, virtual servers, or any combination thereof.

[0083] In this example, UE1 and a scene manager may first establish a communication session, and UE1 may offer an encrypted avatar (or other digital asset) to the scene manager (200). The scene manager and UE2 may then establish a communication session (202). The scene manager may then deliver a scene description for an AR call / experience including both UE1 and UE2 to UE1 and UE2 (204). According to the techniques of this disclosure, the scene description may include data representing content protection for the avatar from UE1, for example, extension to primitive elements according to the techniques described with reference to FIG. 3 or extension to buffer elements according to the techniques described with reference to FIG. 4.

[0084] UE2 may then obtain permission from the DRM server to obtain the protected content from UE1 (206). The DRM server may determine that UE1 has authorized UE2 to access the protected content (208). The DRM server may then provide UE2 with a session key for decrypting the encryption key used to encrypt the protected content from UE1 (210). UE2 may then retrieve avatar data and the protected encryption key for decrypting the avatar data from the AR data server (212). UE2 may then decrypt and render UE1's avatar (214).

[0085] UE1 may authorize other users, such as UE2, or other participants in the AR call in various ways. In some examples, UE1 may receive a request from a DRM server and check that UE1 is in an AR session with UE2. In some examples, UE1 may inform the DRM server about a shared session secret / token and identify other authorized participants using information such as IP addresses or session information protocol (SIP) information. In some examples, UE1 may provide its contact list to the DRM server as a list of pre-authorized users.

[0086] The AR data server can authenticate UE2 using data similar to that of the DRM server. That is, the AR data server may authenticate UE2. For example, the AR data server may receive a session key from UE2 and use the session key to authenticate UE2 as a valid participant in the AR call. Additionally or alternatively, the AR data server may receive a shared session token for the AR call, the shared session token being shared by UE1 and UE2 (and any other bona fide participants in the AR call). The AR data server may use the shared session token to authenticate UE2 as a valid participant in the AR call. Additionally or alternatively, the AR data server may receive a participant list for the AR call, the participant list indicating that each participant in the participant list is authorized to access one or more protected digital assets, and when the participant list includes UE2, authenticate UE2 as a valid participant in the AR call.

[0087] In this way, the techniques of this disclosure enable protection of 3D assets during AR calls and shared experiences. A DRM system can be used to ensure that assets are only used for the lifetime of an AR session. Assets can be encrypted only once with a private key. The session key can be used to decrypt the private key, and the secret can be used to decrypt the 3D asset in a trusted DRM environment.

[0088] 6 is a flowchart illustrating an example method for exchanging protected digital assets for an augmented reality (AR) call, such as an augmented reality (XR) call, in accordance with techniques of this disclosure. In particular, FIG. 6 may be executed by a client device (e.g., a user equipment (UE) device) involved in the AR call. The client device may be a device configured to retrieve and use one or more protected (e.g., encrypted) digital assets, such as an avatar of another user participating in the AR call.

[0089] Initially, a client device may establish an XR session with one or more other devices (250). At least one of the other devices may include protected (e.g., encrypted) 3D object model data, such as avatar data. The client device may receive a scene description for the XR session (252). The client device may then determine one or more encrypted digital assets for the XR session from the scene description (254). Accordingly, the client device may request permission to access the encrypted assets (256). For example, the client device may send a request to a digital rights management (DRM) server associated with the XR session. The network address of the DRM server, such as the URL of the DRM server, may be included in the scene description.

[0090] In response to the request, assuming the DRM server authenticates the client device, the client device may receive (258) a decryption key used to decrypt the protected (encrypted) assets. The decryption key may be a session key used to encrypt all protected assets (or keys associated with protected assets) for the XR session. For example, the same key may be used to both encrypt and decrypt protected assets, and the session key may be used to encrypt the key itself. In this way, only devices involved in the XR session may be sent the session key, and as a result, only devices involved in the XR session can decrypt the keys used to decrypt protected assets using the session key.

[0091] The client device may then use the decryption key to decrypt the encrypted asset (260). Finally, the client device may render and display the asset (262).

[0092] Thus, the method of FIG. 6 represents one example of a method that includes receiving a scene description for an AR call, the scene description including data representing one or more encrypted digital assets for the AR call; requesting permission to access the encrypted one or more digital assets for the AR call; receiving key data used to decrypt the one or more digital assets in response to requesting permission; decrypting the one or more digital assets using the key data to form decrypted digital assets; and rendering the decrypted digital assets during the AR call.

[0093] 7 is a conceptual diagram illustrating an example method for exchanging protected digital assets for an augmented reality (AR) call in accordance with the techniques of this disclosure. In particular, the method of FIG. 7 may be performed by a DRM server to authenticate devices participating in an augmented reality (XR) session (e.g., an AR call) and to distribute decryption keys to the authenticated devices for decrypting the protected digital assets of the XR session.

[0094] Initially, the DRM server may receive a request from a second client device (e.g., UE2) to access an asset of a first client device (e.g., UE1) (280). The DRM server may send the request to the first client device / UE1 (282). The DRM server may receive a response from the first client device (284) that includes permission for the second client device / UE2 to access the asset. In some examples, the permission may itself be a decryption key, while in other examples, the DRM server may separately receive the decryption key from the first client device / UE1 (286). The DRM server may then send the decryption key to the second client device (288).

[0095] In this manner, the method of FIG. 7 represents an example of a method that includes receiving a request from a first client device participating in an AR call to access one or more protected digital assets of a second client device participating in the AR call; receiving authorization from the second client device to provide the first client device with access to the one or more protected digital assets; and, in response to the authorization from the second client device, providing a decryption key associated with the one or more protected digital assets to the first client device.

[0096] Various embodiments of the techniques of this disclosure are summarized in the following clauses: Clause 1: A method for participating in an augmented reality (AR) call, the method comprising: receiving a scene description for the AR call, the scene description including data representing one or more encrypted digital assets for the AR call; requesting authorization to access the encrypted one or more digital assets for the AR call; receiving key data used to decrypt the one or more digital assets in response to requesting authorization; decrypting the one or more digital assets using the key data to form decrypted digital assets; and rendering the decrypted digital assets during the AR call.

[0097] Clause 2: The method of clause 1, wherein requesting permission to access one or more digital assets includes sending a request to a Digital Rights Management (DRM) server.

[0098] Clause 3: The method of clause 2, wherein the scene description includes information that associates a DRM server with one or more digital assets.

[0099] Clause 4: The method of any of clauses 2 and 3, wherein the scene description includes a uniform resource indicator (URI) or uniform resource locator (URL) of the DRM server.

[0100] Clause 5: The method of any of clauses 1 to 4, wherein the key data includes a session key, and wherein decrypting one or more digital assets includes decrypting an encrypted version of the encryption key using the session key to form a decrypted encryption key, and decrypting the one or more digital assets using the decrypted encryption key.

[0101] Clause 6: The method of clause 5, further comprising extracting an encrypted version of the encryption key from the scene description.

[0102] Clause 7: The method of any one of clauses 1 to 6, further comprising retrieving one or more digital assets from an AR data server.

[0103] Clause 8: The method of clause 1, wherein requesting permission to access one or more digital assets includes sending a request to a Digital Rights Management (DRM) server.

[0104] Clause 9: The method of clause 8, wherein the scene description includes information that associates a DRM server with one or more digital assets.

[0105] Clause 10: The method of clause 8, wherein the scene description includes a uniform resource indicator (URI) or uniform resource locator (URL) of the DRM server.

[0106] Clause 11: The method of clause 1, wherein the key data includes a session key, and wherein decrypting one or more digital assets includes using the session key to decrypt an encrypted version of the encryption key to form a decrypted encryption key, and using the decrypted encryption key to decrypt the one or more digital assets.

[0107] Clause 12: The method of clause 11, further comprising extracting an encrypted version of the encryption key from the scene description.

[0108] Clause 13: The method of clause 1, further comprising retrieving one or more digital assets from an AR data server.

[0109] Clause 14: A method for participating in an augmented reality (AR) call, the method comprising: encrypting one or more digital assets for the AR call; providing the one or more digital assets to a scene manager of the AR call; and providing data representing one or more participants in the AR call who are authorized to access the one or more digital assets.

[0110] Clause 15: A method including a combination of any of the methods in clauses 1 to 7 and the method in clause 14.

[0111] Clause 16: The method of any of clauses 14 and 15, wherein encrypting one or more digital assets includes encrypting the one or more digital assets using an encryption key.

[0112] Clause 17: The method of clause 16, further comprising receiving a session key from a digital rights management (DRM) server, encrypting the encryption key using the session key to form an encrypted version of the encryption key, and providing the encrypted version of the encryption key to the scene manager.

[0113] Clause 18: A method according to any of clauses 14 to 17, wherein providing data representing one or more participants in the AR call who are authorized to access one or more digital assets comprises receiving a request from a digital rights management (DRM) server identifying the participants in the AR call, and sending data to the DRM server indicating that the participants are authorized.

[0114] Clause 19: A method as described in any of clauses 14 to 17, wherein providing data representing one or more participants in the AR call who are authorized to access one or more digital assets includes sending a shared session token for the AR call to a digital rights management (DRM) server.

[0115] Clause 20: A method as described in any of clauses 14 to 17, wherein providing data representing one or more participants in the AR call who are authorized to access one or more digital assets includes sending a participant list in the AR call to a digital rights management (DRM) server, the participant list indicating that each participant in the participant list is authorized to access one or more digital assets.

[0116] Clause 21: A method as described in any of clauses 14 to 20, wherein providing data representing one or more participants in an AR call who are authorized to access one or more digital assets includes providing identification information of the one or more participants, the identification information including one of an Internet Protocol (IP) address or Session Information Protocol (SIP) information.

[0117] Clause 22: The method of clause 14, wherein encrypting one or more digital assets includes encrypting the one or more digital assets using an encryption key.

[0118] Clause 23. The method of clause 22, further comprising receiving a session key from a digital rights management (DRM) server, encrypting the encryption key using the session key to form an encrypted version of the encryption key, and providing the encrypted version of the encryption key to the scene manager.

[0119] Clause 24: The method of clause 14, wherein providing data representing one or more participants in the AR call who are authorized to access one or more digital assets includes receiving a request from a digital rights management (DRM) server identifying participants in the AR call, and sending data to the DRM server indicating that the participants are authorized.

[0120] Clause 25: The method of clause 14, wherein providing data representing one or more participants in an AR call who are authorized to access one or more digital assets includes sending a shared session token for the AR call to a digital rights management (DRM) server.

[0121] Clause 26: The method of clause 14, wherein providing data representing one or more participants in the AR call who are authorized to access one or more digital assets includes sending a participant list in the AR call to a digital rights management (DRM) server, the participant list indicating that each participant in the participant list is authorized to access one or more digital assets.

[0122] Clause 27: The method described in Clause 14, wherein providing data representing one or more participants in an AR call who are authorized to access one or more digital assets includes providing identification information of the one or more participants, wherein the identification information includes one of an Internet Protocol (IP) address or Session Information Protocol (SIP) information.

[0123] Clause 28: A device for participating in an augmented reality (AR) call, the device comprising one or more means for performing the methods of any of clauses 1 to 27.

[0124] Clause 29: A device according to clause 28, wherein the one or more means comprise one or more processors implemented in the circuitry.

[0125] Clause 30: A device according to any of clauses 28 and 29, wherein the one or more means comprises a memory for storing one or more digital assets.

[0126] Clause 31: A device according to any of clauses 28 to 30, wherein the apparatus comprises at least one of an integrated circuit, a microprocessor, or a wireless communication device.

[0127] Clause 32: A computer-readable storage medium having stored thereon instructions that, when executed, cause a processor to perform any of the methods of clauses 1 to 21.

[0128] Clause 33: A device for participating in an augmented reality (AR) call, the device comprising: means for receiving a scene description for the AR call, the scene description including data representing one or more encrypted digital assets for the AR call; means for requesting permission to access the encrypted one or more digital assets for the AR call; means for receiving key data used to decrypt the one or more digital assets in response to requesting permission; means for decrypting the one or more digital assets using the key data to form decrypted digital assets; and means for rendering the decrypted digital assets during the AR call.

[0129] Clause 34: A device for participating in an augmented reality (AR) call, comprising: means for encrypting one or more digital assets for the AR call; means for providing the one or more digital assets to a scene manager for the AR call; and means for providing data representing one or more participants in the AR call who are authorized to access the one or more digital assets.

[0130] Clause 35: A method for participating in an augmented reality (AR) call, the method comprising: receiving a scene description for the AR call, the scene description including data representing one or more encrypted digital assets for the AR call; requesting permission to access the encrypted one or more digital assets for the AR call; receiving key data used to decrypt the one or more digital assets in response to requesting permission; decrypting the one or more digital assets using the key data to form decrypted digital assets; and rendering the decrypted digital assets during the AR call.

[0131] Clause 36: The method of clause 35, wherein requesting permission to access one or more digital assets includes sending a request to a Digital Rights Management (DRM) server.

[0132] Clause 37: The method of clause 36, wherein the scene description includes information associating a DRM server with one or more digital assets.

[0133] Clause 38: The method of clause 36, wherein the scene description includes a uniform resource indicator (URI) or uniform resource locator (URL) of the DRM server.

[0134] Clause 39: The method of clause 35, wherein the key data includes a session key, and wherein decrypting one or more digital assets includes using the session key to decrypt an encrypted version of the encryption key to form a decrypted encryption key, and using the decrypted encryption key to decrypt the one or more digital assets.

[0135] Clause 40: The method of clause 39, further comprising extracting an encrypted version of the encryption key from the scene description.

[0136] Clause 41: The method of clause 35, further comprising retrieving one or more digital assets from an AR data server.

[0137] Clause 42: The method of clause 35, wherein the one or more digital assets include one or more GL Transmission Format 2.0 (glTF2) nodes, meshes, primitives, textures, normal maps, height maps, bump maps, shaders, or lights.

[0138] Clause 43: A device for participating in an augmented reality (AR) call, the device comprising: a memory configured to store AR data; and a processing system having one or more processors implemented in circuitry, the processing system being configured to receive a scene description for the AR call, the scene description including data representing one or more encrypted digital assets for the AR call; request permission to access the encrypted one or more digital assets for the AR call; in response to requesting permission, receive key data used to decrypt the one or more digital assets; decrypt the one or more digital assets using the key data to form decrypted digital assets; and render the decrypted digital assets during the AR call.

[0139] Clause 44: The device of clause 43, wherein the processing system is configured to send a request to a digital rights management (DRM) server to request permission to access one or more digital assets.

[0140] Clause 45: The device of clause 44, wherein the scene description includes information associating a DRM server with one or more digital assets.

[0141] Clause 46: The device of clause 44, wherein the scene description includes a uniform resource indicator (URI) or uniform resource locator (URL) of the DRM server.

[0142] Clause 47: The device of clause 43, wherein the key data includes a session key, and wherein, to decrypt one or more digital assets, the processing system is configured to use the session key to decrypt an encrypted version of the encryption key to form a decrypted encryption key, and to use the decrypted encryption key to decrypt the one or more digital assets.

[0143] Clause 48: The device of clause 47, wherein the processing system is further configured to extract an encrypted version of the encryption key from the scene description.

[0144] Clause 49: The device of clause 43, wherein the processing system is further configured to retrieve one or more digital assets from the AR data server.

[0145] Clause 50: The device of clause 43, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0146] Clause 51: The device of clause 43, further comprising a display configured to display the rendered digital asset.

[0147] Clause 52: A method for participating in an augmented reality (AR) call, the method comprising: receiving a request from a first client device participating in the AR call to access one or more protected digital assets of a second client device participating in the AR call; receiving authorization from the second client device to provide the first client device with access to the one or more protected digital assets; and, in response to the authorization from the second client device, providing a decryption key associated with the one or more protected digital assets to the first client device.

[0148] Clause 53: The method of clause 52, further comprising receiving a decryption key from the first client device.

[0149] Clause 54: The method of clause 52, wherein receiving authorization includes receiving data from the second client device representing one or more participants in the AR call who are authorized to access one or more protected digital assets.

[0150] Clause 55: The method of clause 52, wherein receiving authorization includes receiving a shared session token for the AR call, the shared session token being shared by the first client device and the second client device.

[0151] Clause 56: The method of clause 52, wherein receiving authorization includes receiving a participant list in an AR call, the participant list indicating that each participant in the participant list is authorized to access one or more protected digital assets, and the participant list includes the first client device.

[0152] Clause 57: The method of clause 52, wherein receiving authorization includes receiving identification information of the first client device, the identification information including one of an Internet Protocol (IP) address or Session Information Protocol (SIP) information.

[0153] Clause 58: A device for participating in an augmented reality (AR) call, comprising: a memory configured to store decryption keys; and a processing system having one or more processors implemented in circuitry, wherein the processing system is configured to receive a request from a first client device participating in the AR call to access one or more protected digital assets of a second client device participating in the AR call; receive authorization from the second client device to provide the first client device with access to the one or more protected digital assets; and in response to the authorization from the second client device, provide the first client device with one of the decryption keys, the decryption key being associated with the one or more protected digital assets.

[0154] Clause 59: The device of clause 58, wherein the processing system is further configured to receive a decryption key from the first client device.

[0155] Clause 60: A device as described in Clause 58, wherein the processing system is configured to receive data from a second client device representing one or more participants in the AR call who are authorized to access one or more protected digital assets.

[0156] Clause 61: The device described in Clause 58, wherein the processing system is configured to receive a shared session token for an AR call, the shared session token being shared by the first client device and the second client device.

[0157] Clause 62: A device as described in Clause 58, wherein the processing system is configured to receive a participant list in an AR call, the participant list indicating that each participant in the participant list is authorized to access one or more protected digital assets, and the participant list includes a first client device.

[0158] Clause 63: The device described in Clause 58, wherein the processing system is configured to receive identification information of the first client device, the identification information including one of an Internet Protocol (IP) address or Session Information Protocol (SIP) information.

[0159] Clause 64: The device of clause 58, wherein the device comprises a digital rights management (DRM) server device.

[0160] Clause 65: A method for participating in an augmented reality (AR) call, the method comprising: receiving a scene description for the AR call, the scene description including data representing one or more encrypted digital assets for the AR call; requesting permission to access the encrypted one or more digital assets for the AR call; receiving key data used to decrypt the one or more digital assets in response to requesting permission; decrypting the one or more digital assets using the key data to form decrypted digital assets; and rendering the decrypted digital assets during the AR call.

[0161] Clause 66: The method of clause 65, wherein requesting permission to access one or more digital assets includes sending a request to a Digital Rights Management (DRM) server.

[0162] Clause 67: The method of clause 66, wherein the scene description includes information associating a DRM server with one or more digital assets.

[0163] Clause 68: The method of any of clauses 66 and 67, wherein the scene description includes a uniform resource indicator (URI) or uniform resource locator (URL) of the DRM server.

[0164] Clause 69: A method as described in any of clauses 66 to 68, wherein the key data includes a session key, and wherein decrypting one or more digital assets includes using the session key to decrypt an encrypted version of the encryption key to form a decrypted encryption key, and using the decrypted encryption key to decrypt the one or more digital assets.

[0165] Clause 70: The method of clause 69, further comprising extracting an encrypted version of the encryption key from the scene description.

[0166] Clause 71: The method of any of clauses 66 to 70, further comprising retrieving one or more digital assets from an AR data server.

[0167] Clause 72: A method according to any of clauses 66 to 71, wherein the one or more digital assets include one or more GL Transmission Format 2.0 (glTF2) nodes, meshes, primitives, textures, normal maps, height maps, bump maps, shaders, or lights.

[0168] Clause 73: A device for participating in an augmented reality (AR) call, the device comprising: a memory configured to store AR data; and a processing system having one or more processors implemented in circuitry, the processing system being configured to receive a scene description for the AR call, the scene description including data representing one or more encrypted digital assets for the AR call; request permission to access the encrypted one or more digital assets for the AR call; in response to requesting permission, receive key data used to decrypt the one or more digital assets; decrypt the one or more digital assets using the key data to form decrypted digital assets; and render the decrypted digital assets during the AR call.

[0169] Clause 74: The device of clause 73, wherein the processing system is configured to send a request to a digital rights management (DRM) server to request permission to access one or more digital assets.

[0170] Clause 75: The device of clause 74, wherein the scene description includes information associating a DRM server with one or more digital assets.

[0171] Clause 76: A device according to any of clauses 74 and 75, wherein the scene description comprises a uniform resource indicator (URI) or a uniform resource locator (URL) of the DRM server.

[0172] Clause 77: A device described in any of Clauses 73 to 76, wherein the key data includes a session key, and wherein, to decrypt one or more digital assets, the processing system is configured to use the session key to decrypt an encrypted version of the encryption key to form a decrypted encryption key, and to use the decrypted encryption key to decrypt the one or more digital assets.

[0173] Clause 78: The device of clause 77, wherein the processing system is further configured to extract an encrypted version of the encryption key from the scene description.

[0174] Clause 79: A device described in any of clauses 73 to 78, wherein the processing system is further configured to retrieve one or more digital assets from the AR data server.

[0175] Clause 80: A device according to any of clauses 73 to 79, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0176] Clause 81: A device described in any of clauses 73 to 80, further comprising a display configured to display the rendered digital asset.

[0177] Clause 82: A method for participating in an augmented reality (AR) call, the method comprising: receiving a request from a first client device participating in the AR call to access one or more protected digital assets of a second client device participating in the AR call; receiving authorization from the second client device to provide the first client device with access to the one or more protected digital assets; and, in response to the authorization from the second client device, providing a decryption key associated with the one or more protected digital assets to the first client device.

[0178] Clause 83: The method of clause 82, further comprising receiving a decryption key from the first client device.

[0179] Clause 84: A method as described in any of clauses 82 and 83, wherein receiving authorization includes receiving data from the second client device representing one or more participants in the AR call who are authorized to access one or more protected digital assets.

[0180] Clause 85: A method according to any of clauses 82 to 84, wherein receiving authorization includes receiving a shared session token for the AR call, the shared session token being shared by the first client device and the second client device.

[0181] Clause 86: A method according to any of clauses 82 to 85, wherein receiving authorization includes receiving a participant list in an AR call, the participant list indicating that each participant in the participant list is authorized to access one or more protected digital assets, and the participant list includes the first client device.

[0182] Clause 87: A method according to any of clauses 82 to 86, wherein receiving authorization includes receiving identification information of the first client device, the identification information including one of an Internet Protocol (IP) address or Session Information Protocol (SIP) information.

[0183] Clause 88: A device for participating in an augmented reality (AR) call, comprising: a memory configured to store decryption keys; and a processing system having one or more processors implemented in circuitry, wherein the processing system is configured to receive, from a first client device participating in the AR call, a request to access one or more protected digital assets of a second client device participating in the AR call; receive, from the second client device, authorization to provide the first client device with access to the one or more protected digital assets; and, in response to the authorization from the second client device, provide, to the first client device, one of the decryption keys, the decryption key being associated with the one or more protected digital assets.

[0184] Clause 89: The device of clause 88, wherein the processing system is further configured to receive a decryption key from the first client device.

[0185] Clause 90: A device described in any of clauses 88 and 89, wherein the processing system is configured to receive data from a second client device representing one or more participants in the AR call who are authorized to access one or more protected digital assets.

[0186] Clause 91: A device described in any of clauses 88 to 90, wherein the processing system is configured to receive a shared session token for an AR call, the shared session token being shared by a first client device and a second client device.

[0187] Clause 92: A device described in any of clauses 88 to 91, wherein the processing system is configured to receive a participant list in an AR call, the participant list indicating that each participant in the participant list is authorized to access one or more protected digital assets, and the participant list includes a first client device.

[0188] Clause 93: A device described in any of clauses 88 to 92, wherein the processing system is configured to receive identification information of a first client device, the identification information including one of an Internet Protocol (IP) address or Session Information Protocol (SIP) information.

[0189] Clause 94: A device according to any one of clauses 88 to 93, wherein the device comprises a Digital Rights Management (DRM) server device.

[0190] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which correspond to tangible media such as data storage media, or communication media, including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communications protocol. As such, computer-readable media may generally correspond to (1) tangible computer-readable storage media that is non-transitory, or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0191] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included within the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead cover non-transitory tangible storage media. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically and discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0192] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the term "processor," as used herein, may refer to any of the above structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided in dedicated hardware and / or software modules configured for encoding and decoding, or may be incorporated into a combined codec. It is also possible for these techniques to be implemented entirely in one or more circuits or logic elements.

[0193] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). Various components, modules, or units have been described in this disclosure to highlight functional aspects of devices configured to perform the disclosed techniques, but they do not necessarily require realization by different hardware units. Rather, as described above, the various units may be combined in a codec hardware unit or may be provided by a collection of interoperable hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware.

[0194] Various examples have been described. These and other examples are within the scope of the following claims.

Claims

1. 1. A method for retrieving digital assets for an augmented reality (AR) call, comprising: receiving a scene description for an AR call, the scene description including data representing one or more digital assets for the AR call; requesting permission to access the one or more digital assets for the AR call; receiving data of the one or more digital assets in response to requesting authorization; Rendering the one or more digital assets during the AR call; and A method comprising:

2. The one or more digital assets include one or more encrypted digital assets, and requesting permission to access the one or more digital assets includes sending a request to a Digital Rights Management (DRM) server, the method comprising, in response to requesting permission: receiving key data used to decrypt the one or more digital assets; decrypting the one or more digital assets using the data of the key to form a decrypted digital asset; The method of claim 1 further comprising:

3. The method of claim 2 , wherein the scene description includes information that associates the DRM server with the one or more digital assets.

4. The method of claim 2 , wherein the scene description includes a uniform resource indicator (URI) or a uniform resource locator (URL) of the DRM server.

5. the data of the key includes a session key and decrypts the one or more digital assets; decrypting an encrypted version of the encryption key using the session key to form a decrypted encryption key; decrypting the one or more digital assets using the decrypted encryption key; Including, The method of claim 2.

6. The method of claim 5 , further comprising extracting the encrypted version of the encryption key from the scene description.

7. The method of claim 1 , further comprising retrieving the one or more digital assets from an AR data server.

8. The method of claim 1 , wherein the one or more digital assets include one or more GL Transmission Format 2.0 (glTF2) nodes, meshes, primitives, textures, normal maps, height maps, bump maps, shaders, or lights.

9. 1. A device for retrieving digital assets for an augmented reality (AR) call, comprising: a memory configured to store AR data; a processing system including one or more processors implemented in circuitry, said processing system comprising: receiving a scene description for an AR call, the scene description including data representing one or more digital assets for the AR call; requesting permission to access the one or more digital assets for the AR call; receiving data for the one or more digital assets in response to requesting authorization; configured to render the one or more digital assets during the AR call. device.

10. the one or more digital assets include one or more encrypted digital assets, and the processing system is configured to send a request to a Digital Rights Management (DRM) server to request permission to access the one or more digital assets, and in response to the processing system requesting the permission, receiving key data used to decrypt the one or more digital assets; and further configured to decrypt the one or more digital assets using the data of the key to form a decrypted digital asset. The device of claim 9.

11. The device of claim 10 , wherein the scene description includes information that associates the DRM server with the one or more digital assets.

12. The device of claim 10 , wherein the scene description includes a uniform resource indicator (URI) or uniform resource locator (URL) of the DRM server.

13. the data of the keys includes a session key, and to decrypt the one or more digital assets, the processing system: decrypting an encrypted version of the encryption key using the session key to form a decrypted encryption key; configured to decrypt the one or more digital assets using the decrypted encryption key; The device of claim 10.

14. The device of claim 13 , wherein the processing system is further configured to extract the encrypted version of the encryption key from the scene description.

15. The device of claim 9 , wherein the processing system is further configured to retrieve the one or more digital assets from an AR data server.

16. The device of claim 9 , wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

17. The device of claim 9 , further comprising a display configured to display the rendered digital asset.

18. 1. A device for transmitting digital assets for an augmented reality (AR) call, comprising: a memory configured to store one or more assets of AR data from a first device participating in the AR call; a processing system including one or more processors implemented in circuitry, said processing system comprising: receiving the one or more assets of the AR data from the first device participating in the AR call; receiving a request to provide the one or more assets of the AR data to a second device participating in the AR call; configured to, in response to the request, transmit the one or more assets of the AR data to the second device participating in the AR call. device.

19. 20. The device of claim 18, wherein the processing system is configured to transmit the one or more assets of the AR data to the second device in response to authenticating the second device participating in the AR call.

20. the processing system comprising: receiving a session key from the second device; configured to use the session key to authenticate the second device participating in the AR call.

20. The device of claim 19.

21. the processing system comprising: receiving a shared session token for the AR call, the shared session token being shared by the first device and the second device; configured to use the shared session token to authenticate the second device participating in the AR call.

20. The device of claim 19.

22. the processing system comprising: receiving a list of participants in the AR call, the list indicating that each of the participants in the list is authorized to access the one or more digital assets; configured to authenticate the second device participating in the AR call when the participant list includes the second device.

20. The device of claim 19.

23. 20. The device of claim 18, wherein the process is configured to, in response to authenticating the second device participating in the AR call, send to the second device a key used to decrypt the one or more digital assets.

24. 1. A device for participating in an augmented reality (AR) call, comprising: a memory configured to store a decryption key; a processing system including one or more processors implemented in circuitry, said processing system comprising: receiving a request from a first client device participating in the AR call to access one or more protected digital assets of a second client device participating in the AR call; receiving, from the second client device, authorization to provide the first client device with access to the one or more protected digital assets; configured to, in response to the authorization from the second client device, provide to the first client device one of the decryption keys, the decryption key being associated with the one or more protected digital assets. device.

25. 25. The device of claim 24, wherein the processing system is further configured to receive the decryption key from the first client device.

26. 25. The device of claim 24, wherein the processing system is configured to receive data from the second client device representing one or more participants in the AR call who are authorized to access the one or more protected digital assets.

27. 25. The device of claim 24, wherein the processing system is configured to receive a shared session token for the AR call, the shared session token being shared by the first client device and the second client device.

28. 25. The device of claim 24, wherein the processing system is configured to receive a participant list in the AR call, the participant list indicating that each participant in the participant list is authorized to access the one or more protected digital assets, the participant list including the first client device.

29. 25. The device of claim 24, wherein the processing system is configured to receive identification information of a first client device, the identification information including one of an Internet Protocol (IP) address or Session Information Protocol (SIP) information.

30. 25. The device of claim 24, wherein the device comprises a Digital Rights Management (DRM) server device.