Provide detailed device capabilities for split rendering negotiation

By transmitting capability information between AR devices and split rendering servers, the lack of specifications for capturing and sending view location information is resolved, network and server overhead is reduced, and the efficiency of the split rendering process is optimized.

CN122139350APending Publication Date: 2026-06-02TENCENT AMERICA LLC

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2025-01-21
Publication Date
2026-06-02

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

A method and apparatus comprising computer code configured to cause one or more processors to: determine the capabilities of an augmented reality (AR) device, the capabilities indicating either an audio encoder or decoder of the AR device or a video encoder or decoder of the AR device as part of the AR device acting as a split rendering client (SRC); obtain AR data of at least one media component of audio and video based on a request to a server signaling the capabilities of the AR device; provide the AR data, the server acting as a split rendering server (SRS) to perform split rendering of the media components between the SRS and the SRC; and process the AR data based on the capabilities of the AR device signaled to the SRC.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references to related applications

[0001] This application claims priority to U.S. Provisional Application No. 63 / 623,763, filed January 22, 2024, and U.S. Application No. 19 / 028,708, filed January 17, 2025, the contents of which are incorporated herein by reference in their entirety. Technical Field

[0002] This disclosure provides a method for providing device capabilities for negotiation between devices and networks to establish split rendering sessions. Background Technology

[0003] 3GPP has work items related to split rendering of media delivery services, in which client media functions are split between the device and the network edge. As a result, the client runs a lighter, less demanding process and can receive more complex applications and services. Conversely, the edge network receives the media, decodes it, and partially renders it into a simpler form, allowing the client to run a lighter-weight process.

[0004] 5G augmented reality devices require intensive processing, including multiple parallel media decoding and possible media encoding, scene compositing, and augmented reality rendering.

[0005] When an application and / or application service provider decides to run client-side media functionality in a split-rendering manner, the application and / or application service provider must replace the functionality with two new modules: 1. an edge-dependent lightweight media service client, and 2. a media processing application running on 5GMS AS.

[0006] The current specification defines split rendering configuration parameters, including expected view information. It is expected that the device will send this view's position information back to the network. However, the configuration parameters do not define how frequently or in what increments this position information should be captured and sent to the network.

[0007] To address one or more different technical problems, this disclosure provides technical solutions to reduce network overhead and server computational overhead, while also providing options for applying various operations to parsed elements, thereby improving their usability and some aspects of technical signaling characteristics when using these operations.

[0008] Furthermore, for any of these reasons, there is an expectation of technical solutions to such problems that arise in video coding techniques. Summary of the Invention

[0009] A method and apparatus are included, comprising a memory and one or more processors, the memory being configured to store computer program code, and the one or more processors being configured to access and operate according to instructions in the computer program code. The computer program is configured to cause the processors to implement acquisition code, the acquisition code being configured to cause at least one processor to perform a conversion between a stream of visual media files and visual media data according to format rules, the format rules indicating: determining the capabilities of an augmented reality (AR) device, the capabilities indicating either an audio encoder or decoder of the AR device and a video encoder or decoder of the AR device as part of the AR device acting as a split rendering client (SRC); obtaining AR data of at least one media component of audio and video based on a request to a server signaling the capabilities of the AR device; providing the AR data, the server acting as a split rendering server (SRS) to enable split rendering of the media components between the SRS and the SRC; and processing the AR data based on the capabilities of the AR device signaled to the SRC.

[0010] A request for the ability to signal an AR device may include a profile identifier, which is in the form of a Uniform Resource Identifier (URI) and includes the syntax "deviceType".

[0011] Profile identifiers can indicate Extended Reality (XR) runtime and scene manager capabilities.

[0012] The request to signal the capabilities of the AR device may include a list of at least some of the capabilities of the AR device, and at least some of the capabilities of the AR device indicate optional codecs of the AR device.

[0013] At least some of the capabilities of an AR device may also indicate the number of instances and interfaces of optional codecs for the AR device, the extended reality (XR) runtime profile, optional extensions, scene description format, and profiles supported by the scene description format.

[0014] The request to signal the capabilities of the AR device may also include a profile identifier, which is in the form of a Uniform Resource Identifier (URI) and includes the syntax “deviceType”, the syntax for indicating a list of at least some of the capabilities is “deviceDetailedCapabilities”, and the profile identifier may indicate extended reality (XR) runtime and scene manager capabilities in addition to at least some of the capabilities of the AR device. Attached Figure Description

[0015] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which: Figure 1 This is a simplified block diagram of a communication system according to an embodiment.

[0016] Figure 2 This is a simplified illustration of the encoder and decoder environment according to an embodiment.

[0017] Figure 3 This is a simplified block diagram related to the decoder according to an embodiment.

[0018] Figure 4 This is a simplified block diagram related to the encoder according to an embodiment.

[0019] Figure 5 This is a simplified block diagram of an AR system according to an embodiment.

[0020] Figure 6 This is a simplified block diagram of a 5G system according to an embodiment.

[0021] Figure 7 This is a simplified block diagram in a 5G UE environment according to an embodiment.

[0022] Figure 8 This is a simplified block diagram of the EDGAR environment according to an embodiment.

[0023] Figure 9 This is a simplified block diagram of the EDGAR UE environment according to an embodiment.

[0024] Figure 10 This is a simplified diagram of the use of AR according to an embodiment.

[0025] Figure 11 This is a simplified block diagram of a hybrid AR and non-AR system according to an embodiment.

[0026] Figure 12 This is a simplified block diagram of a non-AR UE according to an embodiment.

[0027] Figure 13 These are simplified block diagrams and timing diagrams based on the embodiments.

[0028] Figure 14 This is a simplified block diagram of the 5GMS architecture according to an embodiment.

[0029] Figure 15 This is a simplified block diagram of the MSE framework according to an embodiment.

[0030] Figure 16 This is a simplified block diagram of a 5GMSd sensing application according to an embodiment.

[0031] Figure 17 This is a simplified flowchart based on an embodiment.

[0032] Figure 18 This is a simplified block diagram according to an embodiment.

[0033] Figure 19 These are schematic illustrations based on an embodiment. Detailed Implementation

[0034] The proposed features, discussed below, can be used individually or in any combination in any order. Furthermore, embodiments can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors run a program stored in a non-transitory computer-readable medium.

[0035] Figure 1 A simplified block diagram of a communication system 100 according to an embodiment of the present disclosure is shown. The communication system 100 may include at least two terminals 102 and 103 interconnected via a network 105. For unidirectional data transmission, the first terminal 103 may encode video data at its local location for transmission via the network 105 to the other terminal 102. The second terminal 102 may receive the encoded video data from the other terminal from the network 105, decode the encoded data, and display the recovered video data. Unidirectional data transmission is common in applications such as media services.

[0036] Figure 1 A second terminal pair 101 and 104 is shown, provided to support bidirectional transmission of encoded video, which may occur, for example, during a video conference. For bidirectional data transmission, each terminal 101 and 104 can encode video data acquired at a local location for transmission to the other terminal via network 105. Each terminal 101 and 104 can also receive encoded video data transmitted by the other terminal, can decode the encoded data, and can display the recovered video data on a local display device.

[0037] exist Figure 1In this disclosure, terminals 101, 102, 103, and 104 may be shown as servers, personal computers, and smartphones, but the principles of this disclosure are not limited thereto. Embodiments of this disclosure find application in laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network 105 refers to any number of networks, including, for example, wired and / or wireless communication networks, that transmit encoded video data between terminals 101, 102, 103, and 104. Communication network 105 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks (LANs), wide area networks (WANs), and / or the Internet. For the purposes of this discussion, the architecture and topology of network 105 may be irrelevant to the operation of this disclosure unless explained below.

[0038] As an example of the application of the disclosed topic. Figure 2 The placement of a video encoder and video decoder in a streaming environment is illustrated. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0039] The streaming system may include an acquisition subsystem 203, which may include a video source 201 such as a digital camera, creating, for example, an uncompressed video sample stream 213. This sample stream 213 may be emphasized as having a high data volume compared to an encoded video stream, and may be processed by an encoder 202 coupled to the camera 201. The encoder 202 may include hardware, software, or a combination of hardware and software to implement or enforce aspects of the disclosed subject matter as described in more detail below. The encoded video stream 204 may be emphasized as having a lower data volume compared to the sample stream, and may be stored on a streaming server 205 for future use. One or more streaming clients 212 and 207 may access the streaming server 205 to retrieve copies 208 and 206 of the encoded video stream 204. Client 212 may include video decoder 211, which decodes an incoming copy 208 of the encoded video stream and creates an output video sample stream 210 that can be displayed on display 219 or other presentation device (not depicted). In some streaming systems, video streams 204, 206, and 208 may be encoded according to certain video encoding / compression standards. Examples of these standards are mentioned above and further described herein.

[0040] Figure 3 This can be a functional block diagram of a video decoder 300 according to an embodiment of the present invention.

[0041] Receiver 302 may receive one or more codec video sequences to be decoded by decoder 300; in the same embodiment or another embodiment, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. Encoded video sequences may be received from channel 301, which may be a hardware / software link to a storage device storing the encoded video data. Receiver 302 may receive encoded video data and other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective user entities (not depicted). Receiver 302 may separate the encoded video sequences from other data. To prevent network jitter, buffer memory 303 may be coupled between receiver 302 and entropy decoder / resolver 304 (hereinafter referred to as "resolver"). Buffer 303 may not be necessary or may be made smaller when receiver 302 receives data from a store / forward device with sufficient bandwidth and controllability or from an isochronous synchronization network. For use on business packet networks such as the Internet, a buffer 303 may be required. The buffer 303 can be relatively large and, advantageously, can have an adaptive size.

[0042] Video decoder 300 may include parser 304 to reconstruct symbols 313 from an entropy-encoded video sequence. These symbols may include information for managing the operation of decoder 300, and potential information for controlling a presentation device such as display 312, which is not part of the decoder but may be coupled to it. Control information for the presentation device may be in the form of Supplemental Enhancement Information (SEI) messages or fragments of Video Usability Information (VUI) parameter sets (not depicted). Parser 304 may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence may be based on video coding techniques or standards and may follow various principles known to those skilled in the art, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. Parser 304 may extract a subset of parameters from the encoded video sequence for use in the video decoder, based on at least one parameter corresponding to a group. Subgroups can include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), prediction units (PU), etc. The entropy decoder / parser can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0043] The parser 304 can perform entropy decoding / parsing operations on the video sequence received from the buffer 303 to create symbols 313. The parser 304 can receive encoded data and selectively decode specific symbols 313. In addition, the parser 304 can determine whether to provide specific symbols 313 to the motion compensation prediction unit 306, the scaler / inverse transform unit 305, the intra-frame prediction unit 307, or the loop filter 311.

[0044] Depending on the type of encoded video frames or a subset of encoded video frames (e.g., inter-frame and intra-frame frames, inter-frame and intra-frame blocks) and other factors, the reconstruction of symbol 313 may involve multiple different units. Which units are involved and how they are involved can be controlled by the parser 304 through subgroup control information parsed from the encoded video sequence. For clarity, the flow of such subgroup control information between the parser 304 and the various units described below is not depicted.

[0045] In addition to the functional blocks already mentioned, the decoder 300 can be conceptually subdivided into multiple functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with one another. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into multiple functional units as described below.

[0046] The first unit is the scaler / inverse transform unit 305. The scaler / inverse transform unit 305 receives the quantization transform coefficients as symbols 313 from the parser 304, along with control information, including the type of transform to use, block size, quantization factor, and quantization scaling matrix. The scaler / inverse transform unit 305 can output blocks containing sample values, which can be input into the aggregator 310.

[0047] In some cases, the output samples of the scaler / inverse transform unit 305 may belong to intra-coded blocks, i.e., blocks that do not use prediction information from previously reconstructed images, but can use prediction information from previously reconstructed portions of the current image. Such prediction information may be provided by the intra-picture prediction unit 307. In some cases, the intra-picture prediction unit 307 uses surrounding reconstructed information extracted from the current (partially reconstructed) image 309 to generate blocks of the same size and shape as the blocks being reconstructed. In some cases, the aggregator 310 adds the prediction information generated by the intra-picture prediction unit 307 to the output sample information provided by the scaler / inverse transform unit 305 based on each sample.

[0048] In other cases, the output samples of the scaler / inverse transform unit 305 may belong to blocks of inter-frame coding and potential motion compensation. In this case, the motion compensation prediction unit 306 can access the reference image memory 308 to extract samples for prediction. After motion compensation is performed on the extracted samples according to the symbols 313 belonging to the block, these samples can be added by the aggregator 310 to the output of the scaler / inverse transform unit (in this case, referred to as residual samples or residual signals) to generate output sample information. The extraction of prediction samples from the address in the reference image memory by the motion compensation unit can be controlled by motion vectors, which can be provided to the motion compensation unit in the form of symbols 313, which may have, for example, X components, Y components, and reference image components. Motion compensation may also include interpolation of sample values ​​extracted from the reference image memory, motion vector prediction mechanisms, etc., when using subsample precise motion vectors.

[0049] The output samples of aggregator 310 can undergo various loop filtering techniques in loop filter unit 311. Video compression techniques may include in-loop filtering techniques controlled by parameters included in the encoded video bitstream, which can be used by loop filter unit 311 as symbols 313 from parser 304. However, video compression techniques may also respond to metadata obtained during decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0050] The output of the loop filter unit 311 can be a sample stream, which can be output to the display 312 (which can be a presentation device) and stored in the reference image memory 557 for future inter-frame image prediction.

[0051] Once fully reconstructed, certain encoded images can be used as reference images for future predictions. Once the encoded images have been fully reconstructed and the encoded images (via, for example, parser 304) are identified as reference images, the current reference image 309 can become part of the reference image buffer 308, and new current image memory can be reallocated before reconstructing subsequent encoded images begins.

[0052] The video decoder 300 can perform decoding operations according to a predetermined video compression technique that may be documented in standards such as ITU-T H.265 Recommendation. In the sense that the encoded video sequence follows the syntax of a video compression technique or standard, the encoded video sequence may conform to the syntax specified by the video compression technique or standard used, such as that specified in the video compression technique or standard, particularly in a configuration file documented in the video compression technique or standard. For compliance, it may also be necessary to keep the complexity of the encoded video sequence within the limits defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (measured in megasamples per second, for example), maximum reference image size, etc. In some cases, the limitations set by the hierarchy can be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata managed by the HRD buffer, which is represented as a signal in the encoded video sequence.

[0053] In this embodiment, receiver 302 may receive additional (redundant) data when receiving encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by video decoder 300 to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.

[0054] Figure 4 This may be a functional block diagram of a video encoder 400 according to an embodiment of the present disclosure.

[0055] Encoder 400 can receive video samples from video source 401 (not part of the encoder), which can capture video images that will be encoded by encoder 400.

[0056] Video source 401 can provide a source video sequence in the form of a digital video sample stream to be encoded by encoder (303). This digital video sample stream can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit…), any color space (e.g., BT.601 YCrCb, RGB…), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, video source 401 can be a storage device storing previously prepared video. In a video conferencing system, video source 401 can be a camera capturing local image information as a video sequence. The video data can be provided as multiple individual pictures, which are given motion when viewed sequentially. The pictures themselves can be organized into a spatial pixel array, where each pixel can include one or more samples, depending on the sampling structure, color space, etc., used. Those skilled in the art will readily understand the relationship between pixels and samples. The following focuses on describing samples.

[0057] According to an embodiment, encoder 400 can encode and compress images of a source video sequence into an encoded video sequence 410 in real time or under any other time constraints required by the application. Implementing an appropriate encoding rate is a function of controller 402. The controller controls and is functionally coupled to other functional units described below. For clarity, the coupling is not depicted in the figures. Parameters set by the controller may include rate control related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of pictures (GOP) layout, maximum motion vector search range, etc. Other functions of controller 402 will be readily recognized by those skilled in the art, as these other functions may relate to a video encoder 400 optimized for a particular system design.

[0058] Some video encoders operate within a loop that is readily recognized by those skilled in the art as an "encoding loop." As an oversimplification, an encoding loop may include the encoding portion of an encoder (e.g., source encoder 403) responsible for creating symbols based on the input image to be encoded and a reference image, and a (local) decoder 406 embedded in encoder 400. Decoder 406 reconstructs the symbols to create sample data in a manner similar to how a (remote) decoder also creates sample data (because in the video compression techniques considered in the disclosed subject matter, any compression between the symbols and the encoded video stream is lossless). This reconstructed sample stream is input to a reference image memory 405. Since decoding of the symbol stream produces bit-accurate results independent of the decoder's location (local or remote), the contents of the reference image buffer are also bit-accurately corresponding between the local and remote encoders. In other words, the reference image samples "seen" by the encoder's prediction portion are exactly the same sample values ​​that the decoder will "see" during prediction. This fundamental principle of reference image synchronization (and the drift that occurs, for example, due to channel errors, when synchronization cannot be maintained) is well known to those skilled in the art.

[0059] The operation of the "local" decoder 406 can be combined with the above. Figure 3 The operation of the "remote" decoder 300 is the same as described in the detailed description. However, a brief reference is also provided. Figure 4 When symbols are available and the entropy encoder 408 and the parser 304 are able to encode / decode the symbols into an encoded video sequence without loss, the entropy decoding portion of the decoder 300, including the channel 301, receiver 302, buffer 303 and parser 304, may not be fully implemented in the local decoder 406.

[0060] The observation that can be made at this point is that, in addition to the parsing / entropy decoding present in the decoder, any decoder technique must necessarily exist in the corresponding encoder in essentially the same functional form. The description of encoder techniques can be simplified because encoder techniques are inverses of the fully described decoder techniques. Only in certain areas is a more detailed description required, which is provided below.

[0061] As part of the operation of the source encoder 403, the source encoder 403 may perform motion-compensated predictive coding, which predictively encodes the input frame by referencing one or more previously encoded frames from the video sequence designated as "reference frames." In this way, the encoding engine 407 encodes the differences between pixel blocks of the input frame and pixel blocks of the reference frame, which may be selected as the prediction reference for the input frame.

[0062] The local video decoder 406 can decode encoded video data of a frame that can be designated as a reference frame, based on symbols created by the source encoder 403. The operation of the encoding engine 407 can advantageously be a lossy process. When the encoded video data can be decoded by the video decoder (… Figure 4 When decoded at (not shown), the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder 406 replicates the decoding process, which can be performed by the video decoder on the reference frame, and allows the reconstructed reference frame to be stored in a reference image memory 405, which can be, for example, a cache. In this way, the encoder 400 can locally store a copy of the reconstructed reference frame that shares common content (no transmission errors) with the reconstructed reference frame that will be obtained by the remote video decoder.

[0063] Predictor 404 can perform a prediction search against encoding engine 407. That is, for a new frame to be encoded, predictor 404 can search in reference image memory 405 for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can be used as appropriate prediction references for the new image. Predictor 404 can operate pixel-by-pixel based on sample blocks to find suitable prediction references. In some cases, as determined by the search results obtained by predictor 404, the input image may have prediction references obtained from multiple reference images stored in reference image memory 405.

[0064] The controller 402 can manage the encoding operations of the video encoder 403, including, for example, setting parameters and subgroup parameters for encoding video data.

[0065] The outputs of all the aforementioned functional units can undergo entropy encoding in the entropy encoder 408. The entropy encoder converts the symbols generated by the various functional units into an encoded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art (e.g., Huffman coding, variable-length coding, arithmetic coding, etc.).

[0066] Transmitter 409 can buffer the encoded video sequence created by entropy encoder 408, thereby preparing it for transmission via communication channel 411, which may be a hardware / software link to a storage device capable of storing the encoded video data. Transmitter 409 can combine the encoded video data from video encoder 403 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0067] Controller 402 manages the operation of encoder 400. During encoding, controller 405 can assign a specific encoded image type to each encoded image, but this may affect the encoding techniques applicable to the corresponding images. For example, images can typically be assigned to any of the following frame types: An intra-frame picture (I-picture) is a picture that can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art will understand the variations of I-pictures and their respective applications and characteristics.

[0068] A predictive picture (P-picture) can be a picture that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses at most one motion vector and reference index to predict sample values ​​for each block.

[0069] Bidirectional predictive images (B-images) can be images that can be encoded and decoded using intra-frame or inter-frame prediction, which uses at most two motion vectors and a reference index to predict sample values ​​for each block. Similarly, multiple predictive images can be used to reconstruct a single block using more than two reference images and associated metadata.

[0070] Source images are typically spatially subdivided into multiple sample blocks (e.g., each with 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively coded with reference to other (already coded) blocks, which are determined by the coding assignments of the corresponding images applied to the blocks. For example, blocks of an I-image can be nonpredictively coded, or blocks of an I-image can be predictively coded (spatial prediction or intra-frame prediction) with reference to already coded blocks of the same image. Pixel blocks of a P-image can be nonpredictively coded with reference to a previously coded reference image via spatial prediction or temporal prediction. Blocks of a B-image can be nonpredictively coded with reference to one or two previously coded reference images via spatial prediction or temporal prediction.

[0071] The video encoder 400 can perform encoding operations according to predetermined video coding technologies or standards such as ITU-T H.265 Recommendation. In operation, the video encoder 400 can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technology or standard used.

[0072] In this embodiment, transmitter 409 may transmit additional data while transmitting encoded video. Source encoder 403 may include such data as part of the encoded video sequence. Additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, video availability information (VUI) parameter set fragments, etc.

[0073] Figure 5 Example 500 of an end-to-end architecture for a standalone AR (STAR) device according to an exemplary embodiment shows a 5G STAR user equipment (UE) receiver 600, a network / cloud 501, and a 5G UE (transmitter) 700. Figure 6 This is a further detailed example 600 of one or more configurations for a STAR UE receiver 600 according to an exemplary embodiment, and Figure 7 This is a further detailed example 700 of one or more configurations for a 5G UE transmitter 700 according to an exemplary embodiment. 3GPP TR26.998 defines support for glass-type augmented reality / mixed reality (AR / MR) devices in 5G networks. According to the exemplary embodiments herein, at least two device categories are considered: 1) devices fully capable of decoding and playing complex AR / MR content (standalone AR or STAR); and 2) devices with smaller computing resources and / or smaller physical size (and therefore smaller battery), which can only run such applications if most of the computation is performed on a 5G edge server, network, or cloud rather than on the device (edge-dependent AR or EDGAR).

[0074] According to an exemplary embodiment, as described below, a shared conversational use case can be experienced in which all participants in a shared AR conversational experience have AR devices, and each participant sees other participants in an AR scene, where the participants are overlaid on a local physical scene, and the arrangement of the participants in the scene is consistent across all receiving devices. For example, people in each local space have the same position / seat arrangement relative to each other, and such a virtual space creates the feeling of being in the same space, but the room varies from participant to participant because the room is the actual room or space where each person is physically located.

[0075] For example, according to the target Figure 5 to Figure 7In the exemplary embodiment shown, the immersive media processing function on the network / cloud 501 receives uplink streams from various devices and synthesizes a scene description that defines the arrangement of each participant in a single virtual meeting room. The scene description, along with the encoded media stream, is sent to each receiving participant. The receiving participant's 5G STARUE 600 receives, decodes, and processes the 3D video and audio streams, and renders the 3D video and audio streams using the received scene description and information received from its AR runtime, thereby creating an AR scene of the virtual meeting room with all other participants. While each participant's virtual room is based on their own physical space, the seating / positioning of all other participants in the room is consistent with the virtual room of each other participant in the session.

[0076] According to an exemplary embodiment, see also Figure 8 , Figure 8 An example 800 related to the EDGAR device architecture is shown, where the device itself, such as the 5G EDGAR UE 900, cannot perform heavy processing. Therefore, scene parsing and media parsing for the received content are performed at the cloud / edge 801, and then a simplified AR scene with a small amount of media components is delivered to the device for processing and rendering. Figure 9 A more detailed example of a 5G EDGAR UE 900 according to an exemplary embodiment is shown.

[0077] Figure 10 Example 1000 is shown, in which users A10, B11, and T12 will participate in an AR conference room, and one or more of these users may not have an R device. As shown, user A10 is located in their office 1001, sitting in a conference room with a variety of chairs, and user A10 is using one of the chairs. User B11 is located in their living room 1002, sitting on a two-seater sofa, and their living room also contains one or more sofas for two people, as well as other furniture such as chairs and tables. User T12 is located in an airport lounge 1003, sitting on a bench opposite one of one or more other coffee tables.

[0078] In the AR environment, in office 1001, user A10's AR displays virtual user B11v1 corresponding to user B11 and virtual user T12v1 corresponding to user T12, and both virtual users B11v1 and T12v1 are shown to user A10 as sitting on furniture in office 1001, i.e., office chairs, just like user A10. In living room 1202 of example 1200, user B11's AR displays virtual user T12v2 corresponding to user T12, but virtual user T12v2 is sitting on a sofa in living room 1202, and also displays virtual user A10v1 corresponding to user A10, who is also sitting on furniture in living room 1202, not on an office chair in office 1201. In airport lounge 1203, user T 12's AR representation shows a virtual user A 10v2 corresponding to user A 10, but virtual user A 10v2 is seated at a table in lounge 1203. Virtual user B 11v2 is also shown, seated at a table opposite virtual user A 10v2. In each of these offices 1201, lounges 1202, and airport lounge 1203, the updated scene description of each room is consistent with the other rooms in terms of location / seating arrangement. For example, in each room, user A 10 is shown as being counter-clockwise relative to user 11 or its virtual representation, and user 11 or its virtual representation is also shown as being clockwise relative to user T 12 or its virtual representation.

[0079] However, AR technology is limited in any attempt to incorporate virtual spaces into devices that do not support AR but can resolve VR or 2D video. The embodiments in this paper provide an improved technical process to create virtual scenes consistent with AR scenes when such devices participate in shared AR dialogue services.

[0080] Figure 11 Example 1100 is shown, illustrating an end-to-end architecture with a non-AR device 1101 and a cloud / edge 1102 according to an exemplary embodiment. Figure 12 A further detailed block diagram example of a non-AR device 1101 is shown.

[0081] like Figure 11 and Figure 12As shown, the non-AR UE 1101 is a device capable of rendering 360-degree or 2-D video but without any AR capabilities. However, the edge functionality on the cloud / edge 1102 enables AR rendering of the received scene, rendering the scene along with immersive visual and audio objects in a virtual room selected from a library. The entire video is then encoded and transmitted to device 1101 for decoding and rendering.

[0082] Therefore, multi-view capability is possible; for example, AR processing on the edge / cloud 1102 can generate multiple videos of the same virtual room: multiple videos of the same virtual room generated from different angles and with different viewports. Device 1101 can receive one or more of these videos, switch between them as needed, or send commands to the edge / cloud processing to stream only the desired viewport / angle.

[0083] Furthermore, the ability to change the background is possible, whereby the user associated with device 1101 can select a desired room background from a provided library, such as a meeting room from different conference rooms, or even a living room and layout. The cloud / edge 1102 uses the selected background and creates a virtual room accordingly.

[0084] Figure 13 An exemplary timing diagram 1300 is shown for an exemplary call flow for an immersive AR conversation received by a non-AR UE 1101. For illustrative purposes, only one transmitter is shown in this diagram, and its detailed call flow is not shown.

[0085] The diagram shows an AR application module 21, a media playback module 22, and a media access function module 23. These modules 21, 22, and 23 can be considered as modules that receive non-AR UE 1101. A cloud / edge split rendering module 24 is also shown. A media delivery module 25 and a scene graph synthesizer module 26 for each element in the network cloud 1102 are also shown. A 5G transmitter UE module 700 is also shown.

[0086] S1 to S6 can be considered as the session establishment phase. At S1, the AR application module 21 can request to start a session with the media access function module 23, and at S2, the media access function module 23 can request to start a session with the cloud / edge split rendering module 24.

[0087] At S3, the cloud / edge split rendering module 24 can negotiate a session with the scene graph compositor module 26, and correspondingly, the scene graph compositor module 26 can negotiate with the 5G transmitter UE 700. If the negotiation is successful, at S5, the cloud / edge split rendering module can send an acknowledgment to the media access function module 23, and the media access function module 23 can send an acknowledgment to the AR application module 21.

[0088] Next, S7 can be considered the media pipeline configuration phase, during which each of the media access function module 23 and the cloud / edge split rendering module 24 configures its respective pipeline. Then, after this pipeline configuration, a session can be initiated via signals from the AR application module to the media player module 22 at S8, from the media player module 22 to the media access function module 23 at S9, and from the media access function module 23 to the cloud / edge split rendering module 24 at S10.

[0089] Then, there may be a pose loop phase from S11 to S13. In the pose loop phase, at S11, pose data can be provided from the media player module 22 to the AR application module 21, and at S12, the AR application module can provide pose data 12 to the media access function module 23. After that, the media access function module 23 can provide pose data to the cloud / edge split rendering module 24.

[0090] S14 to S16 can be considered as the shared experience stream phase. During this phase, at S14, the 5G transmitter UE700 can provide a media stream to the media delivery module 25, and at S15, it can provide AR data to the scene graph synthesizer module 26. Then, the scene graph synthesizer module 25 can synthesize one or more scenes based on the received AR data, and at S16, it provides scenes and scene updates to the cloud / edge split rendering module 24. At S17, the media delivery module 25 can also provide a media stream to the cloud / edge split rendering module. According to an exemplary embodiment, this can include: obtaining an AR scene descriptor from a non-AR device that does not render the AR scene, and generating a virtual scene by the cloud device through parsing and rendering the scene description obtained from the non-AR device.

[0091] S18 and S19 can be considered as the media uplink stage. In the media uplink stage, the media player module 22 collects and processes media data from its local user and provides the media data to the media access function module 23 at S18. Then, the media access module 23 can encode the media and provide the media stream to the cloud / edge split rendering module 24 at S19.

[0092] The media downlink phase can be considered to be located between S19 and S20. During the media downlink phase, the cloud / edge split rendering module 24 can perform scene parsing and complete AR rendering. Afterwards, S20 and S21 can be considered to constitute the media stream loop phase. At S20, the cloud / edge split rendering module 24 can provide the media stream to the media access function module 23. Then, the media access function module 23 can decode the media and provide media rendering to the media player 22 at S21.

[0093] By means of such features according to the exemplary embodiments, the non-AR UE 1101 can still utilize its own display capable of rendering VR or 2-D video, even if it does not have a perspective display and therefore cannot create AR scenes. Therefore, the immersive media processing capabilities of the non-AR UE 1101 only generate a common scene description, describing the relative position of each participant with respect to other participants and the scene. As described above, pose information needs to be used to adjust the scene itself at each device before rendering it as an AR scene. An AR rendering process at the edge or in the cloud can parse the AR scene and create a simplified VR-2D scene.

[0094] According to an exemplary embodiment, this disclosure uses a similar split rendering process for non-AR devices (e.g., VR or 2-D video devices) with EDGAR devices, whose characteristics (e.g., edge / cloud AR rendering processes) do not produce any AR scene in this case. Instead, a virtual scene is generated by parsing and rendering a scene description for a given background (e.g., a conference room) received from an immersive media processing function, and then rendering for each participant located at the position described in the scene description within the conference room.

[0095] Furthermore, depending on the ability to receive non-AR devices, the resulting video can be a 360-degree video or a 2-D video, and according to an exemplary embodiment, the location information received from the non-AR device is taken into account to generate the resulting video.

[0096] In addition, each other participant using a non-AR device was added as a 2D video overlay on the 360 / 2D video of the conference room, for example... Figure 10 As shown, the room may have areas specifically for using these overlays, such as areas for furniture that overlay virtual images, like... Figure 10 As shown.

[0097] Furthermore, if needed, audio signals from all participants can be mixed to create a single-channel audio carrying speech in the room, and video can be encoded into a single 360-degree video or 2-D video and transmitted to the device. Optionally, according to an exemplary embodiment, multiple video (multi-view) sources can be created, each capturing the same virtual meeting room from different views and providing these views to the device.

[0098] In addition, the non-AR UE device 1101 can receive 360 ​​video and / or one or more selected multi-view videos, along with received audio, and render them on the device display. Users can switch between different views or change the viewport of the 360 ​​video by moving or rotating the viewing device, so that users can navigate in a virtual room while watching the video.

[0099] While the above embodiments are configured with such a 5G Media Streaming Architecture (5GMS) extension to use edge servers in the 5GMS extended architecture, and while the 5GMS extension specification may have many features, such features are technically not deployable as a set of software development kits (SDKs) on a device or a set of microservices in the cloud. Such technical deficiencies are addressed by embodiments further described below.

[0100] For example, the current Media Service Enabler Technology Report does not define a framework for associating specifications with SDKs, nor does it include any concept of microservices.

[0101] See Figure 14 Example 1400 illustrates a 5G media streaming architecture with edge extension according to an exemplary embodiment. As shown, a User Equipment (UE) 1401 and a Data Network (DN) 1411 are present. UE 1401 may include a 5GMS client 1403 and a 5GMS-aware application 1405, such as the AR or non-AR embodiments described above, and these applications are not limited thereto. The 5GMS client 1403 may also include a media stream processor 1403 and a media session processor. DN 1411 may include a 5GMS Application Server (AS) 1412, a 5GMS Application Function (AF) 1414, and a 5GMS Application Provider 1413. The 5GMS AF 1414 may also communicate with a Network Open Function (NEF) 1415 and a Policy and Charging Function (PCF) 1416. The improvements described herein can be understood in the context of at least one or more of UE 1401, 5GMS-aware application 1405, 5GMS application provider 1413, NEF 1415, and PCF 1416; that is, one or more of these elements can generate their own specifications and provide such specifications to another of these elements upon a compliance request, and these elements can further configure the initial element specifications according to various possibilities such as those further described below, rather than having a monolithic specification (a monolithic specification does not have a definition of a Media Services Enabler (MSE) for each function or group of functions), such processing can be performed through any one or more of the multiple open application programming interfaces (APIs) 1420 and interfaces M1, M2, M4, M5, M6, M7, M8, N33, and N5 shown.

[0102] According to an exemplary embodiment, Figure 15 An example 1500 of the MSE framework is illustrated in two parts: the MSE specification 1501 and the MSE implementation 1502. The MSE implementation 1502 includes an MSE SDK abstraction layer 1510, an MSE SDK (platform-dependent) instantiation 1520, and an MSE microservice 1530. The MSE SDK abstraction layer 1510 may have a configuration API abstraction layer 1513, a control interface 1512, and a media interface 1511. Each of the configuration API abstraction layer 1513, control interface 1512, and media interface 1511 can notify the MSE SDK (platform-dependent) instantiation 1520 and its control interface 1522 and media interface 1521 to the configuration API 1523. The MSE microservice 1530 may also have a configuration API 1533, a control interface 1532, and a media interface 1531.

[0103] According to an exemplary embodiment, MSE specification 1501 defines (a) a media layer and (b) an MSE configuration that provides the following technical advantages. (a) The media layer may indicate any of the following: (i) a functional description of the MSE, including mandatory and optional features; (ii) a control interface, such as provisioning, authentication used by applications, and other functions that interact with the MSE; (iii) a media interface including all input and output formats and protocols; (iv) a network interface including systems and wireless networks; (v) events, notifications, reports, and monitoring; and (vi) error handling. (b) The MSE configuration may indicate: (i) an MSE description document (MDD) (describing (1) the functions supported by the MSE implementation and their configuration parameters, and (2) optionally, performance / cost metrics for different features / options); (ii) an MSE configuration API (MCA) abstraction layer (indicating information for (1) retrieving the description document, (2) configuring the MSE instantiation, and (3) retrieving the status and condition of the MSE instantiation); and (iii) a service API for the MSE configuration API (MCA) abstraction layer. MSE implementation 1502 may include any of the following: (a) MSE SDK abstraction layer 1510 (indicating: (i) a media layer conforming to MSE specification 1501; and (ii) the MSE description document (MDD) and MSE configuration API (MCA) abstraction layer of the aforementioned MSE configuration of MSE specification 1501); (b) MSE SDK (platform dependent) instantiation 1520 (indicating an SDK implementation in a specific environment that conforms to: (i) a media layer conforming to MSE specification 1501; and (ii) an MSE description document (MDD) and conforming to MSE specification 1501). Unlike SDK abstraction layer 1530, only a specific MSE Configuration API (MCA) abstraction layer for the MSE configuration described above in MSE specification 1501 exists; and (c) MSE microservice 1510, which is an MSE implementation of a microservice and indicates: (i) a media layer conforming to the media layer of MSE specification 1501; and an MSE Description Document (MDD) and MSE Configuration API (MCA) abstraction layer for the MSE configuration described above in MSE specification 1501, for the service API of the MSE Configuration API (MCA) abstraction layer. Conformity can be determined, for example, by matching a format or syntax or by indicating a flag indicating such a format or syntax.

[0104] For example, according to this favorable MSE specification (which does not exist in the 3GPP SA4 specification), for any SDK or microservice conforming to the MSE specification, according to an exemplary embodiment, a description of features and their configuration parameters can be retrieved via an external function or service. This external function or service can set a specific configuration for running the SDK and can retrieve the status and condition of the running SDK at any time. The status and condition can indicate availability, running, busy, idle, offline, etc.

[0105] From this perspective, Figure 16 An implementation example 1600 of an exemplary embodiment is shown. In implementation example 1600, there is a 5G Media Streaming Downlink (5GMSd) Aware Application 1405, which communicates with a 5GMSd Client 1402'. The 5GMSd Client 1402' has a Media Session Processor 1404, which communicates with a 5GMSd AF 1414'. Method 1601, notification and error 1602, and status 1603 information can be transmitted between the 5GMSd sensing application 1405' and the 5GMSd client 1402'. Status 1604, subscription and notification 1605, and setting and configuration information 1606 can be transmitted to and from the media session processor 1404. The media session processor 1404 itself can communicate with the 5GMSd AF1414' via link M5d. According to an exemplary embodiment, Figure 16The MSE specification of the Media Session Processor (MSH) example 1600 shown can describe: (a) a media layer (e.g., (i) a functional description ((1) service access information, (2) consumption reports, (3) metric reports, (4) dynamic policies and (5) network auxiliary information) and (ii) M5d, M6d and M7d information); and (b) MSE configuration information, such as (i) an MDD (describes (1) an identifier indicating that the MSE conforms to the media layer, (2) optional features of the functional description and M5d, M6d and M7d and their configuration parameters, and (3) optionally, performance / cost metrics for different features / options), (ii) an API abstraction layer (indicates information for (1) retrieving the description document MDD, (2) configuring the MSE instantiation and (3) retrieving the status and condition of the MSE instantiation), and (iii) a service API for the API abstraction layer. For example, according to an exemplary embodiment, the MSE SDK implementation of the above-mentioned MSE specification located in Android may support the following: (i) the media layer, the above-mentioned media layer of the MSE specification conforming to MSE; and (ii) the above-mentioned MDD of the MSE configuration of the MSE specification, and a specific implementation of the API abstraction layer of the MSE configuration of the MSE specification for MSH.

[0106] According to an exemplary embodiment, see Figure 17Flowchart 1700 describes the features implemented by the MSE SDK in the MDD. The API configured by the MSE allows external Android processes to retrieve this document at S1701 and configure the SDK at S1713 using the configurable parameter set described in the MDD and the state and status of the running SDK. The state and status of the running SDK can be determined at S1714. The language and syntax of the MDD and the general framework of the MCA can be uniformly defined for all SA4 Media Service Enabler specifications at S1701, with only specific code points for each specification defined in this specification. External functions or applications that understand the MDD language and syntax and support the MCA can retrieve information from the MSE implementation at S1701, and if the MSE specification identifier is identified at S1702, the external function or application can understand that it can parse and process the MDD and its configuration parameters. For example, an MSE specification and its identifier can be generated at S1701. After a request for such a specification is made, a conformity check can determine whether the expected identifier is provided. Then, at S1703, the MSE can be deployed based on the generated parameters and various capability and cost metrics at S1704. At S1705, it can be determined whether the SDK is deployed at the device at S1706 or at the server at S1708. If at S1708, the MDD and service API for the MCA can be set up as described above. Alternatively, if at S1710, it can be determined whether to continue using the MDD and MCA in the usual way or using the MDD and a specific MCA. In any case, such an SDK can then be configured by an external device or function at S1713, and the external device or function can request status S1714 at any time based on the implementation configured at S1715.

[0107] Examples of MDDs can be found in ISO / IEC 23090-8. A Functional Description Document (MDD) is a JSON document that describes the functional aspects and features provided by the function, along with its configuration parameters.

[0108] Therefore, an exemplary embodiment provides a method for defining a media service enabler SDK and services, wherein the functions and features of the media service enabler are described in a document that can be retrieved by an external function, application, or service, and in configuration parameters for each feature, wherein an external entity can use the same interface used to retrieve the information to establish a specific configuration of the media service enabler, wherein the complexity or other metrics of various functions of the media service enabler are also described in the document, and wherein the media service enabler specification defines the format of the description document and a unique identifier for the specific media service enabler.

[0109] According to an embodiment, Figure 18Example 1800 provides the 3GPP SR_MSE split rendering architecture and split rendering configuration parameters, including desired view information. Further, the embodiment provides added device capabilities to the split rendering configuration information, such as those shown in Table 1: Table 1

[0110]

[0111] According to the embodiment, detailed device capabilities indicate two levels of capabilities: a high-level device type and a low-level detailed device capability, as shown in Table 2: Table 2

[0112]

[0113] According to an embodiment, during the negotiation between the Split Rendering Client (SRC) and the Split Rendering Server (SRS), the SRC provides the SRS with device capabilities as part of the split rendering configuration. The SRS defines how to split the session between the network and the device by analyzing the device capabilities (whether or not high-level device types or additional detailed capabilities are provided) and the splitRenderingProfile. Because the device capability definition contains more information than the splitRenderingProfile, the SRS can optimize session splitting based on the device's capabilities and does not limit split rendering to the defined split rendering profile.

[0114] Therefore, a method is provided for a device to signal augmented reality device capabilities as part of a split rendering configuration to the network during device-network negotiation to establish a split rendering session, wherein the split rendering client provides optional details related to the number and maximum number of codecs supported by the device, as well as XR runtime and scene manager capabilities, to the device's supported device types / profiles. With this information, the split rendering server can determine what is optimally optional for the split rendering session.

[0115] The XR runtime includes the functionality and hardware components that exist on the XR device. However, these functionalities and hardware components are not directly exposed to XR applications. Instead, the XR runtime provides its functionality and hardware components through the XR system. A single XR runtime can expose more than one XR system for different purposes; for example, a handheld device may have two XR systems, one existing when the user holds the device and one existing when the device is inserted into the HMD. When an XR application starts, it is expected to query which XR systems are available on the XR device and select one of them to create an XR session.

[0116] The split rendering client establishes an XR session locally based on device configuration and user selection. The SR client defines view configurations (e.g., single-view or stereoscopic view), projection formats (e.g., projection, isometric rectangle, quadrilateral patch, or cube map), exchange chain image configurations, etc. Furthermore, XR space and action configurations are negotiated between the SR client and server. This includes defining a common XR space and defining and selecting actions and action sets. The format is extensible to support the exchange of additional / future configuration information.

[0117] The above-described technology can be implemented as computer software that uses computer-readable instructions and is physically stored in one or more computer-readable media, or implemented by one or more specially configured hardware processors. For example, Figure 19 A computer system 1900 suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0118] Computer software can be coded using any suitable machine code or computer language. Any suitable machine code or computer language can be assembled, compiled, linked, or similarly used to create code containing instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), or through interpretation, microcode execution, etc.

[0119] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0120] Figure 19 The components of the computer system 1900 shown are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of this disclosure. The configuration of the components should also not be construed as having any dependencies or requirements relating to any component or combination of components shown in the exemplary embodiments of the computer system 1900.

[0121] Computer system 1900 may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., speech, clapping), visual input (e.g., gestures), and olfactory input (not depicted). Human-machine interface devices may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, images captured from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0122] Human-machine interface input devices may include one or more of the following (only one of each is depicted): keyboard 1901, mouse 1902, touchpad 1903, touch screen 1910, joystick 1905, microphone 1906, scanner 1908, and camera 1907.

[0123] Computer system 1900 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback of touch screen 1910, or joystick 1905, but may also be tactile feedback devices that are not input devices), audio output devices (e.g., speakers 1909, headphones (not depicted)), and visual output devices (e.g., screen 1910 including CRT screens, LCD screens, plasma screens, OLED screens, each screen may or may not have touch screen input capability, each screen may or may not have tactile feedback capability, some of which are capable of outputting two-dimensional visual output or more than three-dimensional output through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays and smoke boxes (not depicted), and printers (not depicted).

[0124] The computer system 1900 may also include human-machine-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 1920 with media such as CD / DVD 1911, finger drives 1922, removable hard disk drives or solid-state drives 1923, conventional magnetic media such as magnetic tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.

[0125] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.

[0126] Computer system 1900 may also include an interface 1999 providing access to one or more communication networks 1998. Network 1998 may be, for example, a wireless network, a wired network, or an optical network. Network 1998 may further be a local area network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a latency-tolerant network, etc. Examples of network 1998 include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks 1998 typically require an external network interface adapter (e.g., a USB port of computer system 1900) attached to certain general-purpose data ports or peripheral buses (1850 and 1951); other network interfaces are typically integrated into the core of computer system 1900 by being attached to system buses as described below (e.g., an Ethernet interface connected to a PC computer system or a cellular network interface connected to a smartphone computer system). Computer system 1900 can communicate with other entities using any of these networks 1998. Such communication can be one-way receiving (e.g., broadcast television), one-way transmitting (e.g., CANBus to certain CANBus devices), or bidirectional, such as using a local area network (LAN) or wide area network (WAN) to other computer systems. Certain protocols and protocol stacks can be used on each of those networks and network interfaces as described above.

[0127] The human-machine interface devices, human-machine accessible storage devices, and network interfaces mentioned above can be attached to the kernel 1840 of the computer system 1900.

[0128] The core 1940 may include one or more central processing units (CPUs) 1941, graphics processing units (GPUs) 1942, graphics adapters 1917, dedicated programmable processing units in the form of field-programmable gate areas (FPGAs) 1943, hardware accelerators 1944 for certain tasks, etc. These devices, along with read-only memory (ROM) 1945, random access memory 1946, and internal mass storage 1947 such as internal non-user-accessible hard disk drives (SDs), etc., can be connected via the system bus 1948. In some computer systems, the system bus 1948 can be accessed via one or more physical connectors to allow for expansion with additional CPUs, GPUs, etc. Peripheral devices can be directly attached to the core's system bus 1948 or attached to the core's system bus 1948 via a peripheral bus 1951. Peripheral bus architectures include PCI, USB, etc.

[0129] The CPU 1941, GPU 1942, FPGA 1943, and accelerator 1944 can execute certain instructions, which can be combined to form the computer code mentioned above. This computer code can be stored in ROM 1945 or RAM 1946. Transient data can also be stored in RAM 1946, while permanent data can be stored, for example, in internal mass storage 1947. Fast storage and retrieval to any storage device can be achieved using a cache, which can be closely associated with one or more CPUs 1941, GPUs 1942, mass storage 1947, ROM 1945, RAM 1946, etc.

[0130] Computer-readable media may have computer code thereon that performs various computer-implemented operations. The media and computer code may be media and computer code specifically designed and constructed for the purposes of this disclosure, or the media and computer code may be of a type known and available to those skilled in the art of computer software.

[0131] By way of example, and not limitation, the architecture corresponding to computer system 1900, particularly kernel 1940, enables functionality because one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) execute software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage as described above, as well as some non-transitory memory of kernel 1940, such as internal kernel mass storage 1947 or ROM 1945. Software implementing various embodiments of this disclosure can be stored in such devices and executed by kernel 1940. Depending on specific needs, computer-readable media may include one or more storage devices or chips. The software can cause kernel 1940, particularly the processors (including CPUs, GPUs, FPGAs, etc.) within kernel 1940, to execute specific processes described herein or specific portions of those processes, including defining data structures stored in RAM 1946 and modifying such data structures according to software-defined processes. Additionally or alternatively, due to logic hardwired or otherwise embodied in the circuitry (e.g., Accelerator 1944), the computer system provides functionality that circuitry can replace or operate with the software to perform a particular process described herein or to perform a particular portion of the particular process described herein. Where appropriate, references to software may include logic, and vice versa. Where appropriate, references to computer-readable media may include circuitry storing software for execution (e.g., integrated circuits (ICs)), circuitry embodying logic for execution, or both. This disclosure includes any suitable combination of hardware and software.

[0132] While several exemplary embodiments have been described in this disclosure, there are modifications, substitutions, and various equivalent alternatives that fall within the scope of this disclosure. Therefore, it should be appreciated that those skilled in the art will be able to design numerous systems and methods that, while not explicitly shown or described herein, embody the principles of this disclosure and thus fall within its spirit and scope.

Claims

1. A video decoding method, the method being executed by at least one processor, and comprising: Determine the capabilities of an augmented reality (AR) device, wherein the capabilities indicate either the AR device’s audio encoder or decoder or the AR device’s video encoder or decoder as part of the AR device used as a split rendering client (SRC); Based on a request to the server to signal the AR device's ability to obtain AR data for at least one media component of audio and video, the AR data is provided, and the server acts as a split rendering server (SRS) to perform split rendering of the media component between the SRS and the SRC. as well as The AR data is decoded based on the ability of the AR device to notify the SRC via signals.

2. The method according to claim 1, wherein, The request to signal the capability of the AR device includes a profile identifier, which is in the form of a Uniform Resource Indicator (URI) and includes the syntax "deviceType".

3. The method according to claim 2, wherein, The configuration file identifier indicates the Extended Reality (XR) runtime and scene manager capabilities.

4. The method according to claim 1, wherein, The request to signal the capabilities of the AR device includes a list of at least some of the capabilities of the AR device, and at least some of the capabilities of the AR device indicate optional codecs of the AR device.

5. The method according to claim 4, wherein, At least some of the capabilities of the AR device also indicate the number of instances and interfaces of the AR device's optional codecs, extended reality (XR) runtime profiles, optional extensions, scene description formats, and profiles supported by the scene description formats.

6. The method according to claim 5, wherein, The request to signal the capabilities of the AR device also includes a profile identifier, which is in the form of a Uniform Resource Indicator (URI) and includes the syntax "deviceType", and the syntax for the list indicating at least some of the capabilities is "deviceDetailedCapabilities".

7. The method according to claim 6, wherein, The profile identifier indicates extended reality (XR) runtime and scene manager capabilities in addition to at least some of the capabilities of the AR device.

8. A video encoding method, the method being executed by at least one processor, and comprising: Determine the capabilities of an augmented reality (AR) device, wherein the capabilities indicate either the AR device’s audio encoder or decoder or the AR device’s video encoder or decoder as part of the AR device used as a split rendering client (SRC); Based on a request to the server to signal the AR device's ability to obtain AR data for at least one media component of audio and video, the AR data is provided, and the server acts as a split rendering server (SRS) to perform split rendering of the media component between the SRS and the SRC. as well as The AR data is encoded based on the ability of the AR device to notify the SRC via signals.

9. The method according to claim 8, wherein, The request to signal the capability of the AR device includes a profile identifier, which is in the form of a Uniform Resource Indicator (URI) and includes the syntax "deviceType".

10. The method according to claim 9, wherein, The configuration file identifier indicates the Extended Reality (XR) runtime and scene manager capabilities.

11. The method according to claim 8, wherein, The request to signal the capabilities of the AR device includes a list of at least some of the capabilities of the AR device, and at least some of the capabilities of the AR device indicate optional codecs of the AR device.

12. The method according to claim 11, wherein, At least some of the capabilities of the AR device also indicate the number of instances and interfaces of the AR device's optional codecs, extended reality (XR) runtime profiles, optional extensions, scene description formats, and profiles supported by the scene description formats.

13. The method according to claim 12, wherein, The request to signal the capabilities of the AR device also includes a profile identifier, which is in the form of a Uniform Resource Indicator (URI) and includes the syntax "deviceType", and the syntax for the list indicating at least some of the capabilities is "deviceDetailedCapabilities".

14. The method according to claim 13, wherein, The profile identifier indicates extended reality (XR) runtime and scene manager capabilities in addition to at least some of the capabilities of the AR device.

15. A method for processing visual media data, the method being executed by at least one processor and comprising: The conversion between visual media files and visual media data streams is performed according to format rules, which indicate: Determine the capabilities of an augmented reality (AR) device, wherein the capabilities indicate either the AR device’s audio encoder or decoder or the AR device’s video encoder or decoder as part of the AR device used as a split rendering client (SRC); Based on a request to the server to signal the AR device's ability to obtain AR data for at least one media component of audio and video, the AR data is provided, and the server acts as a split rendering server (SRS) to enable split rendering of the media component between the SRS and the SRC. as well as The AR data is processed based on the ability of the AR device to notify the SRC via signals.

16. The method according to claim 15, wherein, The request to signal the capability of the AR device includes a profile identifier, which is in the form of a Uniform Resource Indicator (URI) and includes the syntax "deviceType".

17. The method according to claim 16, wherein, The configuration file identifier indicates the Extended Reality (XR) runtime and scene manager capabilities.

18. The method according to claim 15, wherein, The request to signal the capabilities of the AR device includes a list of at least some of the capabilities of the AR device, and at least some of the capabilities of the AR device indicate optional codecs of the AR device.

19. The method according to claim 18, wherein, At least some of the capabilities of the AR device also indicate the number of instances and interfaces of the AR device's optional codecs, extended reality (XR) runtime profiles, optional extensions, scene description formats, and profiles supported by the scene description formats.

20. The method according to claim 19, wherein, The request to signal the capabilities of the AR device also includes a profile identifier in the form of a Uniform Resource Indicator (URI) and includes the syntax "deviceType", the syntax for the list indicating at least some of the capabilities is "deviceDetailedCapabilities", and the profile identifier indicating extended reality (XR) runtime and scene manager capabilities in addition to at least some of the capabilities of the AR device.