Procedures for discovering the capabilities and performance of 5G EDGAR devices
The method and apparatus for controlling 5G EDGAR devices through computer program code manage device features via APIs, addressing network and server overhead, and enhancing usability and signaling efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-04-04
- Publication Date
- 2026-03-18
AI Technical Summary
Existing 5G EDGAR devices lack the ability to respond to capability queries from authorized applications over a wireless network, leading to network and server computation overhead, and there is no framework for integrating edge server capabilities with device capabilities.
A method and apparatus that control a 5G edge-dependent augmented reality (EDGAR) device through computer program code, utilizing APIs to manage device features like media encoding, decoding, capture, rendering, and hardware capabilities, and providing real-time execution and battery efficiency information.
Reduces network and server computation overhead by enabling efficient management of 5G EDGAR device capabilities, improving usability and signaling characteristics.
Smart Images

Figure 0007833035000001 
Figure 0007833035000002 
Figure 0007833035000003
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application is based on U.S. Provisional Patent Application No. 63 / 338,767, filed on May 5, 2022, and U.S. Patent Application No. 18 / 192,328, filed on March 29, 2023, claims the priority thereof, and the disclosures of which are hereby incorporated by reference in their entireties into this specification.
[0002] This disclosure can execute complex parts of an application on the edge, light and provides procedures for the capabilities and dynamic performance of 5G EDGAR devices by an application on a device or an application executed on a network (edge application on the edge or application service provider) such that only processing is performed on the device. detection for.
Background Art
[0003] 3GPP TR 26.998 defines the support for glass - type extended reality / composite reality (AR / MR) devices in 5G networks. Primarily, two device classes are considered: 1) devices capable of fully decoding and playing back complex AR / MR content (stand - alone AR, i.e., STAR), and 2) devices with smaller computing resources and / or smaller physical size (thus, battery), and such applications can be executed only when most of the computing is performed on a 5G edge server, network, or Cloud on rather than on the device (edge - dependent AR, i.e., EDGAR). Recent activities regarding the profiling of EDGAR devices have been initiated in 3GPP.
[0004] Even though the current TR 26.998 defines an architecture for EDGAR devices that receive streaming content by performing some of the large-scale computing required on the cloud / edge, it does not have the ability to receive and respond to capability queries made by any authorized application, such as an application on the device, an application on the edge network, or an application service provider over a wireless network. [Overview of the project] [Problems that the invention aims to solve]
[0005] To address one or more different technical problems, this disclosure provides technical solutions to reduce network overhead and server computation overhead, while also offering options to apply various behaviors to the resolved elements, which may improve their usability and some of their technical signaling characteristics. [Means for solving the problem]
[0006] The invention includes a method and apparatus comprising memory configured to store computer program code, and one or more processors configured to access the computer program code and operate as instructed by the computer program code. The computer program code controls at least one processor to control a 5G edge-dependent augmented reality (EDGAR) device and at least some features of the 5G EDGAR device. detection A first control code configured to cause at least one processor to perform at least some of the features of the 5G EDGAR device index Acquisition code configured to obtain at least one processor for at least some features of a 5G EDGAR device index Accordingly, 5G media streaming (5GMS) services execution This includes a second control code configured to control the following:
[0007] According to an exemplary embodiment, the first control code provides at least one processor with at least several features via one of several application programming interfaces (APIs) of the 5G EDGAR device. detection It is configured to control the 5G EDGAR device in order to do so.
[0008] According to an exemplary embodiment, the first control code is configured to cause at least one processor to control the 5G EDGAR device via a network API, To at least one processor via a network API index Retrieval code configured to obtain the data.
[0009] According to exemplary embodiments, at least some features include at least one parameter among a media encoder and a media decoder.
[0010] According to exemplary embodiments, at least some features include a device capture capability of the 5G EDGAR device, where the device capture capability includes either a camera capability or a microphone capability.
[0011] According to exemplary embodiments, at least some features include a device rendering capability of the 5G EDGAR device, where the device rendering capability includes either a speaker capability or a display capability.
[0012] According to exemplary embodiments, at least some features include the raw hardware capabilities of a 5G EDGAR device, which include any of the following: CPU benchmarks, GPU benchmarks, memory bandwidth, and on-device storage.
[0013] According to an exemplary embodiment, at least some features include the functional efficiency of a 5G EDGAR device, where the functional efficiency indicates a relative amount of power consumption per function and a relative amount with respect to parameters per function.
[0014] According to an exemplary embodiment, at least some features include either the state of the power source of a 5G EDGAR device or the rate of battery depletion of the 5G EDGAR device.
[0015] According to an exemplary embodiment, index indicates either whether a combination of at least some features can be executed in real time on a 5G EDGAR device and the estimated overall battery consumption of that combination.
[0016] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.
Brief Description of the Drawings
[0017] [Figure 1] It is a schematic overview according to an embodiment. [Figure 2] It is a schematic overview according to an embodiment. [Figure 3] It is a schematic block diagram of a decoder according to an embodiment. [Figure 4] It is a schematic block diagram of an encoder according to an embodiment. [Figure 5] It is a schematic block diagram according to an embodiment. [Figure 6] It is a schematic block diagram according to an embodiment. [Figure 7] It is a schematic block diagram according to an embodiment. [Figure 8] It is a schematic block diagram according to an embodiment. [Figure 9] It is a schematic block diagram according to an embodiment. [Figure 10] It is a schematic diagram according to an embodiment. [Figure 11] It is a simplified block diagram according to an embodiment. [Figure 12] It is a simplified block diagram according to an embodiment. [Figure 13] It is a simplified block diagram and a timing diagram according to an embodiment. [Figure 14] It is a simplified block diagram according to an embodiment. [Figure 15] It is a simplified block diagram according to an embodiment. [Figure 16] It is a simplified flowchart according to an embodiment. [Figure 17] It is a simplified flowchart according to an embodiment. [Figure 18] It is a schematic diagram according to an embodiment.
MODE FOR CARRYING OUT THE INVENTION
[0018] The proposed features described below may be used separately or combined in any order. Further, the embodiments may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.
[0019] FIG. 1 shows a simplified block diagram of a communication system 100 according to an embodiment of the present disclosure. The communication system 100 may include at least two terminals 102 and 103 interconnected via a network 105. For unidirectional data transmission, the first terminal 103 may encode video data at a local location for transmission to the other terminal 102 via the network 105. The second terminal 102 may receive the encoded video data of the other terminal from the network 105, decode the encoded data, and display the restored video data. Unidirectional data transmission may be common in media serving applications and the like.
[0020] Figure 1 shows a second pair of terminals 101 and 104 provided to support the bidirectional transmission of coded video, for example, during a video conference. For bidirectional data transmission, each terminal 101 and 104 may code video data captured at its local location for transmission to the other terminal via the network 105. Each terminal 101 and 104 may also receive coded video data transmitted by the other terminal, decode the coded data, and display the restored video data on a local display device.
[0021] In Figure 1, terminals 101, 102, 103, and 104 may be represented as a server, a personal computer, and a smartphone, but the principles of this disclosure are not limited thereto. Embodiments of this disclosure find applications involving laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network 105 represents any number of networks that transmit coded video data between terminals 101, 102, 103, and 104, including, for example, wired and / or wireless communication networks. Communication network 105 may exchange data over circuit-switched and / or packet-switched channels. Typical networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this description, the architecture and topology of network 105 may not be important to the operation of this disclosure unless described below herein.
[0022] Figure 2 shows an example of the application of the disclosed subject matter, illustrating the arrangement of a video encoder and video decoder in a streaming environment. The disclosed subject matter can be equally applied to other video-enabled applications, including, for example, video conferencing, digital TV, and the storage of compressed video on digital media such as CDs, DVDs, and memory sticks.
[0023] The streaming system may include a capture subsystem 203 which can include a video source 201, such as a digital camera, which can create an uncompressed video sample stream 213. The sample stream 213 may be highlighted as having a higher data volume compared to the encoded video bitstream and can be processed by an encoder 202 coupled to the camera 201. The encoder 202 may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject, as will be described in more detail below. The encoded video bitstream 204 may be highlighted as having a lower data volume compared to the sample stream and can be stored in the streaming server 205 for future use. One or more streaming clients 212 and 207 can access the streaming server 205 and retrieve copies 208 and 206 of the encoded video bitstream 204. Client 212 may include a video decoder 211 which decodes the incoming copy 208 of the encoded video bitstream and creates an output video sample stream 210 that can be rendered on a display 209 or other rendering device (not shown). In some streaming systems, video bitstreams 204, 206, and 208 can be encoded according to specific video coding / compression standards. Examples of these standards are mentioned above and are further described herein.
[0024] Figure 3 may be a functional block diagram of a video decoder 300 according to one embodiment of the present invention.
[0025] Receiver 302 may receive one or more codec video sequences to be decoded by decoder 300, which in the same or other embodiments may be one coded video sequence at a time, and the decoding of each coded video sequence is independent of other coded video sequences. Coded video sequences may be received from channel 301, which may be a hardware / software link to a storage device that stores encoded video data. Receiver 302 may receive encoded video data together with other data that may be transferred to their respective user entities (not shown), e.g., coded audio data and / or auxiliary data streams. Receiver 302 may isolate coded video sequences from other data. To combat network jitter, a buffer memory 303 may be coupled between receiver 302 and entropy decoder / parser 304 (hereinafter, "Parser"). Buffer 303 may not be necessary or may be smaller if receiver 302 is receiving data from a storage / transfer device with sufficient bandwidth and controllability, or from an isosynchronous network. For use on best-effort packet networks such as the internet, buffer 303 may be required and can be relatively large, or advantageously, its size can be adapted.
[0026] The video decoder 300 may include a parser 304 for reconstructing symbols 313 from an entropy-coded video sequence. These symbol categories may include information used to manage the operation of the decoder 300, and potentially information for controlling rendering devices such as a display 312, which are not integral parts of the decoder but can be coupled to it. The rendering device control information may be in the form of supplemental extension information (SEI messages) or video usability information parameter set fragments (not shown). The parser 304 may parse / entropy-decode the received coded video sequence. The coding of the coded video sequence may conform to video coding techniques or standards and may follow principles well known to those skilled in the art, including variable-length coding, Huffman coding, context-dependent or non-context-dependent arithmetic coding, etc. From the coded video sequence, the parser 304 may extract from the coded video sequence a set of at least one subgroup parameters of subgroups of pixels in the video decoder, based on at least one parameter corresponding to a group. Subgroups can include groups of pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), and predictive units (PU). The entropy decoder / parser can also extract information from the coded video sequence, such as transform coefficients, quantizer parameter values, and motion vectors.
[0027] The parser 304 may perform entropy decoding / analysis operations on the video sequence received from the buffer 303 to create symbols 313. The parser 304 may receive encoded data and selectively decode specific symbols 313. Furthermore, the parser 304 may determine whether a particular symbol 313 should be provided to the motion compensation prediction unit 306, the scaler / inverse transform unit 305, the intra prediction unit 307, or the loop filter 311.
[0028] The reconstruction of symbol 313 may involve multiple different units, depending on the type of coded video picture or its components (interpicture and intrapicture, interblock and intrablock, etc.), as well as other factors. Which units are involved and how can be controlled by parser 304 using subgroup control information parsed from the coded video sequence. The flow of such subgroup control information between parser 304 and the following multiple units is not illustrated for clarity.
[0029] In addition to the functional blocks already mentioned, the decoder 300 can be conceptually subdivided into several functional units, as described below. In actual embodiments operating under commercial constraints, many of these units can interact closely with each other and be integrated with each other at least partially. However, the following conceptual subdivision into functional units is appropriate for illustrating the disclosed subject matter.
[0030] The first unit is the scaler / inverse unit 305. The scaler / inverse unit 305 receives control information from the parser 304 as symbol 313, including the quantized transformation coefficients and the transformation to be used, block size, quantization factor, and quantization scaling matrix. The scaler / inverse unit 305 can output a block containing sample values that can be input to the aggregator 310.
[0031] In some cases, the output samples of the scaler / inverse transform 305 may also relate to intracoded blocks, i.e., blocks that do not use predictive information from a previously reconstructed picture, but can use predictive information from a previously reconstructed portion of the current picture. Such predictive information can be provided by the intrapicture predictive unit 307. In some cases, the intrapicture predictive unit 307 generates a block of the same size and shape as the block being reconstructed, using the surrounding already reconstructed information fetched from the current (partially reconstructed) picture 309. The aggregator 310, in some cases, adds the predictive information generated by the intrapredictive unit 307 to the output sample information provided by the scaler / inverse transform unit 305, sample by sample.
[0032] In other cases, the output samples of the scaler / inverse unit 305 may relate to an intercoded and potentially motion-compensated block. In such cases, the motion-compensated prediction unit 306 may access the reference picture memory 308 and fetch samples to be used for prediction. After motion-compensating the fetched samples according to the symbols 313 related to the block, these samples may be added to the output of the scaler / inverse unit by the aggregator 310 to generate output sample information (in this case, called residual samples or residual signals). The address in the reference picture memory format from which the motion-compensated unit fetches the prediction samples may be controlled by a motion vector and made available to the motion-compensated unit in the form of a symbol 313 which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values fetched from the reference picture memory when the exact motion vector of a subsample is in use, a motion vector prediction mechanism, etc.
[0033] The output samples of the aggregator 310 can be subjected to various loop filtering techniques in the loop filter unit 311. The video compression technique may include in-loop filtering techniques controlled by parameters contained in the coded video bitstream and made available to the loop filter unit 311 as symbols 313 from the parser 304, but it may also respond to metadata obtained during decoding of earlier parts (in decoding order) of the coded picture or coded video sequence, and may also respond to previously reconstructed and loop-filtered sample values.
[0034] The output of the loop filter unit 311 may be a sample stream that can be output to a display 312, which may be a rendering device, or it may be stored in a reference picture memory 557 for use in future interpicture prediction.
[0035] A particular coded picture, once fully reconstructed, can be used as a reference picture for future predictions. Once a coded picture is fully reconstructed and identified as a reference picture (for example, by parser 304), the current reference picture 309 can become part of the reference picture buffer 308, allowing for the reallocation of new current picture memory before initiating the reconstruction of subsequent coded pictures.
[0036] The video decoder 300 may perform decoding operations according to a predetermined video compression technique documented in a standard such as ITU-T Rec.H.265. The coded video sequence may conform to the syntax specified by the video compression technique or standard used, in the sense that it is faithful to the syntax of the video compression technique or standard, as specified in the video compression technique documentation or standard, specifically in the profile documentation therein. The complexity of the coded video sequence may also be required for compliance to be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limitations set by the level may, in some cases, be further limited by the virtual reference decoder (HRD) specification and metadata for HRD buffer management signaled within the coded video sequence.
[0037] In one embodiment, receiver 302 may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by video decoder 300 to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may take the form of, for example, a time layer, a spatial layer, or a signal-to-noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, forward error correction codes, etc.
[0038] Figure 4 may be a functional block diagram of a video encoder 400 according to one embodiment of the present disclosure.
[0039] The encoder 400 may receive video samples from a video source 401 (not part of the encoder) that can capture video images to be coded by the encoder 400.
[0040] The video source 401 may provide a source video sequence coded by an encoder (303) in the form of a digital video sample stream, which can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and a suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source 401 may be a storage device that stores previously prepared video. In a video conferencing system, the video source 401 may be a camera that captures local image information as a video sequence. The video data may be provided as a series of individual pictures that convey motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each pixel may contain one or more samples depending on the sampling structure, color space, etc., in use. Those skilled in the art will readily understand the relationship between pixels and samples. The following description focuses on samples.
[0041] According to one embodiment, the encoder 400 can encode pictures of a source video sequence in real time or under any other time constraints as required by the application and compress them into a coded video sequence 410. One function of the controller 402 is to ensure an appropriate coding speed. The controller controls and is functionally coupled to other functional units, as described below. The coupling is not illustrated for clarity. Parameters set by the controller may include rate control-related parameters (such as picture skip, quantizer, lambda value of rate distortion optimization technique), picture size, picture group (GOP) layout, maximum motion vector search range, etc. Those skilled in the art will be able to easily identify the other functions of the controller 402, as they may relate to the video encoder 400 optimized for a particular system design.
[0042] Some video encoders operate in what is readily recognizable to those skilled in the art as a “coding loop.” In an overly simplified explanation, the coding loop may consist of an encoding portion of the encoder (e.g., source coder 403) (which is responsible for creating symbols based on the input picture and reference picture to be coded) and a (local) decoder 406 built into encoder 400 that reconstructs the symbols and creates sample data, which (since any compression between the symbols and the coded video bitstream is reversible in the video compression techniques considered in the disclosed subject) would also create for the (remote) decoder. The reconstructed sample stream is input to the reference picture memory 405. Since decoding the symbol stream results in a bit-exact outcome regardless of the decoder's location (local or remote), the contents of the reference picture buffer are also bit-exact between the local and remote encoders. In other words, the predictive portion of the encoder “sees” the exact same sample values as the reference picture samples that the decoder “sees” when using predictions during decoding. This basic principle of the synchronization of a reference picture (and the resulting drift when synchronization cannot be maintained, for example, due to channel errors) is well known to those skilled in the art.
[0043] The operation of the “local” decoder 406 may be the same as that of the “remote” decoder 300, which has already been described in detail above in relation to Figure 3. However, also briefly referring to Figure 4, since symbols are available and the encoding / decoding of symbols to the coded video sequence by the entropy coder 408 and parser 304 can be reversible, the entropy decoding portion of decoder 300, including channel 301, receiver 302, buffer 303, and parser 304, may not be fully performed by the local decoder 406.
[0044] At this point, it can be said that any decoder techniques within the decoder, excluding analysis / entropy decoding, must also necessarily exist in the corresponding encoder in substantially the same functional form. The description of encoder techniques can be omitted as it is the inverse of the comprehensively described decoder techniques. More detailed explanations are necessary only in specific areas, and are provided below.
[0045] As part of its operation, the source coder 403 may perform motion-compensated predictive coding, which predictively codes the input frame by referencing one or more previously coded frames from the video sequence, designated as “reference frames”. In this way, the coding engine 407 codes the difference between the pixel blocks of the input frame and the pixel blocks of the reference frame, which may be selected as a predictive reference to the input frame.
[0046] The local video decoder 406 can decode coded video data of a frame that may be designated as a reference frame, based on symbols created by the source coder 403. The operation of the coding engine 407 may, advantageously, be an irreversible process. If coded video data can be decoded by a video decoder (not shown in Figure 4), the reconstructed video sequence may typically be a copy of the source video sequence with some error. The local video decoder 406 can reproduce the decoding process that may be performed by the video decoder on the reference frame and store the reconstructed reference frame in a reference picture memory 405, which may be a cache, for example. In this way, the encoder 400 can locally store a copy of the reconstructed reference frame having common content as the reconstructed reference frame that will be acquired by a far-end video decoder (without transmission errors).
[0047] The predictor 404 may perform a predictive search for the coding engine 407. That is, for a new frame to be coded, the predictor 404 may search the reference picture memory 405 for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc., which can function as appropriate predictive references for the new picture. The predictor 404 may operate on sample blocks pixel by pixel to find appropriate predictive references. In some cases, the input picture may have predictive references drawn from multiple reference pictures stored in the reference picture memory 405, as determined by the search results obtained by the predictor 404.
[0048] The controller 402 may manage the coding operations of the video coder 403, including, for example, setting parameters and subgroup parameters used to encode video data.
[0049] The outputs of all the aforementioned functional units may be entropy coded by the entropy coder 408. The entropy coder converts the symbols generated by the various functional units into coded video sequences by losslessly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable-length coding, and arithmetic coding.
[0050] Transmitter 409 may buffer the coded video sequence created by entropy coder 408 for transmission over communication channel 411, which may be a hardware / software link to a storage device that will store the encoded video data. Transmitter 409 may merge the coded video data from video coder 403 with other data being transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).
[0051] Controller 402 may manage the operation of encoder 400. During coding, controller 405 may assign a specific coded picture type to each coded picture, which may affect the coding technique that can be applied to each picture. For example, a picture may often be assigned as one of the following frame types:
[0052] An intra-picture (I-picture) can be a picture that can be coded and decoded without using any other frame in the sequence as a source of prediction. Some video codecs enable various types of intra-pictures, including, for example, independent decoder refresh pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.
[0053] A predictive picture (P-picture) can be a picture that can be coded and decoded using intra-prediction or inter-prediction, which uses up to one motion vector and reference index to predict the sample values for each block.
[0054] A bidirectional predictive picture (B-picture) can be a picture that can be coded and decoded using intra-prediction or inter-prediction, which uses up to two motion vectors and reference indices to predict the sample values for each block. Similarly, a multiple predictive picture can use three or more reference pictures and associated metadata for the reconstruction of a single block.
[0055] A source picture can generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each), and each block can be coded. Blocks can be coded predictively by referencing other (already coded) blocks, as determined by the coding assignment applied to each picture in the block. For example, blocks in picture I may be coded unpredictably or predictively by referencing already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks in picture P may be coded unpredictably via spatial prediction or via temporal prediction referencing one previously coded reference picture. Pixel blocks in picture B may be coded unpredictably via spatial prediction or via temporal prediction referencing one or two previously coded reference pictures.
[0056] The videocoder 400 may perform coding operations in accordance with a specified video coding technique or standard, such as ITU-T Rec.H.265. In these operations, the videocoder 400 may perform various compression operations, including predictive coding operations that leverage the temporal and spatial redundancy of the input video sequence. Therefore, the coded video data may conform to the syntax specified by the video coding technique or standard being used.
[0057] In one embodiment, the transmitter 409 may transmit additional data along with the encoded video. The source coder 403 may include such data as part of the encoded video sequence. The additional data may include time / space / SNR enhancement layers, other forms of redundant data such as redundant pictures and redundant slices, supplemental enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, and the like.
[0058] Figure 5 is an example of an end-to-end architecture for a standalone AR (STAR) device according to an exemplary embodiment, showing a 5G STAR user equipment (UE) receiver 600, a network / cloud 501, and a 5G UE (transmitter) 700. Figure 6 is a further detailed example of one or more configurations for the STAR UE receiver 600 according to an exemplary embodiment, and Figure 7 is a further detailed example of one or more configurations for the 5G UE transmitter 700 according to an exemplary embodiment. 3GPP TR 26.998 defines support for glasses-type augmented reality / mixed reality (AR / MR) devices in 5G networks. Furthermore, according to the exemplary embodiments herein, at least two device classes are conceivable: 1) a device capable of fully decoding and playing complex AR / MR content (standalone AR, i.e., STAR), and 2) a device having less computing resources and / or a smaller physical size (and therefore battery) and capable of running such applications only when the majority of the computing is performed on a 5G edge server, network, or cloud rather than on the device (edge-dependent AR, i.e., EDGAR).
[0059] Furthermore, according to an exemplary embodiment, a shared conversation use case can be experienced, as described below, in which all participants in the shared AR conversation experience have an AR device, each participant sees other participants in the AR scene, the participants are overlays in the local physical scene, and the placement of participants in the scene matches on all receiving devices, for example, people in each local space have the same position / seating arrangement relative to each other, creating a sense of being in the same space, however the room is different for each participant, as the room is an actual room or the space in which each person is physically located.
[0060] For example, according to the exemplary embodiment shown in Figures 5-7, an immersive media processing function on the network / cloud 501 receives uplink streams from various devices and constructs a scene description that defines the placement of individual participants in a single virtual meeting room. The scene description and encoded media streams are delivered to each receiving participant. The receiving participant's 5G STAR UE600 receives, decodes, and processes 3D video and audio streams, rendering them using the received scene description and information received from its AR runtime to create an AR scene of the virtual meeting room with all other participants. The virtual room for a participant is based on its own physical space, but the seating / positioning of all other participants in the room matches the virtual room of all other participants in this session.
[0061] As shown in the exemplary embodiment, see also Figure 8, which illustrates an example of an EDGAR device architecture, where the device itself, such as the 5G EDGAR UE900, cannot perform a large amount of processing. Therefore, scene analysis and media analysis of the received content are performed in the cloud / edge 801, and then a simplified AR scene with a small number of media components is delivered to the device for processing and rendering. Figure 9 also shows an AR / MR (mixed reality) application provider 802. Figure 9 shows a more detailed example of the 5G EDGAR UE900 according to the exemplary embodiment.
[0062] However, even with capabilities such as those described in the exemplary embodiments of Figures 5 to 9, one or more technical problems may exist regarding constructing a virtual space scene description common to immersive media functionality. Furthermore, as will be discussed later, such embodiments are technically improved in the context of immersive media processing functionality to generate a scene description provided to all participants so that all participants can experience the same relative placement of participants within the local AR scene.
[0063] Figure 10 shows an example 1000 in which users A10, B11, and T12 are to participate in an AR conference room, and one or more of the users may not have an R device. As shown, user A10 is in his office 1001 and is seated in a conference room with varying numbers of chairs, and user A10 is seated in a chair. User B11 is in his living room 1002 and is seated on a two-seater sofa, and there is also one or more two-seater couches in his living room, as well as other furniture such as chairs and tables. User T12 is on a bench in the airport lobby 1003, but the bench is across one of one or more other coffee tables.
[0064] Furthermore, refer to the AR environment in office 1001 where user A10's AR shows user A10 a virtual user B11v1 corresponding to user B11 and a virtual user T12v1 corresponding to user T12, and user A10 is shown that virtual users B11v1 and T12v1 are sitting in the same furniture and office chairs in office 1001 as user A10. Also, refer to living room 1202 in example 1200 where the AR for user B11 shows a virtual user T12v2 corresponding to user T12, but sitting on a couch in living room 1202, and a virtual user A10v1 corresponding to user A10, but sitting on furniture in living room 1202 instead of an office chair in office 1201. See also airport lobby 1203, where AR for user T12 shows virtual user A10v2, corresponding to user A10 but seated at a table in airport lobby 1203, and virtual user B11v2, also seated at a table opposite virtual user A10v2. And in each of those offices 1201, living room 1202, and airport lobby 1203, the updated scene description of each room is consistent with the other rooms in terms of position / seating arrangement. For example, user A10 is shown as their virtual representation which is counterclockwise relative to user 11, or clockwise relative to user T12, or as their virtual representation in each room.
[0065] However, AR technology is limited in any attempt to incorporate the creation and use of virtual spaces for devices that do not support AR but can analyze VR or 2D video, and embodiments herein provide improved technical procedures for creating virtual scenes that match AR scenes when such devices participate in a shared AR conversation service.
[0066] Figure 11 shows an example of an end-to-end architecture 1100 having a non-AR device 1101 and a cloud / edge 1102 according to an exemplary embodiment. Figure 12 also shows an example of a more detailed block diagram of the non-AR device 1101.
[0067] As shown in Figures 11 and 12, the non-AR UE 1101 is a device that can render 360 or 2D video but does not have AR capabilities. However, the edge functionality on the cloud / edge 1102 makes it possible to AR render received scenes, rendering scenes, and immersive visual and audio objects in a virtual room selected from the library. The entire video is then encoded and delivered to device 1101 for decoding and rendering.
[0068] Thus, the AR processing on the edge / cloud 1102 may have multi-view capabilities, such as being able to generate multiple videos of the same virtual room from different angles and using different viewports. Device 1101 can then receive one or more of these videos and switch between them as needed, or send commands to the edge / cloud processing to stream only the desired viewport / angle.
[0069] Furthermore, the background capabilities can be changed, and the user on device 1101 can select a background for a desired room from a provided library, for example, one of several different meeting rooms, or a living room and layout. The cloud / edge 1102 then uses the selected background and creates a virtual room accordingly.
[0070] Figure 13 shows an exemplary timing diagram 1300 for an exemplary call flow for an immersive AR conversation for a receiving non-AR UE1101. For illustrative purposes only, only one sender is shown in this diagram without showing a detailed call flow.
[0071] The AR application module 21, media playback module 22, and media access function module 23, which can be considered modules of the receiving non-AR UE1101, are shown. The cloud / edge split rendering module 24 is also shown. The media distribution module 25 and scene graph composer module 26 of the network cloud 1102 are also shown. The 5G transmitter UE module 700 is also shown.
[0072] S1 to S6 can be considered as session establishment phases. In S1, the AR application module 21 can request the media access function module 23 to start a session, and in S2, the media access function module 23 can request the cloud / edge split rendering module 24 to start a session.
[0073] The cloud / edge segmented rendering module 24 can perform session negotiation with the scene graph composer module 26, which can negotiate with the 5G transmitter UE700 accordingly, in S3. If successful, in S5, the cloud / edge segmented rendering module can send an acknowledgment to the media access function module 23, and the media access function module 23 can send an acknowledgment to the AR application module 21.
[0074] Subsequently, S7 can be considered a media pipeline configuration stage in which the media access function module 23 and the cloud / edge split rendering module 24 each constitute a pipeline. After this pipeline configuration, sessions may be initiated by signals from the AR application module to the media player module 22 in S8, from the media player module 22 to the media access function module 23 in S9, and from the media access function module 23 to the cloud / edge split rendering module 24 in S10.
[0075] Next, there can be a pause loop stage from S11 to S13. In S11, the pause data can be provided from the media player module 22 to the AR application module 21. In S12, the AR application module can provide the pause data 12 to the media access function module 23. Subsequently, the media access function module 23 can provide the pause data to the cloud / edge split rendering module 24.
[0076] S14-S16 can be considered a shared experience stream stage in which, in S14, the 5G transmitter UE700 provides a media stream to the media distribution module 25, and in S15, provides AR data to the scene graph synthesis module 26. Next, the scene graph synthesis module 25 can construct one or more scenes based on the received AR data and provide the scenes and scene updates to the cloud / edge split rendering module 24 in S16, and the media distribution module 25 can provide the media stream to the cloud / edge split rendering module in S17. This may include, according to an exemplary embodiment, the steps of obtaining an AR scene descriptor from a non-AR device that does not render an AR scene, and generating a virtual scene by a cloud device by parsing and rendering the scene descriptor obtained from the non-AR device.
[0077] S18-S19 can be considered the media uplink stage in which the media player module 22 captures and processes media data from its local user, and in S18 provides that media data to the media access function module 23. Next, the media access module 23 encodes the media and in S19 can provide the media stream to the cloud / edge split rendering module 24.
[0078] The period between S19 and S20 can be considered a media downlink phase in which the cloud / edge split rendering module 24 can perform scene analysis and complete AR rendering, and then S20 and S21 can be considered to constitute a media stream loop phase. In S20, the cloud / edge split rendering module 24 can provide the media stream to the media access function module 23, which then decodes the media and, in S21, can provide the media rendering to the media player 22.
[0079] Such features, as demonstrated by exemplary embodiments, allow a non-AR UE1101 to utilize its display to render VR or 2-D video even if it lacks a see-through display and therefore cannot create AR scenes. Thus, its immersive media processing capabilities generate only a common scene description that describes each participant's relative position to other participants and the scene. The scene itself needs to be adjusted with each device's pose information before being rendered as an AR scene, as described above. Furthermore, an AR rendering process on the edge or in the cloud can analyze the AR scene and create a simplified VR-2D scene.
[0080] According to exemplary embodiments, the disclosure uses a similar segmented rendering process for EDGAR devices for non-AR devices such as VR or 2-D video devices, where characteristics such as edge / cloud AR rendering processes do not generate an AR scene. Instead, it generates a virtual scene by parsing and rendering a scene description received from an immersive media processing function for a given background (such as a conference room), and then rendering each participant in the locations described by the conference room scene description.
[0081] Furthermore, the resulting video can be 360-degree video or 2-D video depending on the capabilities of the receiving non-AR device, and the resulting video is generated taking into account the location information received from the non-AR device according to the exemplary embodiment.
[0082] Additionally, each other participant with a non-AR device can be added as a 2D video overlay to the 360 / 2D video of the meeting room, as shown in Figure 10, and the room may have areas dedicated to the use of these overlays, such as furniture on which virtual images are overlaid, as shown in Figure 10.
[0083] Furthermore, audio signals from all participants can be mixed as needed to create a single-channel audio that carries the sound within the room, video can be encoded as a single 360 video or 2-D video and delivered to the device, and optionally, multiple video (multiview) sources can be created, each capturing the same virtual conference room from a different view and providing those views to the device according to the exemplary embodiment.
[0084] Furthermore, the non-AR UE device 1101 can receive 360 video and / or selected one or more multiview videos along with audio and render them on the device display, allowing the user to change the viewport of the 360 video by switching between different views or moving or rotating the view device, thus enabling navigation within the virtual room while viewing the video.
[0085] The embodiments described above feature such 5G media stream architecture (5GMS) extensions for using edge servers in those architectures, and while their specifications may have many features, such features are technically impossible to deploy as a set of software development kits (SDKs) on devices or as a set of microservices in the cloud, and such technical shortcomings are addressed by embodiments further described below.
[0086] For example, the current Media Services Enabler technical report does not define a framework for associating specifications with SDKs and does not include the concept of microservices.
[0087] Refer to Example 1400 in Figure 14, which shows a 5G media streaming architecture with edge extensions according to an exemplary embodiment. As shown, there is a user device (UE) 1401 and a data network (DN) 1411. The UE 1401 may include, but is not limited to, a 5GMS client 1403 and a 5GMS-enabled application 1405, such as the AR or non-AR embodiments described above. The 5GMS client 1403 may also include a media stream handler 1403 and a media session handler. The DN 1411 may include a 5GMS application server (AS) 1412, a 5GMS application function (AF) 1414, and a 5GMS application provider 1413. The 5GMS AF 1414 may also communicate with a network exposure function (NEF) 1415 and a policy and billing function (PCF) 1416. The improvements described herein can be understood in the context of at least one or more of UE1401, 5GMS-enabled applications1405, 5GMS application providers1413, NEF1415, and PCF1416, that is, rather than having a monolithic specification in which there is no definition of a Media Services Enabler (MSE) for each function or function group, one or more of those elements can generate their own specification and provide such a specification to another element of those elements, which can further configure its initial element specification in accordance with various possibilities as described below, depending on the requirements to be adapted, and such processing can be done via one or more of the multiple Exposure Application Programming Interfaces (APIs)1420 and interfaces M1, M2, M4, M5, M6, M7, M8, N33, and N5 shown.
[0088] Figure 15 shows the new capabilities and performance added to the EDGAR client. detection (CPD) via module 1501 detectionTo obtain this, examples of technical deficiencies caused by the 5G server itself, as well as applications running via the 5G server, are shown. This CPD module 1501 serves to provide a list of features, parameters supported by the features, their performance, and the overall power status of the device (such as battery level). The module can collect statistics on each available function, such as various encoders and decoders, measure resource consumption, i.e., CPU / GPU cycles, memory bandwidth, memory, and battery consumption, and provide detailed characteristics of the functions, such as maximum video encoder profile and level, maximum supported width and height, and other relevant information. In addition, the CPD module 1501 of the EDGAR device architecture enables queries made by any authorized application, such as applications on the device, applications on the edge network, or application service providers over the wireless network.
[0089] Figure 16 shows an example of a CPD module 1501 in an exemplary embodiment, which has a local application programming interface (API) for native applications on the device to access the device, and for providing information to applications on the edge server or to the application service provider itself, rather than the application, via the 5G system interface. For example, the features of a 5G EDGAR device detection The steps to control the 5G EDGAR device include controlling the 5G EDGAR device via a network API and the characteristics of the 5G EDGAR device. index The step to obtain it is via the network API index This includes the step of obtaining [something].
[0090] According to an exemplary embodiment, the CPD module 1501, via its API, provides any of the following functions, which are non-exhaustive examples of its functionality: index,(f) the EDGAR device is configured to provide: (1) a list of features supported by the device: (a) media encoders and decoders, and their maximum supported profiles and levels (image, video, audio); (b) device capture capabilities (camera, microphone, and other sensors), which may include manufacturer specifications, current status such as whether or not they are in use, and application capabilities; (c) device rendering capabilities (speaker, display, and other sensors), which may include manufacturer specifications, current status such as whether or not they are in use, and application capabilities; (d) device network protocol stacks (various communication protocols), which may include manufacturer specifications, current status, and application capabilities; (e) raw hardware capabilities such as CPU and GPU benchmarks, memory, memory bandwidth, and on-device storage, which may include manufacturer specifications, current status, and application capabilities; (f) the efficiency of each function, i.e., the relative energy consumption of that function based on the parameters of each function, which can be determined by the EDGAR device by monitoring while the function is running parameter by parameter; (g) the power status of the device and the battery drain rate; and (2) a combination profile evaluation, such as whether a particular configuration of the function can be run in real time, and the overall battery consumption of that combination.
[0091] Thus, exemplary embodiments of this specification have a local device interface and a network interface that can thereby provide a list of media functions supported by the device. detection and the functionality supported by 5G devices using performance features detectionThe system provides a method for doing so, and raw hardware capabilities include capture, rendering, networking protocol stack, raw hardware capabilities, efficiency of each function, and ultimately power status and consumption. This function can be requested to evaluate whether a particular combination of functions can be performed in real time, and if so, to report power consumption.
[0092] For example, referring to the illustrative flowchart 1700 in Figure 17, the 5G service in S1701 is from the CPD module 1501 in S1702. detection A request can be made, and in response, the CPD module 1501, via its various APIs, in S1704, according to an exemplary embodiment, via its APIs, a non-exhaustive example of functionality is any of the following: detection and indexable: (1) Providing a list of features supported by the device: (a) media encoders and decoders, as well as their maximum supported profiles and levels (image, video, audio), (b) device capture capabilities (camera, microphone, and other sensors), (c) device rendering capabilities (speaker, display, and other sensors), (d) device network protocol stack (various communication protocols), (e) raw hardware capabilities such as CPU and GPU benchmarks, memory, memory bandwidth, and on-device storage, (f) efficiency of each function, i.e., relative energy consumption for each function based on the parameters of that function (this is not limited to any particular codec or any particular media stream, or detection (g) the power status and battery drain rate of the device, and (2) a combination profile evaluation of whether a specific configuration of the function can be performed in real time, and the overall battery consumption of that combination (which may be the object provided to the device inside).
[0093] Next, a report may be generated in S1705 with respect to the index, or in S1707, as one or more combinations of indexed functionalities requested by the networked server itself, rather than by the networked application or a networked application running via the networked server.
[0094] In addition, in S1702 detection Live reports and index It may be provided continuously as such. And in S1706, detection The information provided may be provided to applications running via a network server or networking server, either of which can then decide whether to offload 5G services or any part thereof to or between one of the further 5G networking devices described above in AR and non-AR contexts, but the embodiments herein are not limited in that respect.
[0095] The techniques described above use computer-readable instructions and are implemented as computer software physically stored on one or more computer-readable media, or by one or more hardware processors specifically configured to do so. execution It is possible. For example, Figure 18 shows a particular embodiment of the disclosed subject matter. execution This shows a suitable computer system 1800 for doing so.
[0096] Computer software can be coded using any suitable machine code or computer language that can undergo mechanisms such as assembly, compilation, and linking to create code that includes instructions that can be executed directly by a computer's central processing unit (CPU), graphics processing unit (GPU), etc., or through interpretation, microcode execution, etc.
[0097] Instructions can be executed on various types of computers or computer components, including, for example, personal computers, tablet computers, servers, smartphones, game consoles, and Internet of Things devices.
[0098] The components shown in Figure 18 for the computer system 1800 are essentially illustrative, and embodiments of this disclosure are not limited to those described herein. execution This is not intended to imply any limitations on the scope of use or functionality of the computer software. The configuration of the components should not be construed as having any dependencies or requirements on any one or combination of components shown in the exemplary embodiments of computer system 1800.
[0099] The computer system 1800 may include certain human interface input devices. Such human interface input devices may respond to input from one or more human users via, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, applause), visual input (e.g., gestures), or olfactory input (not shown). The human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., voices, music, ambient sounds), images (e.g., scanned images, photographic images, still images, etc., acquired from a camera), or video (e.g., two-dimensional video, three-dimensional video including stereoscopic video, etc.).
[0100] The input human interface device may include one or more of the following (only one of each is illustrated): keyboard 1801, mouse 1802, trackpad 1803, touchscreen 1810, joystick 1805, microphone 1806, scanner 1808, and camera 1807.
[0101] The computer system 1800 may also include certain human interface output devices. Such human interface output devices may stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (e.g., tactile feedback via a touchscreen 1810 or joystick 1805, although tactile feedback devices that do not function as input devices may also exist), audio output devices (e.g., speakers 1809, headphones (not shown)), visual output devices (e.g., screens 1810 including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touchscreen input capabilities, each with or without tactile feedback capabilities, some of which may be capable of outputting two-dimensional visual output or output beyond three dimensions via means such as stereoscopic image output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).
[0102] The computer system 1800 may also include human-accessible storage devices and associated media, such as optical media including a CD / DVD ROM / RW 1820 with media like a CD / DVD 1811, a thumb drive 1822, a removable hard drive or solid-state drive 1823, legacy magnetic media such as tapes or floppy disks (not shown), and dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown).
[0103] Those skilled in the art should also understand that the term “computer-readable medium” as used in relation to the subject matter of this disclosure does not include transmission media, carrier waves, or other transient signals.
[0104] The computer system 1800 may also include an interface 1899 to one or more communication networks 1898. The network 1898 may be, for example, wireless, wired, or optical. The network 1898 may further be local, wide-area, metropolitan, automotive, and industrial, real-time, latency-tolerant, etc. Examples of the network 1898 include local area networks such as Ethernet and wireless LANs, cellular networks including GSM, 3G, 4G, 5G, and LTE, wide-area digital networks for wired or wireless TV including cable TV, satellite TV, and terrestrial broadcast TV, and automotive and industrial networks including CANBus. Certain networks 1898 generally require a specific general-purpose data port or peripheral bus (1850 and 1851) (e.g., an external network interface adapter attached to a USB port on the computer system 1800), while other networks are generally integrated into the core of the computer system 1800 by connection to a system bus, as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks 1898, the computer system 1800 can communicate with other entities. Such communication may be unidirectional reception only (e.g., broadcast television), unidirectional transmission only (e.g., CANbus to a specific CANbus device), or bidirectional communication to other computer systems using local area or wide area digital networks. Specific protocols and protocol stacks may be used on each of those networks and network interfaces, as described above.
[0105] The aforementioned human interface devices, human-accessible memory devices, and network interfaces can be mounted on the core 1840 of the computer system 1800.
[0106] The core 1840 may include one or more central processing units (CPUs) 1841, graphics processing units (GPUs) 1842, graphics adapters 1817, dedicated programmable processing units in the form of field-programmable gate areas (FPGAs) 1843, and hardware accelerators 1844 for specific tasks. These devices may be connected via a system bus 1848, along with read-only memory (ROM) 1845, random access memory 1846, and internal mass storage 1847 such as internal hard drives and SSDs that are not accessible to the user. In some computer systems, the system bus 1848 may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be connected directly to the core's system bus 1848 or via a peripheral bus 1851. Architectures for peripheral buses include PCI, USB, etc.
[0107] The CPU 1841, GPU 1842, FPGA 1843, and accelerator 1844 can execute certain instructions that, when combined, can constitute the aforementioned computer code. This computer code can be stored in ROM 1845 or RAM 1846. Temporary data can also be stored in RAM 1846, while persistent data can be stored, for example, in internal mass storage 1847. High-speed storage and retrieval to any of the memory devices can be enabled by the use of cache memory, which can be closely associated with one or more CPUs 1841, GPUs 1842, mass storage 1847, ROM 1845, RAM 1846, etc.
[0108] Computer-readable media are various computers execution It may have computer code for performing the operation. The medium and computer code may be specifically designed and constructed for the purposes of this disclosure, or may be of a type that is well known and available to those skilled in the computer software technology.
[0109] For example, but not limited to, an architecture corresponding to computer system 1800, specifically core 1840, can provide functionality as a result of a processor (including CPU, GPU, FPGA, accelerator, etc.) that runs software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage as described above, as well as media associated with specific storage of core 1840 that is non-transient in nature, such as core internal mass storage 1847 or ROM 1845. Software implementing various embodiments of this disclosure can be stored in such devices and executed by core 1840. The computer-readable media can include one or more memory devices or chips, depending on the specific needs. The software can cause core 1840, specifically the processor (including CPU, GPU, FPGA, etc.) therein, to execute specific processes or specific parts of specific processes as described herein, including defining data structures stored in RAM 1846 and modifying such data structures according to processes defined by the software. In addition, or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in a circuit (e.g., accelerator 1844), which may operate in place of or in conjunction with software to perform a particular process or a particular part of a particular process as described herein. References to software may, as necessary, include logic, and vice versa. References to computer-readable media may, as necessary, include circuits that store software for execution (such as integrated circuits (ICs)), circuits that embody logic for execution, or both. This disclosure encompasses any suitable combination of hardware and software.
[0110] While this disclosure describes several exemplary embodiments, there are many variations, substitutions, and alternative equivalents that fall within the scope of this disclosure. Those skilled in the art will therefore understand that numerous systems and methods not expressly shown or described herein can be devised to embody the principles of this disclosure and thus fall within the spirit and scope of this disclosure. [Explanation of Symbols]
[0111] 10 User A, 10v1 Virtual User A, 10v2 Virtual User A, 11 User B, 11v1 Virtual User B, 11v2 Virtual User B, 12 User T, 12v1 Virtual User T, 12v2 Virtual User T, 12 Pose Data, 21 AR Application Module, 22 Media Playback Module, Media Player Module, 23 Media Access Function Module, 24 Cloud / Edge Split Rendering Module, 25 Media Distribution Module, 26 Scene Graph Composer Module, Scene Graph Synthesis Module, 100 Communication System, 101 Terminal, 102 Terminal, Second Terminal, 103 Terminal, First Terminal, 104 Terminal, 105 Communication Network, 201 Video Source, Camera, 202 Encoder, 203 Capture Subsystem, 204 Encoded Video Bitstream, Video Bitstream, 205 Streaming Server, 206 Copy of Encoded Video Bitstream, Video Bitstream, 207 Streaming client, 208 copy of encoded video bitstream, video bitstream, 209 display, 210 video sample stream, 211 video decoder, 212 streaming client, 213 uncompressed video sample stream, 300 video decoder, 301 channel, 302 receiver, 303 buffer memory, encoder, 304 entropy decoder, parser, 305 scaler / inverse unit, 306 motion compensation prediction unit, 307 intrapicture prediction unit, 308 reference picture memory, reference picture buffer, 309 current reference picture, 310 aggregator, 311 loop filter unit, 312 display, 313 symbol, 400 video encoder, video coder, 401 video source, 402 controller, 403 source coder, video coder, 404 predictor, 405 reference picture memory, 406 local video decoder, 407 coding engine, 408 Entropy coder, 409 transmitters, 410 coded video sequences, 411 communication channels, 500 examples, 501 network / cloud, 600STAR User Equipment (UE) Receiver, e.g., 700; 5G Transmitter UE Module, e.g., 800; e.g., 801; Cloud / Edge, 802; AR / MR (Mixed Reality) Application Provider, 900; 5G EDGAR UE, 1000; e.g., 1001; Office, 1002; Living Room, 1003; Airport Lobby, 1100; e.g., 1101; Non-AR Device / Non-AR UE, 1102; Cloud / Edge, 1200; e.g., 1201; Office, 1202; Living Room, 1203; Airport Lobby, 1300; Timing Diagram, 1400; e.g., 1401; User Equipment (UE), 1402; 5GMS Client, 1403; Media Stream Handler, 1405; 5GMS-Enabled Application, 1412; 5GMS Application Server (AS), 1413; 5GMS Application Provider, 1414 5GMS Application Function (AF), 1415 Network Exposure Function (NEF), 1416 Policy and Billing Function (PCF), 1420 Exposure Application Programming Interface (API), 1500 Examples, 1501 Capabilities and Performance detection(CPD) Module, 1700 Flowchart, 1800 Computer System, 1801 Keyboard, 1802 Mouse, 1803 Trackpad, 1805 Joystick, 1806 Microphone, 1807 Camera, 1808 Scanner, 1809 Speaker, 1810 Touchscreen, 1811 CD / DVD, 1817 Graphics Adapter, 1820 CD / DVD ROM / RW, 1822 Thumb Drive, 1823 Removable Hard Drive, Solid State Drive, 1840 Core, 1841 Central Processing Unit (CPU), 1842 Graphics Processing Unit (GPU), 1843 Field Programmable Gate Area (FPGA), 1844 Hardware Accelerator, 1845 Read-Only Memory (ROM), 1846 Random Access Memory, 1847 Internal Mass Storage, 1848 System Bus, 1850 Peripheral Bus, 1851 Peripheral bus, 1898 communication network, 1899 interface, M1 interface, M2 interface, M4 interface, M5 interface, M6 interface, M7 interface, M8 interface, N5 interface, N33 interface, S1 session establishment phase, S2 session establishment phase, S3 session establishment phase, S4 session establishment phase, S5 session establishment phase, S6 session establishment phase, S7 media pipeline configuration phase, S11 pause loop phase, S12 pause loop phase, S13 pause loop phase, S14 shared experience stream phase, S15 shared experience stream phase, S16 shared experience stream phase, S18 media uplink phase, S19 media uplink phase, S20 media stream loop phase, S21 media stream loop phase
Claims
1. A method for media streaming, which is performed by at least one processor, The steps include controlling a 5G edge-dependent augmented reality (EDGAR) device, detecting at least some features of the 5G EDGAR device via one or more application programming interfaces (APIs) of a Capability and Performance Discovery (CPD) module of the 5G EDGAR device, wherein the one or more APIs of the CPD module include a network API from the 5G EDGAR device to the network, a device API from the CPD module to the 5G EDGAR device, and a plurality of functional APIs from the CPD module to the at least some features of the 5G EDGAR device, The steps include obtaining an index of at least some of the features of the 5G EDGAR device from the CPD module via the functional API, A step of controlling the execution of a 5G media streaming (5GMS) service according to the index of at least some of the features of the 5G EDGAR device. Methods that include...
2. The step of controlling the 5G EDGAR device and detecting at least some of the features of the 5G EDGAR device includes the step of controlling the 5G EDGAR device via a network API, The step of obtaining the index of at least some of the features of the 5G EDGAR device includes the step of obtaining the index via the network API. The method according to claim 1.
3. The above-mentioned features include at least one parameter among the media encoder and media decoder, The method according to claim 1.
4. The aforementioned features include the device capture capability of the 5G EDGAR device, and the device capture capability includes either camera capability or microphone capability. The method according to claim 1.
5. The aforementioned features include the device rendering capability of the 5G EDGAR device, and the device rendering capability includes either speaker capability or display capability. The method according to claim 1.
6. The aforementioned features include the raw hardware capabilities of the 5G EDGAR device, which include any of the following: CPU benchmark, GPU benchmark, memory bandwidth, and on-device storage capacity. The method according to claim 1.
7. The aforementioned features include the functional efficiency of the 5G EDGAR device, where the functional efficiency represents the relative amount of power consumption per function and the relative amount of parameters per function. The method according to claim 1.
8. The above-mentioned features include either the power state of the 5G EDGAR device or the battery depletion rate of the 5G EDGAR device, The method according to claim 1.
9. The index indicates either whether the combination of at least some features can be run in real time on the 5G EDGAR device, or the estimated overall battery consumption of the combination of at least some features. The method according to claim 1.
10. An apparatus configured to perform the method described in any one of claims 1 to 9.
11. A computer program that causes at least one processor to perform the method described in any one of claims 1 to 9.
Citation Information
Patent Citations
To provide users with feedback on power consumption within battery-powered electronic devices.
JP2013502181A
Systems and methods for providing zone functionality in networked media systems
US20100292818A1
Edge computing local breakout
US20220038554A1
Methods for discovery of media capabilities of 5g edge
WO2021225765A1