Apparatus for providing a synchronous input
By employing a prediction model to predict user states and synchronizing communication flows, the system addresses the challenges of low-latency communication and data synchronization in metaverse applications, achieving efficient and high-quality data representation.
Patent Information
- Application Number
- JP2024566302
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-22
- Filing Date
- 2023-05-11
- Publication Date
- 2025-05-30
AI Technical Summary
Current communication systems face challenges in achieving low-latency communication over long distances, especially in metaverse applications where high-quality data representation and synchronization of data streams from multiple users are required.
The proposed solution involves using a prediction model to predict the state of a user based on past sensor inputs, allowing for low-latency interaction by interacting with the prediction model rather than the user directly. Additionally, data compression is achieved by downloading the complex model during initialization and restricting sensor data sampling. The system also synchronizes communication flows based on determined communication parameters.
This approach enables low-latency communication over long distances, reduces communication overhead, and ensures high-quality data representation in metaverse applications, while also synchronizing data streams from multiple users effectively.
Smart Images

Figure 2025516579000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a communication system that requires low-latency communication between remote users. The present invention is suitable for (but not limited to) next-generation real-time communication systems and metaverse implementations where users can define a virtual reality space in which they can interact with a computer-generated environment and other users.
Background Art
[0002] Among the technical specification groups of 3GPP (registered trademark) services and system aspects (TSG SA), the main purpose of 3GPP (registered trademark) TSG SA WG1 (SA1) is to consider and study new and enhanced services, features, and capabilities of the 5G system, and identify any corresponding stage 1 requirements to be met by 3GPP (registered trademark) specifications. These service requirements are documented in the standard specifications under the responsibility of SA1. A related study is TR22.847, "Study on supporting tactile and multi-modality communication services (TAMMCS)". This study includes eight use cases and related requirements of the so-called "tactile Internet" (TI).
[0003] The International Telecommunication Union (ITU) defines TI as an Internet network that combines ultra-low latency with very high availability, reliability, and security. Mobile Internet enables the exchange of data and multimedia content while on the move. The next step is the Internet of Things (IoT) that enables the interconnection of smart devices. TI is the next evolution that enables real-time control of IoT. By enabling tactile sensations and haptics, it will add a new dimension to human-machine interaction and at the same time bring a revolution to machine interaction. TI enables humans and machines to interact with their environment in real time while on the move or within a specific spatial communication range.
[0004] In the IEEE publication P1918.1, "Tactile Internet: Application Scenarios, Definitions and Terminology, Architecture, Functions, and Technical Assumptions", in order to avoid adverse effects on the user experience, it is required that the cellular 5G communication system support a mechanism to assist in synchronizing multiple streams of a multimodal communication session (e.g., tactile, audio, and video). Furthermore, the 5G system should be able to support the interaction with applications on the user equipment (UE), support the data flow for grouping information within a single tactile and multimodal communication service, and apply the policies provided by third parties to the flows associated with the applications. The policies can include the set of UEs and data flows, the expected quality of service (QoS) handling and associated trigger events, as well as other adjustment information.
[0005] Figure 3 shows the scenario addressed by the present invention. Two persons A and B are trying to interact (communicate) in a metaverse or real-time communication application. For this purpose, persons A and B have corresponding rendering devices, e.g., VR devices and corresponding sensor devices. Persons A and B are separated by a distance d. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION
[0006] There are mainly two problems in this scenario: First, when the metaverse application requires a maximum latency of 1 ms, since information propagates at 300,000 km / s, the maximum distance between persons A and B is d = 300 km. Second, since high-quality data representation is required, the sensor device needs to sample a high-quality representation of the person to be sent to the other party's rendering device, which leads to high data rate requirements.
[0007] Furthermore, more than three people participate in the metaverse, and different people are in different locations. For example, assume that users U0, U1, and U2 are interacting with each other at different locations L0, L1, and L2. User U0 receives data from users U1 and U2. When the data arrives at U0, it is necessary to synchronize the data streams generated at U1 and U2 and the data stream of the UE local to U0.
[0008] Therefore, the third problem is how to synchronize the data streams transmitted from UEs at different locations.
[0009] The present invention aims to mitigate the above problems.
[0010] Another object of the present invention is to enable metaverse interaction between remotely located people and reduce communication overhead.
Means for Solving the Problems
[0011] According to a general definition of the present invention, a first person interacts with a second person in the metaverse through a prediction model of the second person located near the first person. The prediction model predicts the state at time t based on sensor inputs sampled at times t(0)=t - d*c, t(1)=t - d*c - T,..., t(N - 1)=t - d*c - (N - 1)*T) (where c is the speed of light).
[0012] This approach enables any one or more of the following: First, low latency: Instead of direct interaction, Person A (B) interacts with the prediction model of Person B (A). The model of Person B (A) predicts the actions of Person B (A) based on the inputs collected by sensor device B. The prediction model of Person B (A) is used in rendering device A (B).
[0013] The model of a person is constructed when that person participates in the metaverse and can be deployed at an appropriate location when that person wants to interact with other people. This appropriate location must be near the person.
[0014] Second, data compression: It is achieved, for example, by a) downloading the complex model of a person to an appropriate location during initialization and b) restricting the sensor data that needs to be sampled for an appropriate prediction model leading to a realistic representation of that person.
[0015] Another general definition of the present invention proposes the synchronization of (predicted) communication flows from different user devices at different locations as follows, for example: First, determining communication parameters (e.g., latency) between users. Second, synchronizing the flows based on the communication parameters. Third, applying this synchronization to the use of the prediction model.
[0016] According to a first aspect of the present invention, an apparatus for providing synchronized inputs to a third device is proposed. This apparatus a. A memory for storing communication parameters shared between the third device and the first device and between the third device and the second device, b. A communication unit for receiving a first communication flow from the first device and a second communication flow from the second device, and the first communication flow and the second communication flow are synchronized based on the communication parameters.
[0017] According to a second aspect of the present invention, a system is proposed that includes at least one communication device of the first aspect of the present invention and at least one remote second device for transmitting the received communication flow.
[0018] According to a third aspect of the present invention, an apparatus for providing a derived prediction model of a first device to a third device is proposed. This apparatus a. a storage unit for storing a general prediction model of the first device, and b. a communication unit for receiving communication parameter characteristics of a communication link between the first device and the third device, and c. a calculation unit capable of determining a derived prediction model of the first device, is provided, and the derived prediction model is obtained based on the general prediction model and the received communication parameters.
[0019] According to another general definition of the present invention, an apparatus for providing synchronized inputs to a third device is proposed. This apparatus a. a memory for storing communication parameters shared between the third device and the first device and / or between the third device and the second device, and b. a communication unit for receiving a first communication flow from the first device and / or a second communication flow from the second device, is provided, and the communication flows are synchronized based on the communication parameters.
[0020] In a first variant of this general definition of the present invention, the communication parameters include - the distance between the third device and the first device and / or the distance between the third device and the second device, - the latency from the first device to the third device and / or the latency from the second device to the third device, or - other communication parameters shared between the first device and the third device and / or between the second device and the third device, is included.
[0021] In a second modification, the above device further comprises a calculation unit that receives, as an input, communication parameters between the first device and the third device and executes a prediction model that predicts a control input of the third device, and the control input includes predicted communication parameters between the first device and the third device.
[0022] In a third modification, the above device comprises a calculation unit that receives, as an input, at least the communication flow of the first device and executes a prediction model that predicts a control input of the third device.
[0023] In a fourth modification, the prediction model is - a model derived from a general prediction model that follows parameters shared at least between the first device and the third device, and - a generation model, and is at least one of them.
[0024] In a fifth modification, the communication parameters are obtained by executing a protocol on the first device.
[0025] In a sixth modification, the communication parameters are set by a management device.
[0026] In a seventh modification, the communication parameters are - the latency between the third device and the first device and / or the second device, - QoS, - the distance between the third device and the first device and / or the second device, - the computational requirements for processing the communication, - the computational capacity for processing the communication, - the memory requirements for processing the communication, - the memory capacity for processing the communication, - the available bit rate, - the number of communication parties, - The (relative) position of the third device with respect to the first device and / or the second device, - The (relative) velocity of the third device with respect to the first device and / or the second device, - The (relative) acceleration of the third device with respect to the first device and / or the second device, - The (relative) rotation of the third device with respect to the first device and / or the second device, can be at least one of the above.
[0027] Under this other general definition, the present invention also targets a system comprising at least one third device including the devices defined in the above definitions and their modifications, and at least one remote first device for transmitting the communication flow received by the third device.
[0028] Under this further other general definition, a method for providing synchronized input to a third device is also proposed. This method includes a. Storing in a memory communication parameters shared between the third device and the first device and / or between the third device and the second device; b. Receiving, by a communication unit, a first communication flow from the first device and / or a second communication flow from the second device; c. Synchronizing the communication flows based on the communication parameters.
[0029] Under this further general definition, a device for using a prediction model of a first device is also proposed. This device includes a. A storage unit for storing the prediction model of the first device; b. A communication unit for obtaining communication parameter characteristics of a communication link between the first device and the third device; c. A computing unit capable of using the prediction model of the first device, and the output of the prediction model is obtained based on the prediction model and the communication parameters.
[0030] According to the first modification example, the output of the prediction model enables a synchronized communication flow between the first device and the third device.
[0031] According to the second modification example, the output of the prediction model is a derived prediction model, and the number of required input parameters of the derived prediction model is less than the number of required input parameters of the prediction model.
[0032] Furthermore, under this general definition, a method for using the prediction model of the first device is also proposed. This method includes a. a step of storing the prediction model of the first device, and b. a step of obtaining the communication parameter characteristics of the communication link between the first device and the third device, and c. a step of using the prediction model of the first device, wherein the output of the prediction model is obtained based on the prediction model and the communication parameters.
[0033] It should also be understood that the preferred embodiments of the present invention can be dependent claims corresponding to the independent claims or any combination of the above embodiments.
[0034] It should be understood that some or all of the above aspects can be realized by a computer program including instructions that enable the implementation of the method targeted by the present invention when executed on a computer.
[0035] These and other aspects of the present invention will become apparent from the embodiments described below and will also be described with reference to the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0036]
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9A-D
Figure 10
Figure 11
Figure 12
Embodiments for Carrying Out the Invention
[0037] Next, embodiments of the present invention will be described based on a cellular communication network environment such as 5G. However, the present invention can also be used in relation to other wireless technologies where a TI or a metaverse application is provided or can be introduced. The present invention is also applicable to other applications such as video streaming services, video broadcast services, or data storage.
[0038] Throughout this disclosure, the abbreviation "gNB" (5G term) or "BS" (base station) is intended to mean a wireless access device such as a cellular base station or a WiFi (registered trademark) access point, or an ultra-wideband (UWB) personal area network (PAN) coordinator. The gNB can be composed of a central control plane unit (gNB-CU-CP), a plurality of central user plane units (gNB-CU-UP), and / or a plurality of distributed units (gNB-DU). The gNB is part of a radio access network (RAN) and provides an interface to functions within a core network (CN). The RAN is part of a wireless communication network. This implements a radio access technology (RAT). Conceptually, it exists between communication devices such as mobile phones, computers, or any remotely controlled machine and provides a connection to its CN. The CN is the core part of the communication network and provides a number of services to customers interconnected via the RAN. More specifically, it manages communication streams via the communication network and, in some cases, other networks.
[0039] Furthermore, the terms "base station" (BS) and "network" may be used synonymously in this disclosure. This means that, for example, when it is written that the "network" performs a particular operation, that operation may be performed either by the CN functions of a wireless communication network or by one or more base stations that are part of such a wireless communication network, and vice versa. It can also mean that part of the function is performed by the CN functions of a wireless communication network and part of the function is performed by the base station.
[0040] Furthermore, the term "metaverse" is understood to refer to a persistent shared set of interactive spaces where users can interact with each other in parallel with mutually perceivable virtual features (i.e., augmented reality (AR)), or where those spaces are entirely composed of virtual features (i.e., virtual reality (VR)). VR and AR are sometimes generally referred to as "mixed reality" (MR).
[0041] Furthermore, the term "data" is understood to refer to a representation in a known or agreed-upon form of information that is stored, transferred, or otherwise processed. The information can consist of, in particular, one or more channels of synchronous audio, video, images, tactile, motion, or other forms of multimedia information. Such multimedia information can be obtained from sensors (e.g., microphones, cameras, motion detectors, etc.) or can be partially or fully synthetic (e.g., a real actor in front of a synthetic background).
[0042] The term "data object" is understood to refer to one or more data sets as defined above, accompanied by one or more data descriptors that optionally provide additional semantic information about the data that affects processing procedures at the transmitter and receiver. The data descriptors can be used, for example, to describe how the data is classified by the transmitter and how it should be rendered by the receiver. As an example, data representing an image or video sequence is decomposed into a set of data objects that collectively describe the entire image or video, and these are processed (e.g., compressed) individually in an optimal manner for the object and its semantic context, substantially independently of other data objects.
[0043] Furthermore, the term "data object classification" is understood to refer to the process by which data is divided or segmented into multiple data objects. For example, an image can be divided into multiple parts, such as a forest in the background and a person in the foreground (as will be described later in connection with, for example, FIG. 8). Data objects are classified using data object classification criteria. In the present disclosure, such criteria include at least one of a measure of the semantic content of the data object, the context of the data object, a class of compression techniques that are optimal for retaining sufficient semantic content for a given context, and the like.
[0044] Furthermore, "compression technique" is understood to refer to a method of reducing the size of data in order to improve transmission and storage efficiency. For example, there is a method of deleting redundant data and data that is considered semantically imperceptible to the end user, and efficiently encoding the remaining data so that a faithful or nearly faithful semantic representation of the original data can be reconstructed.
[0045] Furthermore, "compression or reconstruction model" is understood to refer to a tool that can be used to assist in the compression and reconstruction of data and a repository of data objects. For example, the model includes algorithms used for the analysis and compression of data objects and data objects that can be used as the basis for generating compression techniques. Advantageously, the model can be shared or owned by a transmitter and a receiver and / or updated or optimized according to the semantic content of the data being transferred.
[0046] Note that in the present disclosure, only the blocks, components, and / or devices related to the proposed data distribution function are shown in the accompanying drawings. Other blocks are omitted for the sake of brevity. Furthermore, blocks designated by the same reference numeral are intended to have the same or at least similar functions, and thus their functions will not be described later.
[0047] Figures 1A and 1B schematically show a network architecture (e.g., IEEE P1918.1 architecture) considered for implementing the metaverse. This architecture includes an actuator gateway (AG), actuator node (AN), controller node (CN), control plane entity (CPE), gateway node (GN, GNC corresponds to GN&CN), human-system interface node (HN), network controller (NC), sensor / actuator (S / A), computing and storage entity (SE), sensor gateway (SG), sensor node (SN), tactile device (TD), tactile edge (TE), tactile service manager (TSM), user plane entity (UPE), access interface (A), first tactile interface Ta (TD-TD communication), second tactile interface Tb (TD-GNC communication), open interface (O), service interface (S), network side (N), network domain (ND), bidirectional information exchange (BIE), external application service provider (EASP), and dedicated low-latency network (LLNW).
[0048] The architectures of Figures 1A and 1B provide an overall communication architecture defined by a general-purpose approach that can be executed on any network including 5G. These architectures cover various modes of the interconnected network domain between two TEs (TE A, TE B). Each TE is composed of one or more TDs, and the TDs of TE A are for information. For example, tactile / haptic Communicate through ND to TD of TE B to meet the requirements of a given TI use case. ND can be any of a shared wireless network (e.g., 5G wireless access and core network), a shared wired network (e.g., Internet core network), a dedicated wireless network (e.g., point-to-point microwave or millimeter-wave link), or a dedicated wired network (e.g., point-to-point dedicated line or fiber optic link). Each TD can support one or more of the functions of sensing, actuation, tactile feedback, or control via one or more corresponding entities. The S or A entity refers to a device that executes the sensing function or the actuation function respectively without a network formation module. SN or AN refers to a device that has an air interface network connection module and executes the sensing function or the actuation function respectively. To connect S to SN or A to AN, an SG entity or an AG entity needs to be used respectively. These gateways provide a general-purpose interface for connecting to third-party sensing and actuation devices and another interface for connecting to SN and AN. Also, TD can function as an HN that can convert human input into tactile output and as a CN that has the necessary network connection module and executes a control algorithm to process the operation of the SN and AN systems.
[0049] GN is an entity with enhanced network formation capabilities that exists at the interface between TE and ND, and is mainly responsible for user plane data transfer. GN is accompanied by NC, which is responsible for control plane processing including admission and congestion control, service provisioning, resource management and optimization, and connection management intelligence to achieve the QoS required for the TI session. GN and CN (collectively labeled GNC) can exist either on the TE side (shown in Figure 1A) or the ND side (shown in Figure 1B), depending on the network design and configuration. GNC is a central node and is essential for compatibility with other new standards such as the 3GPP (registered trademark) 5G NR specification to facilitate interoperability with various possible network domain options. Placing GNC under ND, for example in 5G, is intended to support the option of absorbing the functions of GNC into the management and orchestration functions already included in ND. In Figures 1A and 1B, ND is shown to be composed of radio access points or base stations logically connected to CPE and UPE within the network core.
[0050] Users within the region of interest (ROI) are surrounded by a set of TDs linked to TE. TDs may include rendering actuators and / or sensors. The rendering actuator has the task of creating a metaverse environment around the user, such as VR glasses, 3D television (TV), holographic devices, etc. The sensor TD is a device that captures the user's behavior and / or environment, including video cameras, audio devices such as microphones, and tactile sensors. Generally, TDs are UEs as seen from the 5G system.
[0051] The TD within the ROI can be connected to the user's TE, for example, by wire or wirelessly. In the case of wireless, the UE can connect to a base station such as a 5G gNB or a WiFi (registered trademark) access point. The network formation infrastructure and computing resources of the TE coexist within the ROI or are located in a nearby edge server (at a distance less than the maximum edge distance) to ensure fast response.
[0052] To support the implementation of the following embodiments, at least one of the three communication functions is introduced. First, provide latency-based flow synchronization (LBFS), a function that can be executed on a device within the TE. This can also be deployed in a receiving TD that can determine communication parameters with a transmitting TD (or TS) and synchronize the communication flow based on these communication parameters, particularly the relative latency between TDs (or TEs). Second, the edge application of the TE is configured to execute an environment / person latency-dependent configurable prediction model (LDCPM) within the metaverse sessions of different TEs. Third, provide a model management and configuration function that can register a general-purpose model of the devices and / or people within the ROI and / or TE, save it in a database, and deploy the LDCPM reconfigured when determining communication parameters.
[0053] Another pioneering node in the architectures of FIGS. 1A and 1B is the SE, which provides both computing resources and storage resources to improve the performance of the TE and meet the delay and reliability requirements of end-to-end communication. The SE executes advanced algorithms that employ AI technology, in particular, to offload processing operations (such as tactile rendering, motion trajectory prediction, and sensory correction) that consume too many resources and too much energy to perform at the TD. The goal is to enable real-time connection awareness using predictive analytics while overcoming challenges and uncertainties along the path between the source TD and the destination TD, dynamically estimate the network load and rate fluctuations over time to optimize resource utilization, and share learned environmental experiences between different TDs. On the other hand, the SE also significantly impacts reducing the end-to-end traffic load and thus provides an intelligent caching function that can reduce data transmission delays. The SE can be present locally to the TE to improve the response speed to requests from the TD or GNC, or remotely in the cloud to provide services to the TE and ND. Furthermore, the SE can be either centralized or distributed. Each of these options has its own advantages and disadvantages in terms of delay, reliability, capacity, cost, and practicality. Communication between two TEs can be unidirectional or bidirectional and is based on a client-server model or a peer-to-peer model and can belong to any of the above use cases with corresponding reliability and delay requirements. For this purpose, the TSM plays an important role in defining the service characteristics and requirements between two TEs and disseminating this information to the main nodes of the TE and ND. The TSM also supports functions such as registration and authentication and provides an interface to the EASP of the TI.
[0054] The A interface provides the connection between the TE and the ND. This is the main reference point for the information exchange of the user plane and the control plane between the ND and the TE. Depending on the architecture design, the A interface can be either between the TD and the ND or between the GNC and the ND. Furthermore, the T interface provides the connection between the entities within the TE. This is the main reference point for the information exchange of the user plane and the control plane between the entities of the TE. The T interface is divided into two sub-interfaces, Ta and Tb, to support different modes of the TD connection. The Ta interface is used for the communication between TDs, and the Tb interface is used for the communication between TD and GNC when the GNC exists in the TE. Furthermore, the O interface provides the connection between any architecture entity and the SE, and the S interface provides the connection between the TSM and the GNC. The S interface transmits the control plane information. Finally, the N interface refers to any interface that provides the internal connection between the ND entities. This is usually treated as part of the network domain standard and can include sub-interfaces for both the user plane entities and the control plane entities.
[0055] Two broad categories of haptic information, namely, tactile or kinesthetic, are introduced, and these may be combined. Tactile information refers to the perception of information by various mechanical receptors in the human skin, such as surface texture, friction, and temperature. Kinesthetic information refers to the information perceived by the human body's skeleton, muscles, and tendons, such as force, torque, position, and velocity.
[0056] The first differentiating point of TI and related standards compared to 5G Ultra-Reliable Low-Latency Communication (URLLC, ITU-R M.2083) is related to the fact that TI must be developed to be able to meet its requirements over distances longer than the 150 km distance interval for a round trip (or 100 km for fiber) due to its 1 ms propagation time. Such capabilities can be achieved by network-side support functions incorporated into the TI architecture, as envisioned in the standard operation of IEEE1918.1. These capabilities can, for example, model remote environments using artificial intelligence (AI) approaches and may in some cases be partially or fully present in the TI end device (i.e., the client of TI / haptic information).
[0057] The second differentiating point is related to the fact that TI brings an application with unique characteristics implied by its application, and it is expected that this application can be deployed as an overlay network on top of a network or a combination of networks. It is not intended to be applied only in the context of 5G URLLC as a basic communication means.
[0058] Furthermore, in the above architectures of FIGS. 1A and 1B, data streams such as tactile feedback must also be synchronized, and the user expects to "feel" or "experience" an event when it occurs, regardless of whether the visually represented event is audible. Therefore, the synchronization of audio data, video data, and tactile data becomes very important. This can be achieved incidentally by receiver buffering, thereby completely removing the challenges of the communication network in achieving the required latency (e.g., jitter).
[0059] To meet strict end-to-end (E2E) QoS requirements, the architecture must also provide advanced operation and management functions such as lightweight signaling protocols, distributed computing and caching by predictive analytics, intelligent adaptation to load and network conditions, and integration with external application service providers (ASPs).
[0060] As a result, reliability, latency, and scalability can be considered as key performance indicators (KPIs), and ubiquity (rapid "latching" from TD to TI infrastructure), ad-hoc (minimal maintenance of TI network domains), and hybrid (scalable and minimal maintenance of TI rendezvous devices) can be considered as three main approaches for the bootstrapping of TI services and the instantiation of the architecture. The design of TI is essentially built on the concept of E2E sustenance managed in an operation setting mandated at the edge. That is, the TD at the edge declares its communication parameters and operation parameters (e.g., expected latency and reliability), communicates it to the TI architecture, and the TI architecture allocates the resources necessary to meet such requirements in both the bootstrap setup and E2E communication.
[0061] Figure 2 schematically shows a state diagram representing the operating states of a finite state machine of the overall operation for implementing a metaverse application.
[0062] The TD device can start from the registration phase (REG), which is defined as the act of establishing communication with the TI architecture. Under the skewed TI paradigm, registration is performed at the GNC, which may include TI components from the ND such as the TSM. The selected application can provide a user interface for the user / ROI to register, for example, in the TSM. Registration can be performed either by the application itself or via the TSM, and registration may include registering the TDs that the user / ROI has. Registration may include registering the TD as part of the 5G system (5GS) to access the functions provided by the 5GS, such as quality of service (QoS), low latency, edge computing, or synchronization of communication flows, as part of the communication infrastructure.
[0063] At registration, the TSM can assign the TE to a user, ROI, and / or TD that is close to or suitable for its communication parameters. In some cases, the sensor TD may generate an output that is supplied to the rendering TD within the same ROI.
[0064] The "latching" point of the TD to start registration is sometimes called the TI anchor. At this stage, the TD probes the TI architecture to call for E2E communication and does not perform any other function other than latching to the TI architecture. In both the ad-hoc model and the hybrid model, this step may include involving the TSM (potentially via the GNC in the former) to establish registration.
[0065] The next state depends on the type of TD. In the case of a low-end SN / AN, the TD can have a "parent" designated in its immediate vicinity, and the TD is initially associated (ASS) with this. This parent TI node then ensures reliable operation and assists in connection establishment and error recovery. If the TD device operates independently, this is an optional step. Some mission-critical TDs and new TDs may need to be authenticated (AUT) without a parent (Ap) before being permitted to participate / start a TI session (SS).
[0066] Another stage is an optional state where the TD (NATD) communicates with an authentication agent within the TI infrastructure to perform authentication. The TSM is an entity that can perform this task, for example, with the assistance of the SE as required, or along with a large amount of traffic. Next, the TD starts its E2E control synchronization (Ctrl Sync) and probes to establish a link to the end TE. In this state, the TD cannot communicate operational data but can focus on connection setup and relaying maintenance parameters. This includes setting the parameters of the interfaces along the E2E path that assist the ND in selecting the optimal path across the network to provide the required connection parameters. This state encompasses the path establishment phase and the route selection phase of TI operations. Usually, multiple tiers of the TI architecture are involved and communicate to ensure that a path that meets the minimum requirements set in the "setup" message is actually available and reserved.
[0067] When the TD participating in the TI session is a haptic node (HN) targeting haptic communication, the next state may involve the specific communication and establishment of haptic-specific information prior to actual data communication. This state may include the determination of the codec, session parameters, and messaging format specific to this current TI session. Different use cases may require different haptic exchange frequencies, but all haptic communications are expected to start from the haptic synchronization state (H-Sync) to determine the initial parameters. Future changes to the codec and other haptic parameters are treated as data communication in the "operation" state (OP). This ensures that all haptic communications can perform the initial setup regardless of future updates to the parameters that may be included in the operation data payload.
[0068] After that, all TD components transition to the operation state. In this state, the E2E path is established, all connection setup requirements are met, and the TE is ready to exchange TI information. During operation in this state, a TD may detect intermittent network errors (ERR), in which case the TD transitions to the "recovery" mode (REC), and the specified protocol takes over the error checking and potential correction mechanism to attempt to re-establish reliable communication. When it is determined that the error is intermittent and resolved, the TD returns to the operation state. If the error persists for some reason, the TD returns to synchronization control to re-detect whether the E2E path is actually available under the operation requirements set by the edge user.
[0069] Finally, when the TI operation is completed successfully, the TD transitions to the "termination" phase (TERM), in which all resources that have been dedicated to this TD can be released and returned to the TI management plane. If this was initially processed by the NC, the resources return to the NC. More typically, the TSM is involved in the provisioning of TI resources.
[0070] Figure 3 shows two persons P A and PB Schematically shows an exemplary metaverse scenario in which an attempt is being made to interact with the metaverse. For this purpose, two persons have corresponding rendering devices RDA and RDB (e.g., virtual reality (VR) devices) and corresponding sensor devices SDA and SDB. Two persons P A and P B are separated by a distance d. Since high-quality data representation is required, the sensor devices SDA and SDB need to sample a high-quality representation of the person to be transmitted to the other's rendering device. Therefore, high-quality communication with reduced communication overhead is required.
[0071] The following presents embodiments for enabling high-quality communication and interaction with low communication overhead achieved by compression.
[0072] FIG. 4 schematically shows a block diagram of a network architecture for implementing some of the embodiments.
[0073] The transmitter (Tx) 22 is understood to be a device (e.g., a 5G UE) that detects or generates data to be compressed. The data or compressed data is transferred to the network (NW) 20 via an access link (AL) (e.g., a 5G New Radio (NR) wireless link). Further, the receiver (Rx) 28 (e.g., a 5G UE) is understood to be a device that renders data or compressed data. The data or compressed data is transferred from the network 20 via an access link (AL) (e.g., a 5G NR wireless link).
[0074] Furthermore, a network (NW) 20 is provided, which is understood to be any type of arrangement or entity used to transfer, store, process, and / or otherwise handle data or compressed data. The network 20 includes a plurality of logical components distributed across a plurality of physical devices. In an embodiment, network edge servers 24, 29 and a core network (CN) 21 are provided.
[0075] The network edge servers 24, 29 are understood to be devices that are physically close to a wireless access network (not shown) and provide data processing and storage services to user devices (e.g., UEs) that interact or communicate. The physical proximity provides ultra-low latency communication with the user devices. In an exemplary application, the transmit edge server (TxES) 24 and the receive edge server (RxES) 29 are provided to the transmitter 22 and the receiver 28, respectively, and provide a data storage function, a compression / decompression database, a compression / decompression function, and a negotiation function for negotiation with peer devices.
[0076] Furthermore, the core network 21 is understood to include the remaining portion of the network 20 that is used to transfer data between the transmitter 22 and the receiver 28, optionally via the respective edge servers 24, 29.
[0077] Furthermore, a shared storage (S) 23 is provided as a virtual device that represents the memory shared by both the transmitter 22 and the receiver 28, and / or the compression and decompression functions of the edge servers 24, 29. This can have multiple copies synchronized as needed and be physically located in one or more locations.
[0078] Communication parameters are understood to be parameters that affect the performance of communication or impose requirements on communication. Communication parameters include at least one of latency, QoS, distance between communication parties, computational requirements for processing communication, computational capabilities for processing communication, memory requirements for processing communication, memory capabilities for processing communication, available bitrate, number of communication parties, etc. Some of these parameters are related to each other. For example, the latency between two user devices depends on the distance between the devices, but also on other aspects such as the computational requirements and / or capabilities for processing communication. In particular, when a prediction model is involved in communication, the communication latency is affected by the available / required computational capabilities of both devices.
[0079] In some embodiments, only some of the above communication parameters are mentioned without loss of generality. For example, if only latency is mentioned in a particular embodiment, this should be understood to be latency or other communication parameters, particularly other communication parameters that affect the latency of the communication link.
[0080] In embodiments, by using data compression, sufficient image quality (e.g., a realistic representation of a person) can be provided on the receiving side. Such data compression may include conventional data compression methods. Further, specific data compression methods and corresponding devices or systems for compressing and decompressing data will be described in the following embodiments. Note that although these embodiments are beneficial for the specific applications mentioned in this specification, they can be implemented independently, for example, in contexts other than the metaverse, for other applications.
[0081] Embodiments may relate to a first type of system where a compression device obtains some input data and stores it efficiently in a storage unit. This is useful in cloud-based settings where it is necessary to store some data efficiently. Here, the compression device or encoder is used as a device (which may be a software product) that performs data compression. For example, the device receives data, splits (classifies) the data into one or more data objects according to appropriate criteria, compresses the data objects again according to appropriate criteria, and can store the compressed data objects together with any necessary compression models and other metadata required later to enable reconstruction of the source data. Appropriate criteria include taking into account the (semantic) accuracy required for reconstruction, the available storage space, and the processing resources for reconstruction. For the purposes of the present disclosure, the criteria also include taking into account the semantic content of the data and data objects and the compression models used.
[0082] Optionally, a decompression device (or decoder) may be provided, which is understood to be a device (which may be a software product) that retrieves compressed data from storage and uses an appropriate (decompression) compression model to reconstruct and render the original data with minimal loss of meaning. For this, a compression model repository can be used. This is understood to be a database that includes the tools and data objects used for data compression and is advantageously available to both the compression device and the decompression device. A subset of the repository's tools and / or data objects can be combined to form a compression model optimized in some way for a given sample or data type.
[0083] The storage unit is understood to be a device that is accessible by a compression device during storage and by a decompression device during retrieval, and that stores compressed data along with an associated compression model and metadata. It may also hold a compression model repository. The storage unit can take many forms. For example, it can be a web-based server or a physical device such as a memory stick, hard drive, or optical disk.
[0084] Other embodiments may relate to a second type of system where the compression / sending device exchanges data with a second decompression / receiving device in an efficient manner. This is useful in settings where a sending device (e.g., a cloud for a streaming service) desires to share data efficiently with a receiving device (e.g., a TV). In some cases, the sending device (e.g., the transmitter 22 in FIG. 4) has sensors capable of sensing / capturing data such as VR glasses, mobile terminals, or user equipment. In some cases, the receiving device (e.g., the receiver 28 in FIG. 4) has sensors capable of reproducing / rendering data such as VR glasses, mobile terminals, or user equipment. In some cases, an edge server (e.g., the edge servers 24, 29 in FIG. 4) is associated with one of the sending / receiving devices that takes over and / or shares some of the functions. In some cases, some of the sending / receiving devices are part of a telecommunications system such as a 5G system.
[0085] In connection with other embodiments of the second type of system, the compression transmission device is typically understood to be a compression device that compresses data in a substantially streaming manner, taking into account latency, computational overhead, or communication overhead at either the transmission device or the reception device as part of its compression criteria, and provides the compressed data to the transmission channel. Further, the decompression reception device is understood to be a decompression device that decompresses the data arriving at the transmission channel and typically renders it in real time and also typically taking into account latency, computational overhead, or communication overhead as part of its rendering technique.
[0086] (Optional) Edge servers can be provided adjacent to the compression transmission device to assist the compression transmission device by ensuring compression (post) - processing, on - the - fly provision or update of the compression model, or timely provision of compressed data by other means. Similarly, (optional) edge servers can be provided adjacent to the decompression reception device to assist the decompression reception device by ensuring decompression (pre) - processing, on - the - fly provision or update of the compression model, or timely rendering of decompressed data by other means.
[0087] Furthermore, (optional) sensors (e.g., audio, video) are provided to the compression transmission device and are typically, but not necessarily, configured as a device or an array of devices that capture certain aspects of a scene in real time. Examples include cameras, microphones, motion sensors, etc. Some devices can capture stimuli outside the human sensory range (e.g., infrared cameras, ultrasonic microphones) and "down - convert" that stimuli into a form perceivable by humans. Some devices include an array of sensor elements that give an extended or more detailed impression of the environment (e.g., multiple cameras capturing a 360° viewpoint, multiple microphones capturing a stereo or surround - sound sound field). Sensors of different modalities can also be used together (e.g., audio and video). In such cases, it is necessary to synchronize different data streams. The compression transmission device with sensors can be a VR / AR headset or simply a UE.
[0088] Furthermore, an (optional) rendering device (e.g., audio, video), which is typically a device or an array of devices that provides the decompression receiving device and renders certain aspects of the scene in real time, can be provided. Examples include a video display or projector, headphones, loudspeakers, tactile transducers, etc. Some rendering devices include an array of rendering elements (e.g., multiple video monitors, a loudspeaker array for rendering stereo or surround sound audio) that provide an enhanced or more detailed impression of the captured scene. Rendering devices of different modalities can also be used together (e.g., sound and video). In such cases, the rendering subsystem must ensure that all stimulus channels are rendered synchronously.
[0089] Furthermore, an (optional) communication manager, which is an entity that manages communication and can be either centralized or decentralized, can be provided to the second type of system. The communication manager aims to optimize communication in terms of latency, overhead, etc. The communication manager can be an entity within a communication network such as 5GS or an external entity such as an application function.
[0090] As a further (optional) element of the second type of system, the compression model repository includes information (such as data, machine learning (ML) models used to obtain the data, etc.) that is useful for reconstructing the data based on, for example, a prompt.
[0091] In the following embodiments (at least some of which can be combined to further improve performance), data compression and reconstruction are based on a prompt of a model (such as a latent diffusion model) from text to image that can be learned on the fly to represent objects not previously visible without the need to retrain the entire reconstruction model. This technique ("textual inversion") is performed quickly and iteratively as described by Rinon Gal et al. in "An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion" (available at https: / / TExtual-inversion.github.io / ).
[0092] Furthermore, in the embodiments, data compression and reconstruction are based on a system in which generation compression based on textual inversion is induced by the input image so that the reconstruction remains similar to the input image, as described by Zhihong Pan et al. in "EXTREME GENERATIVE IMAGE COMPRESSION BY LEARNING TEXT EMBEDDING FROM DIFFUSION MODELS" (available at https: / / ARxiv.org / pdf / 2211.07793.pdf).
[0093] Furthermore, in the embodiments, image classification can be performed quickly on low-performance devices, as described by Salma Abdel Magid et al. in "Image Classification on IoT Edge Devices: Profiling and Modeling" (available at https: / / ARxiv.org / pdf / 1902.11119.pdf).
[0094] Typically, diffusion models are good at reproducing items that form part of their training dataset but are poor at reproducing items that are not. Since the training dataset is obtained from the public Internet and diffusion models are costly to retrain, this poses a problem for reproducing inputs that are a priori unknown.
[0095] For example, a diffusion model can reproduce an image of a general person from a prompt such as "young man", but cannot reproduce an image of any one specific person (excluding some edge cases such as famous people).
[0096] The loss of realism between the reproduced output and the original observation is called "meaning loss" and differs in some respects from the distortion introduced by traditional codecs, especially by being more dependent on the original object being observed.
[0097] In addition to meaning loss, spatial and temporal stability are also issues. In related technologies (Neural Radiance Field, NERF), recent research has addressed techniques for improving spatial and temporal stability.
[0098] Recent techniques (such as those described above) have attempted to address the problem of meaning loss. The state-of-the-art in this field is represented by text inversion, i.e., dynamically learning an embedding that represents an object not previously seen without retraining the entire diffusion model. Thus, a guidance image is used to ensure that the learned embedding appropriately represents the observed reality (represented by the guidance image).
[0099] Figure 4 shows the idea underlying the present invention for enabling low-latency interaction between devices or people far apart, for example, in a metaverse application, while reducing communication overhead.
[0100] In accordance with a general definition of the present invention, a first person (or device) interacts, for example, with a second person (or device) in a metaverse application by means of a prediction model of the second person located near the first person. Next, the prediction model predicts the state or sensor input at time t based on sensor inputs sampled at previous times, for example, times t(0)=t - d*c, t(1)=t - d*c - T, …, t(N - 1)=t - d*c - (N - 1)*T) (where c is the speed of light and T is a given sampling period). Various types of prediction models are applicable.
[0101] This approach enables the following: First, low latency: Instead of direct interaction, person A (B) interacts with the prediction model of person B (A). The prediction model of person B (A) predicts the actions of person B (A) based on the inputs collected by sensor device B. The output of the prediction model of person B (A) is used in the rendering device of person A (B).
[0102] The model of person A is constructed when participating in the metaverse. This model is deployed at an appropriate location when person A wants to interact with another person B. This appropriate location is preferably near person B (for example, running on a local edge server).
[0103] Second, data compression is achieved by a) downloading a complex model of the person to an appropriate location during initialization, and b) restricting the sensor data that needs to be sampled for an appropriate prediction model leading to a realistic representation of that person, as seen, for example, in a generative model.
[0104] Another idea underlying the present invention relates to the synchronization of (predicted) communication flows from different user devices at different locations by: First, determining communication parameters (such as latency) between users. Second, synchronize the flow based on (communication) parameters. Third, optionally, apply this synchronization to the use of a prediction model.
[0105] Communication parameters are parameters that affect the performance of communication or impose requirements on communication. Communication parameters include the following: - Latency, - QoS, - Distance between communication parties, - Computational requirements for processing communication, - Computational capacity for processing communication, - Memory requirements for processing communication, - Memory capacity for processing communication, - Available bitrate, - Number of communication parties, - …
[0106] Note that some of these parameters are related to each other. For example, the latency between two devices depends on the distance between the devices, but also on other aspects such as the computational requirements and / or capabilities for processing communication, or the computational capacity for rendering the received data. In particular, when a prediction model is involved in communication, the communication latency is affected by the available / required computational capacity of both devices.
[0107] Note that in some embodiments, only some of the above communication parameters are mentioned without loss of generality. For example, if only latency is mentioned in a particular embodiment, this should be understood as latency or other communication parameters, particularly other communication parameters that affect the latency of the communication link.
[0108] Note that in some embodiments, in addition to communication parameters, other parameters such as transmission time, required data rate, ML model, compression rate, etc. are also considered in the adaptation of communication functions. In particular, the above parameters are related to (relative) position, speed, or orientation. For example, when an object / person (participating in a metaverse session) moves closer / farther at a higher speed, a higher communication speed is required to perform accurate rendering, so these parameters are relevant. Similarly, a more powerful model may be required for better prediction. Similarly, when the relative orientation of a person / device (participating in a metaverse session) changes, a higher communication speed may be required to perform better rendering or prediction of the data required for rendering.
[0109] General description of the system The present invention will be described in the context of a system focused on enabling applications such as metaverse applications, immersive teleconference systems, or immersive real-time communication among a large number of persons A, B, …, i, …, N.
[0110] As shown in FIG. 5, each person is surrounded by devices D connected to each other via communication links (e.g., 5G communication links). The communication link from device i (person i) to device j (person j) is characterized by a set of (communication) parameters P_ji that includes parameters related to the physical distribution of the devices (e.g., distance, speed, orientation) and actual communication parameters such as average latency (distance-dependent), jitter, bandwidth, reliability, QoS requirements, privacy requirements, etc. Note that some of these parameters are affected, for example, by allocating more resources (e.g., reliability), while other parameters (e.g., latency) cannot be changed. This set of communication parameters can be represented as a square matrix.
Number
[0111] Throughout the system, the following definitions related to IEEE 1918.1 and Figure 1 are used. Users within the region of interest (ROI) are surrounded by a set of tactile devices (TDs) linked to the tactile edge (TE). The TD may include a rendering actuator and / or a sensor. The TD rendering actuator has the task of creating a metaverse environment around the user and includes VR glasses, 3DTV, holographic devices, etc. The sensor TD is a device that captures the user's actions and / or environment and includes audio devices such as video cameras, microphones, and tactile sensors. The TDs within the ROI can be connected to the user's TE, for example, by wire or wirelessly. In the case of wireless, the UE can be connected to a base station such as a 5G gNB or a WiFi (registered trademark) access point. The network formation infrastructure and computing resources of the TE coexist within the ROI or are located in a nearby edge server (at a distance less than dmaxedge) to ensure fast response.
[0112] Generally, the TD is a UE from the perspective of a 5G system.
[0113] To support the implementation of the embodiments in the present invention, a communication function is introduced: First, latency-based flow synchronization (LBFS) is a function that can be deployed either on the devices within the TE or in the receiving TD that determines the communication parameters with the transmitting TD (or TE) and synchronizes the communication flow based on these communication parameters, particularly the relative latency between TDs (or TEs).
[0114] Second, a latency-dependent / communication parameter-dependent configurable prediction model (LDCPM) within the metaverse session in different TEs (e.g., provided by the edge application of the TE).
[0115] Thirdly, model management and setting functions where one or more of the following are possible: (1) registering general models of ROI / device / person in the TE, (2) saving it in the database, and (3) deploying the (reconfigured) LDCPM after determining the communication parameters.
[0116] This function is shown in FIG. 6. Here, First, although LDCPM management is shown as part of the 5GC system, it may be a complete or partial external function. The LDCPM model is shared by third-party applications, and LDCPM management saves the general LDCPM model in the application and deploys the (configured) LDCPM to the TD / TE as needed. For example, in some embodiments, the LDCPM operates on the TE and the LBFS operates on the TE. For example, in some embodiments, latency-dependent / communication parameter-dependent configurable prediction models are determined centrally, while in other embodiments, they are determined by entities that execute them based on, for example, currently measured communication parameters, such as the receiving TE or TD entity.
[0117] The LBFS or other related components shown in the present invention, for example, the components in FIG. 7 that can set / determine the latency / distance / parameters between TEs, can be generalized as blocks that can set / determine the (communication) parameters between TEs (or TDs), such as the orientation of the device, and act based on them, as shown in FIG. 8. For example, synchronize the communication flow as seen in some embodiments of the present invention.
[0118] The transmitting TE and the receiving TE may be the same. That is, the transmitting TD and the receiving TD are in the same TE and can thus be close to each other.
[0119] User creation, TD registration The application provides a user interface for the user / ROI to register, for example, in the TSM.
[0120] Registration can be performed either on the application itself or via the TSM.
[0121] Registration may include registering the TDs that the user / ROI has.
[0122] Registration may involve registering TDs as part of the communication infrastructure, for example, as part of a 5G system (5GS), so as to access the functions provided by the 5GS such as service quality, low latency, edge computing, or synchronization of communication flows.
[0123] At the time of registration, the TSM assigns a TE to a user / ROI / TD that is close to or suitable for its communication parameters.
[0124] In some cases, the sensor TD generates an output that is supplied to the rendering TD within the same ROI.
[0125] TD registration is also performed, for example, as part of the initial primary authentication when registering a UE in the 5GS. The TD can disclose its capabilities or NFs in the 5GS, for example, by examining the capabilities of the TDs stored in the UDM, UDR, or other AFs. Based on the location of the TD (the location is determined by the 5GS / TSM), the TD is assigned to a given TE. Based on the capabilities of the TD, the 5GS / TSM assigns computing or communication capabilities to the TD.
[0126] Registration, creation, deployment, and use of models The model of the user or ROI refers to one or more general models of the ROI / device / person within the TE. For example, it refers to models for each relevant TD sensor or models for each (person) feature that requires rendering.
[0127] The model can generate an appropriate (predicted) representation of the ROI or person and depends on different types of artificial intelligence / machine learning models / networks such as generative models.
[0128] User / ROI registration involves creating and registering a (general-purpose) prediction model M for the user's ROI, or the user himself / herself predicts the state of the ROI (or the user). For example, given N sensing inputs of past ROI / persons, for example, samples generated by sensor TD at times (t - cd, t - cd - T, t - cd - 2T, …, t - cd - NT) when the sampling period is evenly distributed, predict the state of the ROI / user at time t.
[0129] Other sampling distributions are also possible. For example, the samples have a higher density at more recent times and a lower density at older times: (t - cd, t - cd - T1, t - cd - 2T1, …, t - cd - NT1 / 2, t - cd - NT1 / 2 - T2, t - cd - NT1 / 2 - 2T2, …, t - cd - N(T1 + T2) / 2 Here, t1 < t2.
[0130] Generally, the model may be a function. M(D, cd, T, N)
[0131] Here, D represents the data sampled by sensor TD (within the transmitting TE) and used in the model of the transmitting user / ROI / TD / TE deployed at the receiving TE, cd refers to the communication parameter (e.g., cd, i.e., latency L) between the transmitting TE and the receiving TE, T represents the sampling frequency (or multiple sampling frequencies) of the sensing data, and N represents the number of samples used in the inference model.
[0132] This model may also be based on an AI / ML model such as a generative model that enables obtaining / generating a given data output (e.g., audio / video / ...) based on a (text) prompt. The prompt represents the user and is used to generate the user's data representation to be rendered. The prompt includes metadata indicating the location / orientation of the (generated) data representation of the user. This model definition also conforms to the aforementioned function, but in this case, D refers to the prompt + metadata information. For example, the prompt ["Bob", "Movement_Vector", "Orientation", t0] is a sensor sample, indicating that it is necessary to use the generative model to obtain the data representation of Bob moving with a given "Movement_Vector" and having a given "Orientation" at time t0. For example, a given model can use N such data samples.
[0133] The model is created by at least one of the following: First, placing sensors (TD) around the user or the user's surroundings, Second, asking the user to perform a specific movement, Third, measuring the actual movement using, for example, additional (calibrated) sensors, Fourth, measuring the movement with TD, Fifth, training a model that uses the output of TD as input using the calibrated measurements.
[0134] The prediction model may refer to a deep learning model such as a recurrent neural network model like a long short-term memory (LSTM) model. The prediction model may refer to a single model for the entire TE, or may refer to multiple models (e.g., a model for each TD within the TE or for each user (feature) within the TE).
[0135] In some cases, the trained model is trained to make predictions "far" in the future, i.e., it represents a large latency L = c.dmax linked to the maximum distance dmax between two TEs. In such a model, the TD in the transmitting TE sends samples of the environment to the TD in the receiving TE. The receiving TE is set with the prediction model of the transmitting TE and uses, when supplying the model, N received samples, e.g., the last N received samples. When the transmitting TE and the receiving TE are at a distance d < dmax, the receiving TE needs to delay the stream generated by the model (dmax - d).
[0136] In some cases, the trained model is trained to make predictions at any time in the future, i.e., it represents an arbitrary latency L linked, for example, to the distance d between two TEs or to the computational load of these TEs. In such a model, the TD in the transmitting TE sends samples of the environment to the TD in the receiving TE. The receiving TE is set with the prediction model of the transmitting TE and uses, when supplying the model, at most N received samples, e.g., the last N received samples. In this case, the generated data / stream can be used directly at the receiving TE.
[0137] In some cases, the number of samples (input to the model) required for prediction may be fixed, but when the latency L (i.e., d) is small, the number of samples exchanged may be reduced. For example, only samples sampled every other one at the transmitting TD may be communicated. The receiving TE can infer the missing samples (e.g., by interpolation) before supplying them to the model. This approach can reduce the communication overhead. Figure 7 shows a possible implementation of this approach.
[0138] In Figure 7, the TD sensor samples the environment of the transmission TE at a period T. Given the latency (or distance or communication parameter, or other parameter of the sampled scene) value set (e.g., by the TSM or application) or determined (e.g., distributed by the transmission TE and the reception TE) between the transmission TE / reception TE, the transmission TD can determine the compression parameter, e.g., the subsampling frequency of the samples sampled by the transmission TD sensor. Next, the compressed samples are transmitted from the transmission TD to the reception TD by the transmission TD. The samples arrive with a delay of dc seconds. The reception TD has set or determined the same latency (or distance) between the TEs and can use this to decompress the input signal (e.g., determine and apply the interpolation rate of the received signal). Thereafter, the model obtains N samples corresponding to a period of N*T seconds in order to infer a signal that can control the actuator at the current time, i.e., without delay, at the reception TD.
[0139] For example, when the TD sensor returns the prompt of the generation model so that the prompt can be used to reproduce the expression of the remote user, when the latency is large, the transmission TD (TE) chooses to transmit samples more frequently so that the receiving side can better predict the future state of the remote user.
[0140] For example, when the TD sensor returns the prompt of the generative model so that the expression of the remote user can be reproduced using that prompt, when the user is moving quickly or when the user is quickly changing direction (e.g., rotating), the transmitting TD (TE) selects to send samples more frequently / at a higher data rate so that the receiving side can better predict the future state of the remote user and / or can better render it. In some cases, the number of samples required for prediction is variable. For example, when L (i.e., d) is small, a smaller number of samples N can be used for the deployed model. This can be achieved when the transmitting TE has multiple models trained for different (communication) parameters (e.g., different latency L / distance d) and selects the optimal model based on the communication parameters (e.g., distance). This can reduce the CPU needs at the receiving TE. An alternative is for the transmitting TE to train a general-purpose model, which is then converted to an adaptation / compression / fitting model when deployed to the receiving TE characterized by specific communication parameters such as the distance d (latency L). In a potential instantiation, the general-purpose model receives N samples as input, but the compression / fitting model receives < N samples as input when d < dmax. This can be achieved, for example, by combining a decompression algorithm with the general-purpose model of FIG. 7 into a single block "adaptation model" shown in FIG. 8. This adaptation / compression / fitting model is created at the transmitting TE (or transmitting TD) and is deployed directly to each receiving TE (or receiving TD), for example, after the TE (or TD) - to - TE communication parameters (e.g., latency / distance) are determined. This compression / fitting model can be created at a central entity after the TE (or TD) - to - TE communication parameters (e.g., latency / distance) are determined, and this central entity may trigger the deployment to each receiving TE (or receiving TD).
[0141] The transmitting TE and the receiving TE may be the same. That is, the transmitting TD and the receiving TD are in the same TE, and thus, they may be close to each other. The components shown in FIG. 7 as being able to set / determine the latency / distance between TEs can be generalized as blocks that can set / determine communication parameters between TEs (or between TDs, or within a TE) such as the orientation of the device, as shown in FIG. 8.
[0142] To obtain such an adaptation / compression model, there are multiple embodiments as follows. For example, First, combine the interpolation block with the general model of FIG. 7, or Second, use some techniques as shown in FIG. 9, or Third, cause the model to predict the output after T seconds and apply it multiple times when the delay is k.
[0143] In some cases, the model has an input N. As shown in FIG. 9a, given N samples, a prediction is made within a given time cd. When the transmitting TE and the receiving TE are far from that cd, the receiving TE masks some of the inputs (e.g., k inputs) of the model to indicate that they are not available and chooses to give a prediction at cd + kN seconds using a reduced model with N - k input samples. This is shown in FIG. 9b, where k = 3. Naturally, in this case, since some inputs are lost, the accuracy decreases. When the transmitting TE and the receiving TE are very close to each other, the receiving TE masks some of the inputs (e.g., m inputs) of the model to indicate that they are not available and chooses to use a reduced model with N - m input samples. This is shown in FIGS. 9c and 9d, where m = 11 and m = 9. The model in FIG. 9a is a general-purpose (high-precision) model that can be reused for multiple receiving TEs at different distances from the transmitting TE by masking or padding the inputs.
[0144] In FIGS. 9a - 9d, the masked input is shown as having an input value of 0, but other values are also possible. For example, the same value as the previous unmasked input, interpolation of the unmasked input, etc. are also possible.
[0145] This technique can also be used for data compression. For example, when the receiving TE is close, the transmitting TE decides to transmit only every other sample from the TD sensor and notifies the receiving TE that it is necessary to mask every other input in the model.
[0146] Given the fact that some inputs are (always) masked, this technique can be used to simplify a general - purpose model to a compressed model by combining operations, because it can combine operations.
[0147] In another embodiment, the general - purpose model is, for example, an LSTM network that can predict the value of a given fixed delay T, as shown in FIG. 10a. In this embodiment, when the receiving TE (or TD) is at a distance d = T / c from the transmitting TE (or TD), one iteration is required. when the receiving TE (or TD) is at a distance d = 2T / c from the transmitting TE (or TD), the model needs to be iterated twice, and the output of the first iteration is used as the input for the second iteration. when the receiving TE (or TD) is at a distance d = 3T / c from the transmitting TE (or TD), the model needs to be iterated three times, the output of the first iteration is used as the input for the second iteration, and the outputs of the first two iterations are used as the input for the third iteration.
[0148] Generally, when the receiving TE (or TD) is at a distance d = kT / c from the transmitting TE (or TD), it is necessary to iterate the model k times, the output of the first iteration is used as the input for the second iteration, …, and the output of the first (k - 1) iterations is used as the input for the last iteration.
[0149] Furthermore, or alternatively, in other prediction models, a prompt is used to play the content on the receiving side. The prompt is extracted from the input content on the sending side. When the prediction model is shared, the same content can be generated from the prompt on the receiving side. If the content is an object or a person, and the object or person is moving at a given speed, for example, the prediction model on the receiving side predicts where the object / person is located considering the speed of the person and the communication latency.
[0150] For example, assume that the transmitter observes the person "Oscar" on the left side of the image and Oscar is moving right at a speed of 10 m / s. Assume that the transmitter can extract OscarPrompt from the image. Assume that the width of the image is 3 m and it has a width in pixels of 1920 pixels. Assume that the receiver receives data from the transmitter with a latency of 30 ms. Assume that the transmitter sends "”OscarPrompt”, current_location:left_size;current_speed:10m / s" to the receiver, and the receiver takes the prompt "Oscar" to reproduce the image of Oscar and shows / display / renders the image of Oscar starting from pixel (10 m / s * 0.03 s) * 1920 pixels / 3 m = 192 pixels considering the speed information and latency information (from the transmitter to the receiver).
[0151] In this example, when the receiver receives data from two transmitters, for example, “”OscarPrompt”, current_location:left_size;current_speed:10m / s”, latency 30 ms “”AlicePrompt”, current_location:left_size;current_speed:10m / s”, latency 10 ms
[0152] Even if the data from Alice is received 20 ms earlier than the data from Oscar, the receiver will display the images of both Oscar and Alice superimposed. The reason is that when the model synchronizes the communication flow, for example, when predicting the current location of the person / object / content being rendered / displayed, it takes into account the communication delay / latency.
[0153] Configuration of the communication infrastructure In an embodiment, a metaverse application utilizing a communication infrastructure can configure the communication infrastructure.
[0154] This configuration can be performed via a TSM that adjusts the underlying network.
[0155] 5G (or 6G) TSM exists in 5GS (or 6GS).
[0156] The TSM interacts with the 5G TSM when both entities exist. This means that the TSM functions as an encompassing entity that orchestrates application communication (e.g., metaverse applications), while the 5GS TSM has responsibilities for 5GS. This is relevant even when multiple 5GSs (e.g., the 5GS of the home network and the 5GS of the visited network) are involved. Also, the 5G TSM of the home network functions as the encompassing entity, and the 5GS TSM of the visited network "reports" to it (the 5GS of the home network).
[0157] This configuration may be performed by a policy.
[0158] This policy may include configuration items of each TD within each TE (e.g., participating in a communication link or session).
[0159] The application adds an entry corresponding to the new TD or TE to the policy, for example, whenever a new TE (TD) joins a (new) (metaverse) communication session or at that time, or whenever the permission is changed or at that time.
[0160] TSM distributes the policy of the new TD (TE) to all existing TDs (TEs) that are already participating in the metaverse communication session.
[0161] TSM distributes a policy containing entries of all existing TDs (TEs) that are already participating in the metaverse communication session to the new TD (TE).
[0162] This setting may be a one-time setting or may be a metaverse session setting for a metaverse session among several TEs (for example, some users (A, B,..., i,...)).
[0163] The setting includes, for example, a policy that specifies at least one of the following: First, for example, the number of users, the QoS target depending on the relative latency, Second, the resynchronization delay applicable to the data flow originating from the transmitting TE (or the transmitting TD within the transmitting TE), and the resynchronization flow applicable to the receiving TE (or the receiving TD within the receiving TE): TAU_i, TAU_j, TAU_bm, TAU_am, as described in other embodiments. Third, the latency between TEs and the need for continuous monitoring of the update rate of the parameters, as described in other embodiments. Fourth, the latency delay requirement for each TD within the TE. Therefore, the compression or model is correspondingly adapted. Fifth, the prediction model for each TD within the TE. Therefore, TSM can deploy the model or its compressed model to other TDs / TEs in the communication session. Sixth, the need for QoS equalization and, if applicable, the method as in other embodiments.
[0164] Similarly, the communication infrastructure may also notify the metaverse application of communication parameters, the relative position of the TD, and / or other parameters such as the settings of the metaverse application.
[0165] Launch of the user's application in TDi / TE / i, learning / configuration of parameters such as communication parameters In an embodiment, when a user launches a metaverse application in which one or more (remote) users participate, the communication infrastructure needs to learn a set of parameters such as communication parameters with other users / ROIs. This can be done mainly in the following two ways: First, end-to-end learning from a new TD (or TE) to other TDs (or TEs). In this case, the new TD of the TE is set by the TSM or application using the identifiers of other existing TDs / TEs participating in the current communication / metaverse session. The existing TDs / TEs are also notified about the possibility of the new TD / TE. The existing TDs / TEs also obtain information about the new TD / TE from a repository or the like. A TD (or TE) can perform distributed learning of parameters such as communication parameters (e.g., distance or latency), i.e., distributed learning from TD to TD or from TE to TE. For example, two remote TDs measure the round-trip time and divide it by two to determine the latency (distance) of the communication link. For example, two remote TDs share their current speed / direction with each other so that they can obtain, for example, their relative speed / direction.
[0166] Second, TSM supported based on communication of parameters such as latency from TD / TE to (5G) TSM. In this case, the new TD of the new TE is registered in the system, and parameters (e.g., latency) from the new TD / TE are measured from / to the relevant point (e.g., the central) network formation infrastructure through which (all) communications of (all) participating TEs pass. For example, the remote TD measures the round-trip time with the central network formation infrastructure and divides it by 2 to determine the latency (distance) of the communication link. Alternatively, the central network formation infrastructure may perform this action. Given this measurement value, TSM calculates the end-to-end communication parameters between each pair of TD / TE, and a) sets the new TD of the new TE or the new TE with the communication parameters of the existing TD / TE, and b) sets the existing TD / TE with this communication parameter with the new TD / TE. For example, TSM collects information regarding the orientation of the TDs and determines their relative orientation. Next, TSM uses this information to adapt the communication parameters.
[0167] Other parameters and / or communication parameters are learned in a similar way, and / or the parameters and / or communication parameters are made public. For example, the (available) computing power of the TE / TD is shared with TSM.
[0168] TSM uses the learned parameters and / or communication parameters or makes them public / exchanges / shares them with applications (e.g., application functions inside and outside the 5G system).
[0169] The launch of the user's application in TDi / TE / i triggers the deployment of the prediction model in the remaining TD / TE In an embodiment, when the user starts a metaverse session, the model of the user's TE / TD is shared with TSM by the application.
[0170] Next, TSM determines which TD / TE already associated with the metaverse session requires the model.
[0171] TSM directly sets TD / TE with parameters between them (such as communication parameters / latency / distance), or determines parameters for TD / TE in pairs.
[0172] Next, TSM selects appropriate compression parameters.
[0173] TSM can calculate a compressed version of the deployed model depending on, for example: First, the communication parameters / latency / distance between TD / TE. For example, if the communication parameters between the source TD / TE and the target TD / TE indicate a large distance / latency, the deployed model becomes more complex with respect to the number of past samples used for prediction. Second, computing power: For example, if the computing power is not sufficient for a high-precision model, a simplified model that requires fewer computing resources for prediction is deployed. Third, other parameters such as the relative position / speed / orientation between TD / TE.
[0174] When the model is successfully deployed to all target TD / TEs, the user / network formation infrastructure is notified of this, which serves as an indication to start communication.
[0175] The metaverse application is also notified when the above occurs, enabling the user to start a metaverse session.
[0176] Alternatively, TSM exposes the above parameters, such as communication parameters or a subset thereof, to the metaverse application / user, and an appropriate model is obtained.
[0177] Operation - unicast flow and multicast flow In a unicast communication flow, it is necessary to maintain NM unicast flows for each sensing TD (N devices) to each actuator / rendering TD (M devices). This becomes less efficient as N and M increase.
[0178] A more efficient approach is the multicast approach, where each sensing TD multicasts its flow, and this flow is distributed to each of the subscribed rendering TDs. Even though it is still important to consider that multicast flows reach different rendering TDs / TEs at different times, this involves N multicast flows, and the TD / TE that receives the multicast flow at an early stage uses a compressed model of the transmitting TD / TE, while the TD / TE that receives the multicast flow at a later stage requires, for example, a less compressed model.
[0179] Therefore, the multicast flow includes multiple data streams adapted to compression models with different compression ratios.
[0180] Operation - synchronization of data flows In an embodiment, the transmitting TE synchronizes the outgoing communication flows originating from the sensing TDs within the TE. This involves, for example, a policy by which the TD or TE is set to determine a method for synchronizing communication flows originating from multiple sensing TDs within the TE that include sensors with different sampling frequencies, such as video sensors, audio sensors, or tactile sensors. Since these sensors may not be fully synchronized or the sampling instants may not be adjusted, the network formation infrastructure within the TE may resynchronize the communication flows between them. For this purpose, First, each of the data flows i originating from TD i within the TE is delayed by a time TAU_i, Second, the value TAU_i is set by the policy, Third, the TE / TD is set at the sampling instant, Fourth, set TE / TD with communication resources that enable synchronous acquisition of sampled data.
[0181] Furthermore, TE j aligns all its outgoing flows with the remaining data flows in the system by delaying all data flows by delay TAU_j.
[0182] To achieve this, the metaverse applications executed at sensing TD need to expose the data / parameters (e.g., related to sampling frequency, instant (sampling time), and / or clock, etc.) of each sensing TD to the underlying communication infrastructure (e.g., 5GS). The exposure is performed, for example, via TD itself or via third-party application communication with TSM.
[0183] When this information becomes available, a method for synchronizing the outgoing communication flows and a policy for determining the delay values TAU_i and TAU_j applied to the transmitting TD / TE for each outgoing communication flow are deployed in the network formation infrastructure (e.g., 5G).
[0184] In another embodiment, the receiving TD within the receiving TE is instructed to synchronize the incoming flow. Flow synchronization may also be applied to the TE for all receiving TDs.
[0185] This means that TD (or TE) is set with a policy that requires applying delays TAU_bm and / or TAU_am to each incoming flow.
[0186] Delay TAU_bm is applied before passing the data to the (compressed) model for inference of the (control) signal of a given TD.
[0187] This delay can finely synchronize the input from the received input data stream.
[0188] The delay TAU_am is applied to the (control) signal obtained from the model, for example, in the following cases. (1) While the receiving TE is located at a distance d less than dmax from the transmitting TE, a general-purpose model that generates a (control) signal corresponding to the maximum latency / distance or the current time is applied at the receiving TE, and thus, TAU_am = c(dmax - d), or (2) The predicted signal becomes negligible in the future.
[0189] As in the previous embodiment, all communication flows originating from the transmitting TE may also need to apply the delay TAU_j (if this delay is applied at the receiving TE, TAU_j may not be necessary at the transmitting TE).
[0190] Since the input to the model depends on the communication parameters between the receiving TD (TE) and the transmitting (remote) TD / TE, the output of the model (of the remote environment) corresponds to the current state of the remote environment at the transmitting TE.
[0191] All or part of the setting parameters (e.g., delay) can be set at the TD or TE.
[0192] All or part of the setting parameters (e.g., delay) can be stored, for example, in a database, a communication infrastructure, or an external metaverse application.
[0193] All or part of the setting parameters (e.g., delay) are exchanged between the communication infrastructure (e.g., a telecommunications network) and the external metaverse function or are made public to the external metaverse function.
[0194] Operation - flow synchronization based on relative latency without using the prediction model In related embodiments, the incoming flows (from multiple TEs) are synchronized based on a setting policy, particularly communication parameters, particularly the relative distance (or latency).
[0195] In particular, when TD1 (in TE1) receives two communication flows F21 and F31 characterized by latencies L21 and L31 such that L21 < L31 from TD2 (in TE2) and TD3 (in TE3), flow F21 is delayed by the time (L31 - L21) and thus synchronized with flow F31.
[0196] Operation - flow synchronization based on relative latency using the prediction model In related embodiments, since the latency within TE1 can be considered zero, the actions of user 1 within TE1 need to be rendered with a latency L31. This can be unpleasant for the user.
[0197] Therefore, instead of delaying communication flows that arrive at the TE at an early stage (e.g., communication flow F21), a prediction model linked to TE2 deployed within TE1, as shown in other embodiments (e.g., embodiments related to model registration, creation, deployment, and use), given previous samples received at TE1, serves to predict the timing of the state of TE2 at the current time.
[0198] Operation - continuous latency monitoring / prediction (e.g., for communication via satellite link) Depending on the situation, the latency / distance between the user / ROI is not static and varies depending on the underlying network formation infrastructure. For example, the communication link is based on a satellite link, and the distance between satellite links varies when the satellites are not in geostationary orbit. Similarly, communication links can be congested, which results in higher latency. Therefore, it is necessary to handle such latency-variable communication links.
[0199] To address this need, in embodiments that can be used in combination with other embodiments, the following methods need to be applied (either alone or in combination) (periodically). For example, the frequency is specified in the policy set by the TSM: First, the UE or the network formation infrastructure (i.e., 5GS) continuously monitors communication parameters (e.g., end-to-end latency) in a distributed manner and updates communication policies. For example, two remote TDs measure the round-trip time and divide it by 2 to determine the latency (distance) of the communication link.
[0200] Second, the communication infrastructure (e.g., 5GS) continuously monitors, predicts, and / or makes available to the UE or the network formation infrastructure the communication parameters (e.g., end-to-end latency) between the UE and the network formation infrastructure.
[0201] Third, the communication infrastructure (e.g., 5GS) requires the storage of the expected / average communication parameters (e.g., latency, congestion level, jitter level) of a communication link (e.g., between two routers) in order to estimate the communication parameters of the end-to-end communication link based on the selected communication path.
[0202] Fourth, the communication infrastructure (e.g., 5GS) discloses to an application function (e.g., a metaverse application function) either within the communication infrastructure such as 5G CN or outside the communication infrastructure the expected or predicted end-to-end communication parameters. This application function calculates a model adapted using the communication parameters, which can be shared and deployed in the communication infrastructure.
[0203] Fifth, the communication infrastructure (e.g., 5GS) communicates to an entity (e.g., TE) that executes a prediction model the expected or predicted end-to-end communication parameters (e.g., latency) with / among the remote peers with which the entity is communicating so that the entity can use the end-to-end communication parameters. Adapting the basic model of a remote communication partner to a fitted model of the remote communication partner that can be used to calculate predicted values based on an input stream received from the remote communication partner, and / or Calculating predicted values using the basic model of the remote communication partner in combination with an input stream received from the remote communication partner.
[0204] Sixthly, the communication infrastructure (e.g., 5GS) adjusts the basic model of the first communication partner. The usage method of the basic model is (always) adapted to the communication parameters between the first communication partner and the second communication partner provided by the communication infrastructure. For example, The communication infrastructure adapts or continues to adapt the basic model so that it is adapted based on the real-time communication parameters between the first communication partner and the second communication partner. The communication infrastructure adapts or continues to adapt the input to the basic model generated from the input data stream received from the second communication partner based on the real-time communication parameters between the first communication partner and the second communication partner.
[0205] The above technologies can also be applied / used to other parameters (e.g., the relative position / speed / orientation of TD).
[0206] Operation - continuous monitoring of the relative positions of the transmitter / receiver and adaptation of communication settings In some scenarios, when the relative position of the transmitter / receiver UE (such as TD) changes rapidly, it is natural to require a higher data rate to ensure accurate mutual rendering. This is particularly important when the UE is incorporated into, for example, AR / VR glasses. This is because the user's movement affects the movement of the UE and the user's perspective. When two remote users wearing AR / VR glasses participate in a remote (metaverse) application, if the relative position of the two is stable, the required communication flow rate can be kept low, but a higher flow rate is required when the relative position changes.
[0207] Relative position refers to relative speed, relative location, relative acceleration, relative orientation, etc.
[0208] Therefore, in an embodiment, the transmitting TD (TE) and the receiving TD (TE) share their positions (location, speed, acceleration, orientation, etc.). Thereby, the TD (TE) can determine their relative positions and thus can determine whether the required data rate is sufficient.
[0209] In a further embodiment, the transmitting TD (TE) and the receiving TD (TE) share their positions (location, speed, acceleration, orientation, etc.) with a central management entity. Thereby, the management entity can determine the relative positions of the transmitting TD (TE) and the receiving TD (TE) and thus can determine whether the required data rate is sufficient.
[0210] In a further embodiment, the TD (TE) / central management entity determines the data rate of a given application (e.g., a metaverse application) based on the relative positions of the transmitting TD (TE) and the receiving TD (TE).
[0211] In a further embodiment, the telecommunication system has an interface for sharing the position (location, speed, acceleration, orientation, etc.) of the TD and / or the relative position of the TD with an external application, and / or for receiving settings from the external application regarding the required data rate.
[0212] In a further embodiment, a device (e.g., a TD) or a management entity executes a prediction model for predicting the positions of other TDs and / or the relative positions with other TDs in order to estimate the required data rate.
[0213] In a further embodiment, the transmitting / receiving TD and / or the management entity orchestrates / adjusts / triggers / adapts the following using the current relative position or the predicted relative position: - Resource allocation in the RAN of the transmitting UE and / or the receiving UE - Data rate of the encoded data, - …
[0214] In a further embodiment, the device (e.g., TD) or management entity performs one or more of the following tasks: - Task 1: Requesting and / or receiving an instruction regarding communication parameters (e.g., between both UEs or between both RANs) and / or (relative) position parameters of the UE from the RAN or core network service (e.g., AMF, SMF, LMF, NWDAF) of the transmitting UE and / or receiving UE, - Task 2: Requesting and / or receiving an instruction of available resources within the RAN from the RAN or core network service (e.g., AMF, SMF, LMF, NWDAF) of the transmitting UE and / or receiving UE (e.g., which time slots are available, which resources are available to provide a given throughput with a given latency, which QoS / QoE is possible and / or guaranteed, which CPU capabilities are available), - Task 3: Instructing the necessary resources (e.g., which specific time slots are preferably allocated to the communication link between UEs, which CPU resources are required) to the RAN of the transmitting UE and / or receiving UE, - Task 4: Requesting the remote end device to adapt the transmission quality (e.g., by adapting the data rate), - Task 5: Publishing information to an external management function (e.g., AF) and / or receiving commands regarding Tasks 1 to 4 from the external management function. - Task 6: Executing some of these tasks in a specific order (e.g., first execute Task 1, then execute Task 2, and then execute Task 3), and / or - Task 7: After determining a change in the parameters of Task 1 (e.g., communication parameters (higher latency) or relative position (e.g., the orientation of the UE changes)), execute some of these tasks (e.g., Task 3 and / or Task 4).
[0215] This synchronizes the allocated resources within the RAN of the transmitting UE and / or the receiving UE, enabling, for example, the provision of necessary services without excessive latency.
[0216] Operation - QoS equalization Depending on the metaverse scenario, a user may have different / better communication parameters than other users. For example, as shown in FIG. 11, User 2 is between User 1 and User 3. Therefore, User 2 receives the communication flow from User 1 (or User 3) at twice the speed (half the latency) compared to when User 1 (or User 3) receives the communication flow from User 3 (or User 1).
[0217] When the technologies described in the above embodiments are applied, the communication flows resulting from the TD and TE of each corresponding transmitting user reach in synchronization with the receiving TD / TE of each user. This also requires the use of a prediction model. However, since User 2 is still close to User 1 and User 3, the predicted signal quality is still better for User 2 than for User 1 and User 3. This is because the input received by User 2 needs to be applied to a prediction model that requires predictions not too far in the future, while for example, User 1 needs to apply asynchronous received input to a prediction model that requires predictions far in the future.
[0218] In such a situation, the metaverse application instructs the (5G) TSM and TSM, the underlying network formation infrastructure, to equalize the QoS provided to the users involved in the metaverse communication session. This includes the following. For example, First, for example, imposing a penalty on a user who has a communication advantage due to its location by delaying the communication flow originating from such a user (in the example, User 2), and / or Second, provide a near-optimal model to a user (TE / TD) in a more favorable situation (in the example, User 2), and / or Third, allocate more computing resources or a more powerful model to some users (in the example, Users 1 and 3) in less favorable situations, and / or Fourth, set the parameters of the above actions.
[0219] This may sound paralogical, but it is useful in applications such as games. Otherwise, such users would always be at an advantage compared to other users.
[0220] To achieve this goal, the metaverse application instructs the core network formation infrastructure to provide QoS equalization. The core network formation infrastructure then sets up the communication flow, for example, by deploying policies in the edge network formation infrastructure to enforce QoS equalization.
[0221] Operation using split rendering The above embodiments are also applicable to an architecture using split rendering. Split rendering means that high-load rendering processing is performed by a device with high computing resources (for example, a tactile edge device (TED) such as an edge server), and at a later stage, user-specific or device-specific low-load rendering is performed locally, for example, on a tactile device (TD). With split rendering, the calculation can be offloaded to the TED to simply maintain the TD. When a split rendering architecture is used, the prediction model (as described in the embodiments related to model registration, for example) may be executed on the TED.
[0222] One or more prediction models may be executed for each user. One of these prediction models is, for example, for predicting a user's volumetric video (VV) representation so that the user can be represented photorealistically. TED may execute multiple prediction models (e.g., per-user prediction models) and may require synchronization of the eta streams of multiple users at different locations / TEs having multiple tactile sensors (e.g., based on embodiments related to "operation-data flow synchronization"). When TED executes prediction models for multiple remote users, TED renders the combined, time-synchronized, and time-predicted representations (e.g., VV-rendered representations) of all users participating in the metaverse session. "Time-synchronized" means that the generated data streams are aligned, i.e., they follow a common clock. A time-synchronized data stream received (from another remote TE) may arrive after a given time Delta compared to the local clock of the local TE. Thus, "time-predicted" means that the representation is predicted at a future time Delta to synchronize with the local clock of the local TE / TED. This time Delta depends on the latency or communication parameters between each pair of remote TE / TEDs.
[0223] TED consumes information about the local user (e.g., the local rendering device (e.g., TD) associated with the user in the local environment). For example, TED consumes the height, position, and orientation of the VR / AR glasses worn by the user. Using this information, TED can obtain a TD-specific representation of the environment consumed by the user's rendering device (TD). For example, this representation is a 2D representation of volumetric video rendering at the edge server from the perspective of the user's rendering device (e.g., VR / AR glasses). For example, TED uses the consumed / received data regarding the position of the VR / AR glasses to determine the relative position to another user's VR / AR glasses and adapts communication parameters such as the required bitrate accordingly.
[0224] In such an environment, the local TED requests the communication system to allocate communication resources so that the TD within the environment can continuously provide inputs regarding, for example, its pose. This is done in a time-deterministic manner. For example, 5GS allocates H resource blocks every m milliseconds to transmit data regarding the pose of the head. When this is done, the delay when the TED receives the user's pose is T = Tsensing + m + Tflight, where Tsensing is the processing delay from sensing the pose to transmitting the value, m is the delay due to discrete measurements, and Tflight is the propagation time from the TD to the TED. This includes allocating deterministic uplink communication resources by a mechanism similar to, for example, semi-persistent scheduling, whereby the TD can continue to transmit inputs in a reliable time-deterministic manner. This may require the TED to consider T by, for example, running a pose prediction model that enables predicting the actual current pose of the user given past samples.
[0225] Similarly, the TED requests the communication system to allocate communication resources so that the TD within the environment can continuously receive TD-specific expression inputs generated at that TED. The TED may also need to consider the transmission delay T = Trendering + m + Tflight in the local TE. Here, Trendering is the time required to update the rendering at the local TD after receiving the data, m is the delay due to discrete transmission times, and Tflight is the propagation time from the TED to the TD.
[0226] In this embodiment, the uplink communication including information regarding the TD (e.g., pose), i.e., the latency from the TD to the TED, as well as the latencies of the downlink communication and local rendering, are part of the communication parameters considered when synchronizing data streams from other users in other locations or applying prediction models.
[0227] Enhancement of the system architecture for next-generation real-time communication TR23.700-87 v1.0.0 describes the enhancement of the 5G system architecture for next-generation real-time communication, including the following. ● Enhancement of the IMS network architecture necessary to support AR telephony communication for various types of AR-capable UEs. ● IMS procedures including signaling and media processing that need to be changed to support AR telephony communication.
[0228] Solutions #8 and #9 in TR23.700-87 address these architecture enhancements. In TR23.700-87, ● It has been decided to use the data channel architecture as a baseline for supporting AR telephony communication. ● If the UE requires network support for media rendering, the architecture and procedures specified in Solution #9 are used. ● Otherwise, if the UE can perform media rendering without network support, the procedures specified in Solution #8 are adopted as the baseline for the terminal rendering process.
[0229] In an embodiment, the systems and functions described in Solution #8 of TR23.700-87 are extended to support some of the above embodiments. For this purpose, Figure 6.8.2-1 of TR23.700-87 and the following three procedures: (1) IMS multimedia telephony call, (2) establishment of the bootstrap data channel (DC), and (3) establishment of the application DC, including the related communication flows between two UEs, can be further improved by the embodiments of the present application.
[0230] AR telephony communication is exchanged via RTP. In an embodiment, these procedures can be extended by the following: ● Incorporate the ability to determine the end - to - end latency between UEs during the establishment of an application DC or during application data exchange, and / or ● Incorporate the ability to determine end - to - end parameters between UEs, such as relevant location or communication parameters, during the establishment of an application DC or during application data exchange, and / or ● Incorporate, for example, the deployment of a prediction model of the peer - side UE (or UE environment) after the establishment of an application DC is completed, and / or ● Set / synchronize communication flows (e.g., set QoS), and / or incorporate the ability to calculate resources (e.g., allocate CPU / storage resources to an edge server), policies (e.g., latency, sensor rate to apply), and / or compression algorithms.
[0231] In a further embodiment, the system and functions of Solution #9 of TR23.700 - 87 are extended to support some of the above - described embodiments. Figure 6.9.2.2 - 1 of TR23.700 - 87 shows the communication flow between two UEs in which an AR media processing network function (ARMR) responsible for extended reality (AR) communication media transmission and media rendering functions has a network rendering process. ARMR includes the following: ● AR rendering logic: Controls the application - based rendering logic for AR communication. ● AR media processing functions: Include a vision engine and a 3D rendering engine. The vision engine and the 3D rendering engine create a spatial map and render scenes, virtual human models, and 3D object models according to the field of view, pose, position, etc. transmitted from the UE using a data channel.
[0232] In Figure 6.9.2.2 - 1, ARMR plays the role of TE, TED, or TSM described in the above - mentioned embodiments. This solution / architecture / protocol of TR23.700 - 87 can be extended to perform the following: ● For example, incorporate the ability to determine end - to - end latency between UEs during the negotiation procedure for AR media rendering, during the AR session media renegotiation procedure for network rendering, or during network rendering, and / or ● For example, incorporate the ability to determine end - to - end parameters between UEs, such as relevant location or communication parameters, during the establishment of the application DC or during application data exchange, and / or ● In particular, with respect to step 20 (transmitting AR media from UE - A to DCMF via the application DC) and step 21 (transferring AR media from DCMF to ARMF), or with respect to step 24 (transmitting rendered audio / video from UE - A to P - CSCF / IMS - AGW via RTP) and step 25 (transferring rendered audio / video to ARMF), incorporate the function of synchronizing incoming data streams of different remote users located at different locations, and / or ● Incorporate the deployment of a prediction model of the UE (or UE environment) in the ARMF. ● Set / synchronize the communication flow (e.g., set QoS), and / or incorporate the ability to calculate resources (e.g., allocate CPU / storage resources to an edge server), policies (e.g., delay, sensor rate to apply), and / or compression algorithms.
[0233] In a further embodiment, the UE or ARMF responsible for UE UE1 synchronizes the incoming data streams (e.g., DSi received from UEi at UE1, where i = 2, …, N). The synchronization determines the delay Di of each DSi, determines the DSj that arrives with the largest delay Dj, delays DSi by the amount Dj - Di (where the delay can be performed by buffering DSi), and optionally passes all synchronized data streams to the corresponding prediction models PMi (where all PMi are set to predict for the same latency). Each PMi receives the (delayed) DSj as input and calculates a predicted value. This embodiment represents a distributed approach where each UE performs the above synchronization (similarly, each UE is linked to an ARMF that performs such synchronization and / or prediction instead of the UE). Since all PMs predict their outputs based on the same delay Dj, this embodiment applies QoS equalization.
[0234] In a further embodiment, as an alternative to the above embodiment, the data stream DSi is not synchronized with other DSs and is passed to PMi. Here, PMi is set to perform a prediction of the latency Di (instead of Dj). This embodiment applies QoS equalization. In this embodiment, since PMi predicts its output based on the delays Di (i = 1, …, N), QoS equalization is not applied.
[0235] In a further embodiment, a centralized network entity, such as an application server (AS), such as a centralized ARMF or an ARMF within a home network, or an AR application server as shown in FIG. 6.9.1.3-1 of TR23.700-87, is responsible for such synchronization and / or prediction of output on behalf of the relevant UE. In this embodiment, it is important to consider that the (communication) parameters between UEi and AS are UEi-specific. For example, the distance between UEi and AS (and thus the minimum latency / delay Di) is UEi-specific. The AS receives the data streams of UEi (i = 1,..., N) with a delay Di, and the AS can synchronize them, for example, as shown in the above embodiment. Next, the AS applies the received data stream (optionally synchronized) to a prediction model PMi that predicts the output (e.g., Oi) at least at time Di further into the future, such that when Oi is supplied to UEi, Oi is supplied on time based on the communication delay Di between the AS and UEIII.
[0236] In a further embodiment, the UE performs synchronization and / or prediction of the output on behalf of the relevant UE itself. Thereby, the UE is set by an ARMF or other network function, or by an AR application server.
[0237] Exemplary use cases The above embodiments can support multiple use cases.
[0238] The first use case is to enable a real-time teleconference service.
[0239] In this first use case, three users are participating in an immersive metaverse teleconference using 5GS. Users Bob, Lucas, and Yon are located in the United States, Germany, and China respectively. Each user is served by a local metaverse edge computing server (MECS) hosted on 5GS, and each server is located near the user it serves. When a user joins the metaverse videoconference, that user's avatar is loaded onto the metaverse edge computing servers of the other users. For example, the metaverse edge computing server near Bob hosts the avatars of Yon and Lucas.
[0240] The large distance between users (e.g., the distance between the United States and China is approximately 11,640 km) determines a minimum communication latency (e.g., 11,640 / c = 38 ms). This latency also varies due to multiple reasons such as latency introduced by hardware components (variable processing times) such as congestion, sensors, and rendering devices. This value is too high and variable for a truly immersive metaverse teleconference experience, so each deployed avatar includes one or more predictive models that are predictive models of the person it represents and that enable the local edge server to render the synchronized predicted (current) avatar representation of the remote user.
[0241] Figure 12 shows this exemplary scenario, in which the MECS at Location 3 (United States) executes the predictive models of the remote users (Yon and Lucas), receives the sensing data received from all users (Yon, Lucas, and Bob) as input, and generates the synchronized predicted (current) avatar representation of the users to be rendered on Bob's local rendering device.
[0242] In this first use case, the following preconditions and assumptions apply to this use case: 1. Up to three different MNOs operate the 5GS that provides the metaverse teleconference service. 2. Users Bob, Lucas, and Yon are subscribed to the metaverse teleconference service. 3. Each user (e.g., Bob) decides to participate in a teleconference session.
[0243] In this first use case, the following service flows need to be provided: 1. Each user (e.g., Bob) decides to participate in a teleconference session and agrees to the deployment of an avatar. 2. The metaverse sensors in each user sense the real-time records of each user. The sensed real-time representation of each user is distributed to the metaverse edge computing servers of other users within the metaverse teleconference session. 3. Each metaverse edge computing server applies the incoming data stream representing each user located remotely to the corresponding avatar prediction model considering the current communication parameters / performance (e.g., latency), and creates a combined and synchronized current representation of the remote users as input to the rendering device within the local TE.
[0244] In this first use case, the main post-condition is that each user enjoys an immersive metaverse teleconference.
[0245] In this first use case, there are new requirements necessary to support this use case. That is, the 5G system must provide a means to synchronize the data streams of multiple metaverse (sensor and rendering) devices locally associated with different users in different locations. The second requirement is the need to support the distribution and execution of the prediction model at the edge server.
[0246] The second use case is to enable real-time teleconference services.
[0247] In this second use case, three users are participating in an immersive metaverse teleconference using 5GS. Users Bob, Lucas, and Yon are located in the United States, Germany, and China respectively. Each user is served by a local metaverse edge computing server hosted on 5GS, and each server is close to the user it serves. When a user participates in a metaverse teleconference, the avatar of each user is loaded onto the metaverse edge computing servers of other users. For example, the metaverse edge computing server near Bob hosts the avatars of Yon and Lucas.
[0248] The large distance between users (for example, the distance between the United States and China is about 11,640 km) determines a minimum communication latency (for example, 11,640 / c = 38 ms). This latency also varies for multiple reasons, such as delays introduced by hardware components such as congestion, sensors, and displays. Since this value is too high for a truly immersive metaverse teleconference experience, each deployed avatar includes one or more prediction models of the person it represents, which enable the predicted (current) representation of the person to be rendered at the edge server.
[0249] Each metaverse edge computing server can combine data streams from other users to create a meaningful real-time representation of the user. Since the rendering capabilities of the user's metaverse rendering device are limited, a split rendering approach is applied. In the split rendering approach, the metaverse edge computing server performs the computationally intensive rendering and distributes a personalized view to the user's metaverse rendering device.
[0250] In this second use case, the following preconditions and assumptions apply to this use case: 1. Up to three different MNOs operate 5GS that provides a metaverse teleconference service. 2. Users Bob, Lucas, and Yon are subscribed to the metaverse teleconference service. 3. Each user (e.g., Bob) decides to participate in a teleconference session.
[0251] In this first use case, the following service flows need to be provided: 1. Each user (e.g., Bob) participates in a teleconference session and decides to consent to the deployment of an avatar. 2. The metaverse sensor in each user senses the real-time recording of each user. The sensed real-time representation of each user is distributed to the metaverse edge computing servers of other users within the metaverse teleconference session. 3. Each metaverse edge computing server applies the real-time data stream representing each remote user to the corresponding avatar prediction model in consideration of the current communication parameters / performance (e.g., latency), and renders the combined and synchronized current representation of the remote users. 4. Each metaverse edge computing server continues to collect the current pose of the local user, such as the position, height, and head orientation in the local user's room, using the metaverse sensors around the local user. 5. Each metaverse edge computing server continues the following processing: (1) The rendered combined and synchronized current representation of the remote users while considering the following points (2) The current pose of the local user Thereby, obtain the view of the remote users for the local user that can be distributed to the metaverse rendering device of the local user.
[0252] In this second use case, the main post - condition is that each user enjoys an immersive metaverse teleconference.
[0253] Even in this second use case, there are new requirements necessary to support this use case. That is, the 5G system must provide a means to synchronize the data streams of multiple metaverse (sensor and rendering) devices locally associated with different users at different locations. The second requirement is the need to support the distribution and execution of prediction models at the edge server.
[0254] Therefore, a method and a corresponding edge server for executing such a method can be defined, and this method includes receiving a prediction model at the edge server, receiving data from remote user equipment and / or tactile devices within the area of interest, rendering the rendered data from the received data, for example, by inferring missing samples, and sending the rendered data to the destination user equipment.
[0255] General considerations Although the present invention has been described in the context of virtual spaces such as the metaverse, its application is not limited to operations of such a kind. Low - latency systems such as industrial IoT systems will also benefit from the teachings of the present invention and its embodiments.
[0256] Other variations of the disclosed embodiments can be understood and effected by those skilled in the art in practicing the invention according to the claims, from a study of the drawings, the disclosure, and the appended claims. In the claims, the term "comprising" does not exclude other elements or steps, and the singular does not exclude the plural. A single processor or other unit may perform the functions of several items recited in the claims. The mere fact that certain means are recited in mutually different dependent claims does not mean that these means cannot be used advantageously in combination. The above description details particular embodiments of the invention. However, it will be understood that the invention may be practiced in many ways and is not limited to the disclosed embodiments, however detailed the above description may be in the text. Note that the use of particular terms in describing particular features or aspects of the invention should not be construed as implying that the term is redefined herein to include any particular features of the features or aspects of the invention with which the term is associated. Further, the expression "at least one of A, B, and C" shall be understood as being disjunctive, i.e., as "A and / or B and / or C".
[0257] A single unit or device may perform the functions of several items recited in the claims. The mere fact that certain means are recited in mutually different dependent claims does not mean that these means cannot be used advantageously in combination.
[0258] The described operations, such as those shown in the above embodiments, may be implemented as program code means of a computer program and / or as dedicated hardware of related network devices or functions. The computer program can be stored and / or distributed on a suitable medium, such as an optical storage medium or a solid-state medium, supplied together with or as part of other hardware, but can also be distributed in other forms, such as via the Internet or other wired or wireless communication systems.
Claims
1. An apparatus for providing synchronized inputs to a third device, comprising: a. a memory for storing communication parameters shared between the third device and a first device and / or between the third device and a second device; b. a communication unit for receiving a first communication flow from the first device and / or a second communication flow from the second device; wherein the communication flow is synchronized based on the communication parameters.
2. The apparatus according to claim 1, wherein the communication parameters include: the distance between the third device and the first device and / or the distance between the third device and the second device; the latency from the first device to the third device and / or the latency from the second device to the third device; or other communication parameters shared between the first device and the third device and / or between the second device and the third device.
3. The apparatus according to claim 1 or 2, further comprising a computing unit for receiving, as an input, the communication parameters between the first device and the third device and for executing a prediction model to predict a control input of the third device, wherein the control input includes predicted communication parameters between the first device and the third device.
4. The apparatus according to any one of claims 1 to 3, further comprising a computing unit for receiving, as an input, at least the communication flow of the first device and for executing a prediction model to predict a control input of the third device.
5. The apparatus according to claim 4, wherein the prediction model is at least one of: a model derived from a general-purpose prediction model that follows at least the parameters shared between the first device and the third device; and a generation model.
6. The apparatus according to any one of claims 1 to 5, wherein the communication parameters are obtained by executing a protocol on the first device.
7. The apparatus according to any one of claims 1 to 6, wherein the communication parameters are set by a management device.
8. The communication parameters include: the latency between the third device and the first device and / or the second device; QoS The distance between the third device and the first device and / or the second device, The computational requirements for processing the communication, The computational capabilities for processing the communication, The memory requirements for processing the communication, The memory capabilities for processing the communication, The available bitrate, The number of communication parties, The (relative) position of the third device with respect to the first device and / or the second device, The (relative) velocity of the third device with respect to the first device and / or the second device, The (relative) acceleration of the third device with respect to the first device and / or the second device, The (relative) rotation of the third device with respect to the first device and / or the second device, The apparatus according to any one of claims 1 to 7, which is at least one of the above.
9. A system comprising at least one third device including the apparatus according to any one of claims 1 to 8, and at least one remote first device for transmitting a communication flow received by the third device.
10. A method for providing synchronized inputs to a third device, comprising: a. Storing in a memory communication parameters shared between the third device and the first device and / or between the third device and the second device; b. Receiving, by a communication unit, a first communication flow from the first device and / or a second communication flow from the second device; c. Synchronizing the communication flow based on the communication parameters. The method includes the above steps.
11. An apparatus for using a prediction model of a first device, comprising: a. A storage unit for storing the prediction model of the first device; b. A communication unit for obtaining communication parameter characteristics of a communication link between the first device and the third device; c. A computing unit capable of using the prediction model of the first device. The apparatus is provided with the above components. The output of the prediction model is obtained based on the prediction model and the communication parameters.
12. The apparatus according to claim 11, wherein the output of the prediction model enables a synchronized communication flow between the first device and the third device.
13. The output of the prediction model is a derived prediction model, and the number of required input parameters of the derived prediction model is less than the number of required input parameters of the prediction model. The apparatus according to claim 11 or 12.
14. A method for using a prediction model of a first device, comprising: a. Saving the prediction model of the first device; b. Obtaining communication parameter characteristics of a communication link between the first device and a third device; c. Using the prediction model of the first device, wherein the output of the prediction model is obtained based on the prediction model and the communication parameters. A method comprising the above steps.
15. A computer program comprising code means for generating the steps of the method according to claim 10 or claim 14 when executed on a computer device.