Data compression with controllable semantic loss
The compression system segments and compresses data into semantic objects, using a shared model adjusted by user preferences to reduce semantic loss and communication overhead in augmented and virtual reality environments, achieving high-quality, personalized content delivery.
Patent Information
- Application Number
- JP2025535916
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-24
- Filing Date
- 2023-12-18
- Publication Date
- 2026-01-21
AI Technical Summary
Current video streaming platforms lack the ability to provide personalized content and maintain consistency in the placement of reproduced objects both in space and time, while also experiencing semantic loss due to data compression in augmented reality and virtual reality environments.
A compression system that segments input data into semantic objects based on criteria, applies compression techniques, and shares a reconstruction model with the receiver, which includes a shared compression model determined by reconstruction performance and user preferences, allowing for personalized data generation and reduced communication overhead.
Enables high-quality interactions with reduced semantic loss and personalized content by optimizing compression at a per-object level, improving spatial and temporal consistency of regenerated scenes.
Smart Images

Figure 2026502123000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a data compression system for data transfer (such as low latency communication between remote users) that can be implemented in, but not limited to, Augmented Reality / Virtual Reality (AR / VR) or Metaverse applications, which can be defined as a virtual reality space in which users can interact with computer-generated environments and / or other users. [Background technology]
[0002] Within the 3GPP® Technical Specification Group Services and System Aspects (TSG SA), the primary objective of 3GPP® TSG SA WG1 (SA1) is to consider and study new and enhanced services, features, and capabilities of 5G systems and identify the corresponding Stage 1 requirements to be met by 3GPP® specifications. These service requirements are documented in normative specifications under the responsibility of SA1. A related study is TR22.847, "Study on supporting tactile and multi-modality communication services (TAMMCS)." This study includes eight use cases and related requirements for the so-called "Tactile Internet" (TI).
[0003] The International Telecommunication Union (ITU) defines TI as an internet network that combines ultra-low latency with extremely high availability, reliability, and security. The mobile internet has made it possible to exchange data and multimedia content even while on the move. The next step is the Internet of Things (IoT), which allows smart devices to be interconnected. TI is the next generation technology that enables real-time control of the IoT. By enabling the sense of touch and tactile sensation, TI adds a new dimension to human-machine interaction and at the same time revolutionizes machine interaction. TI enables humans and machines to interact with their environment in real time while on the move and within a certain spatial communication range.
[0004] IEEE Publication P1918.1, "Tactile Internet: Application Scenarios, Definitions and Terminology, Architecture, Functions, and Technical Assumptions," requires that cellular 5G communication systems support mechanisms to assist synchronization between multiple streams (e.g., haptic, audio, and video) of a multimodal communication session to avoid adverse effects on the user experience. Furthermore, 5G systems will be able to support interactions with applications on user equipment (UE) or data flows that group information within a single haptic and multimodal communication service, and a means to apply third-party policies to flows associated with the application. Policies include sets of UEs and data flows, trigger events associated with expected quality of service (QoS) treatment, and other coordination information.
[0005] In scenarios such as augmented reality or virtual reality (AR / VR), metaverse, or communications, virtual reality spaces where users can interact with computer-generated environments and other users, it is crucial to compress the exchanged representations of the user and the environment so that the rendering of the relevant user representations, environments, and / or virtual spaces is as realistic as possible. However, after the rendering process, there is a problem in that dependencies, concepts, and / or other link information of the relevant user representations, environments, and / or virtual spaces are lost due to compression, which is called semantic loss.
[0006] Furthermore, in scenarios such as video streaming, users receiving a video stream (or other type of data) want to be provided with personalized content so that they can enjoy videos that perfectly suit their preferences.
[0007] However, current video streaming platforms do not allow content personalization; for example, one can select series A, B, C, ..., but there is only one version of series A, B, or C. If series A is selected, all users see the exact same content. Therefore, it is desirable to provide a way to control, trade-off, or optimize semantic loss (certainly not for subsections of the input, such as specific parts of an image), and provide consistency in the placement of reproduced objects both in space (e.g., in images) and time (e.g., in videos or audio). Therefore, it is also desirable to provide a way to provide personalized content. Summary of the Invention [Problem to be solved by the invention]
[0008] The object of the present invention is to solve the above-mentioned desirable characteristics by reducing communication overhead and enabling high quality interactions through personalized data generation. [Means for solving the problem]
[0009] This object is achieved by an apparatus according to claims 1 and 9, a system according to claim 15, a method according to claims 21 and 22, a computer program according to claim 23 and a bitstream according to claim 24.
[0010] According to a first aspect (directed to the compression or encoding side of a compression system, e.g., a transmission device, transmitter, or computing device), there is provided an apparatus configured to identify or segment input data to generate instances related to the identified semantic objects according to one or more criteria, and apply compression techniques to the identified semantic objects. In one example, the input observation data is computer-generated data, e.g., a computer game or video streaming.
[0011] According to a second aspect (directed to a decompression or decoding side of a compression system, e.g., a receiving device, receiver, or computing device), there is provided an apparatus configured to receive compressed semantic objects and apply a decompression technique based on a compression model and a type of the received compressed semantic data object to obtain decompressed data.
[0012] According to a third aspect, a system comprises a transmitting device comprising the apparatus of the first aspect and a receiving device comprising the apparatus of the second aspect, wherein the transmitting device is configured to share a compression model with the receiving device, the shared compression model being determined or updated based on reconstruction performance and / or receiver preferences.
[0013] According to a fourth aspect (directed to the compression or encoding side of a compression system), a method comprises: segmenting the input observation data to generate instances related to the identified semantic objects according to one or more criteria; and applying a compression technique to the identified semantic objects.
[0014] According to a fifth aspect (directed to the decompression or decoding side of a compression system), a method comprises: receiving a compressed semantic object; applying a decompression technique based on a compression model and a type of the received compressed semantic objects to obtain decompressed data, wherein the compression model and / or the compressed semantic objects depend on at least one of a decompression performance requirement, a semantic loss requirement, and a user preference requirement.
[0015] According to a sixth aspect there is provided a computer program product comprising code means for generating the steps of the method of the third or fourth aspect when executed on a computing device.
[0016] According to a seventh aspect, there is provided a bitstream generated by the method of the third aspect and comprising at least one description prompt representing compressed semantic objects belonging to a class of data objects that can be represented or reconstructed by the compression model and the description prompt.
[0017] Therefore, a compression system for semantic loss control and negotiation applicable to generative compression techniques is proposed. The system uses the insight that various objects in an observed scene have significantly different impacts on the overall semantic loss to enable per-object loss optimization. By operating at a per-object level, it is possible to optimize semantic loss with greater control than known systems.
[0018] Thus, data compression can be achieved by classifying input data into one or more types of data objects according to one or more criteria and applying a compression technique to each of the object types to obtain compressed data. The compressed data can then be decompressed by applying a decompression technique to each of the input compressed data object types to obtain decompressed data.
[0019] Additionally, the input data may be analyzed to determine whether at least one selected portion of the input data can be compressed using a first compression scheme according to at least one first criterion, and upon such determination, the at least one selected portion of the input data may be compressed using the first compression scheme and the remaining portion of the input data may be compressed using at least one different compression scheme.
[0020] By incorporating consistency data generated from a comparison of the generated scene with reality before or after transmission or storage, semantic loss can be immediately improved. Furthermore, by using per-subject and global consistency data, the spatial and temporal consistency of the regenerated scene can be improved.
[0021] Semantic losses can be negotiated between the encoder and a third party (e.g., a network function) to jointly optimize bandwidth and realism. The shared reconstruction model can be dynamically updated to account for successfully learned prompts available to both the transmitter (encoder) and receiver (decoder), thereby reducing the bandwidth of future transmissions of the same subject.
[0022] In predictive mode, prompts can be generated in advance of the relevant observation, while still providing control over semantic loss.
[0023] According to a first option, which can be combined with any one of the first to seventh aspects, description prompts are generated for semantic objects belonging to a class of data objects that can be represented or reconstructed with minimal semantic loss by the compression model and description prompts, thereby achieving significant data compression since only description prompts need to be sent instead of the semantic objects.
[0024] According to a second option, which can be combined with the first option or any one of the first to seventh aspects, an image is synthesized using the generated description prompt, instance segmentation is performed on the synthesized image, and at least one of crop parameters, translation parameters, and scale parameters is determined, and the determined parameters are output together with the description prompt, whereby the original image is reconstructed from the synthesized image by using the additional output parameters.
[0025] According to a third option, which can be combined with the first or second option or any one of the first to seventh aspects, the compression technique is based on at least one of the ease of reconstruction of the identified semantic objects, semantic loss requirements, personalization requirements, and compression needs via a compression model, thereby providing a flexible object-oriented compression technique.
[0026] According to a fourth option, which can be combined with any one of the first to third options or any one of the first to seventh aspects, a temporal data association of generated description prompts with frames of input observation data is performed over time, and parameters of a global motion model are determined for each instance associated with the generated description prompt, thereby simplifying the reconstruction of motion associations by using the global motion model.
[0027] According to a fifth option, which can be combined with any one of the first to fourth options or any one of the first to seventh aspects, data objects belonging to a data object type that cannot be represented or reconstructed with minimal semantic loss by description prompts and a compression model are compressed by at least one of generating putative description prompts suitable for preserving sufficient semantic content for a given context, developing description prompts based on text reversal induced by the data object, and using conventional compression techniques, thereby allowing the amount of compression of the remaining, less easily reconstructed regions of the observed scene to be flexibly adjusted depending on the type of data object.
[0028] According to a sixth option, which can be combined with any one of the first to fifth options or any one of the first to seventh aspects, the compressed data objects are labeled according to the compression technique and / or reconstruction model used, which facilitates decompression / decoding at the receiving device.
[0029] According to a seventh option, which can be combined with any one of the first to sixth options or any one of the first to seventh aspects, the observation data comprises at least one of images, video, audio, and haptic information, thereby making the proposed compression scheme applicable to all aspects of metaverse applications.
[0030] According to an eighth option, which may be combined with any one of the first to seventh options or any one of the first to seventh aspects, a sensor is used to acquire the input observation data, and therefore the proposed compression scheme can be applied to any kind of semantic object related to any measurable or detectable scene.
[0031] According to a ninth option, which can be combined with any one of the first to eighth options or any one of the first to seventh aspects, the compression policy is configured by / negotiated with the communication manager, and the compression policy determines at least one of an amount of semantic loss tolerated, a desired compression ratio, a desired computational overhead, a desired storage overhead, and a desired communication overhead. This allows the compression policy to control the selection of a compression method for an identified object type, and enables network-based compression policy control to be implemented. In an example of the eighth option, the decompression technique relies on at least one of a personalized compression model, a personalized policy, and a personalized prompt to obtain personalized decompressed data.
[0032] According to a tenth option, which can be combined with any one of the first to ninth options or any one of the first to seventh aspects, at least one of the compressed semantic objects is decompressed based on a shared reconstruction model and description prompts, so that compression at the receiving end can be achieved by simply transferring a reference to the compression model.
[0033] According to an eleventh option, which can be combined with any one of the first to tenth options or any one of the first to seventh aspects, the (shared) compression model is determined and updated based on reconstruction performance, thereby establishing a feedback loop for controlling the shared compression model in order to optimize reconstruction performance.
[0034] According to a twelfth option, which may be combined with any one of the first to eleventh options or any one of the first to seventh aspects, the compressed model is retrieved from a compressed model repository. Thus, by providing access to the compressed model repository, the flexibility of the proposed object-based compressed model can be increased.
[0035] According to a thirteenth option, which can be combined with any one of the first to twelfth options or any one of the first to seventh aspects, the description prompts are generated predictively, whereby the previous evolution of the observed scene can be evaluated in order to predict object types or movements and predict corresponding description prompts to improve compression performance.
[0036] According to a fourteenth option, which can be combined with any one of the first to thirteenth options or any one of the first to seventh aspects, the decompressed data, which is decompressed based on the predicted description prompts, is compared with the input observation data to determine a correction factor, which can thus be used as a measure of prediction performance.
[0037] According to a fifteenth option, which can be combined with any one of the first to fourteenth options or any one of the first to seventh aspects, the correction factor is used to retrain the shared reconstruction model, thereby allowing the shared reconstruction model to adapt to the compression history in order to improve future compression performance.
[0038] According to a sixteenth option, which may be combined with any one of the first to fifteenth options or any one of the first to seventh options, a correction factor is transmitted to the receiver, and thus the correction factor can be transmitted together with the predicted prompt in order to reduce semantic loss while maintaining good compression performance.
[0039] According to a seventeenth option, which can be combined with any one of the first to sixteenth options or any one of the first to seventh aspects, at least one of the compressed semantic objects is decompressed based on the shared reconstruction model and the description prompts, so that the decompressor only needs a reference to the shared reconstruction model to properly reconstruct the observed scene based on the description prompts.
[0040] According to an eighteenth option that can be combined with any one of the first to seventeenth options or any one of the first to seventh aspects, a decompression policy is provided on the decompression side, so that decompression can be controlled based on the policy.
[0041] According to a nineteenth option, which can be combined with any one of the first to eighteenth options or any one of the first to seventh aspects, the decompression policy is negotiated with or configured by the communication manager, whereby the decompression performance can be controlled from the network side.
[0042] According to a twentieth option, which may be combined with any one of the first to nineteenth options or any one of the first to seventh aspects, the compression model is shared with the transmitting device, so that compression at the receiving side can be achieved by simply transferring a reference to the compression model.
[0043] According to a twenty-first option, which can be combined with any one of the first to twentieth options or any one of the first to seventh aspects, the received description prompt is used for predictive decompression, whereby previous decompressions can be evaluated to accelerate the decompression process.
[0044] According to a 22nd option, which can be combined with any one of the 1st to 21st options or any one of the 1st to 7th aspects, the predictive decompression is based on communication parameters such as delay with the transmitter, thus allowing the decompression to be adapted to the communication performance.
[0045] According to a 23rd option, which can be combined with any one of the 1st to 22nd options or any one of the 1st to 7th aspects, predicted decompressed data is compared with the obtained decompressed data to determine a correction factor, which allows the predicted decompression performance to be measured and real-time corrections to be made.
[0046] According to a 24th option, which can be combined with any one of the 1st to 23rd options or any one of the 1st to 7th aspects, a correction factor is transmitted to a transmitter that receives the compressed semantic object, so that a correction can be performed at the transmitter side to achieve a better match with the expected decompression.
[0047] According to a 25th option, which can be combined with any one of the 1st to 24th options or any one of the 1st to 7th aspects, a correction factor is received from a transmitter that has received the compressed semantic object, and the correction factor is used to correct the obtained decompressed data, so that the correction can be performed at the receiving side and a better match with the decompression can be achieved.
[0048] According to a 26th option, which can be combined with any one of the 1st to 25th options or any one of the 1st to 7th aspects, an image is synthesized using the generated description prompt, instance segmentation is performed on the synthesized image, and at least one of crop parameters, translation parameters, and scale parameters is determined, and the determined parameters are output together with the description prompt, thereby enabling processing of multiple objects.
[0049] According to a 27th option, which can be combined with any one of the 1 to 26 options or any one of the 1 to 7 aspects, a temporal data association of the generated description prompts with frames of input observation data is performed over time, and parameters of a global motion model are determined for each instance associated with the generated description prompt, thus taking into account movement of a camera or other sensor device during scene observation.
[0050] According to a 28th option, which can be combined with any one of the 1 to 27 options or any one of the 1 to 7 aspects, the semantic loss is determined based on an instance rate distortion function, and the total loss is summed or weighted summed over multiple semantic objects identified in the input observation data, and the object loss is composed of the object rate and object distortion, which depends on the object color, object shape, and texture parameters. Thus, the instance rate distortion for each object can be taken into account. The reason for the weighted sum is that the distortion factors for the objects have different importance, for example, the object color is less important than the object shape.
[0051] According to a 29th option, which can be combined with any one of the 1st to 28th options or any one of the 1st to 7th aspects, a receiving device receives a personalized compressed (decompressed) model, and the receiving device receives the personalized model by (1) notifying the transmitting device of its preferences or (2) the transmitting device determining that the model is compatible with the receiving device's preferences.
[0052] According to a 30th option, which may be combined with any one of the 1st to 29th options or any one of the 1st to 7th aspects, a sending device is capable of creating and / or editing at least one content based on a compression model and an interpreted programming language, and sharing said content with at least one receiving device. It is understood that the apparatus according to claims 1 and 9, the system according to claim 15, the method according to claims 21 and 22, the computer program according to claim 23, and the bitstream according to claim 24 have similar and / or identical embodiments, in particular as defined in the dependent claims.
[0053] It is understood that a preferred embodiment of the invention can also be any combination of the dependent claims or the above embodiments with the respective independent claim.
[0054] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. [Brief explanation of the drawings]
[0055] [Figure 1A] FIG. 1 illustrates a block diagram of an alternative network architecture for a metaverse implementation in a cellular network. [Figure 1B] FIG. 1 illustrates a block diagram of an alternative network architecture for a metaverse implementation in a cellular network. [Figure 2]FIG. 1 is a schematic diagram illustrating a state diagram representing the operation of an exemplary system implementing a metaverse application. [Figure 3] FIG. 1 is a diagram illustrating an exemplary metaverse implementation. [Figure 4] FIG. 1 illustrates a schematic block diagram of a network architecture for implementing various embodiments. [Figure 5] FIG. 1 illustrates a schematic block diagram of different layers involved in a compression system, according to various embodiments. [Figure 6] 6A-6C are schematic diagrams of flow diagrams of compression and decompression processes according to various embodiments using the different layers of FIG. 5. [Figure 7] 1A-1C are schematic illustrations of flow diagrams of various processes involved in compression and decompression processes, according to various embodiments. [Figure 8] FIG. 2 shows a schematic diagram of the processing steps and outputs of an exemplary embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0056] Next, an embodiment of the present invention will be described based on a cellular communication network environment such as 5G. However, the present invention can also be used in combination with other wireless technologies in which TI or metaverse applications are provided or can be implemented. The present invention can also be applied to other applications such as video streaming services, video broadcasting services, or data storage.
[0057] Throughout this disclosure, the abbreviations "gNB" (5G terminology) or "BS" (base station) are intended to mean a wireless access device such as a cellular base station, a Wi-Fi access point, or an ultra-wideband (UWB) personal area network (PAN) coordinator. The gNB consists of a centralized control plane unit (gNB-CU-CP), multiple centralized user plane units (gNB-CU-UP), and / or multiple distributed units (gNB-DU). The gNB is part of the radio access network (RAN) and provides an interface to functions in the core network (CN). The RAN is part of a wireless communication network and implements a radio access technology (RAT). Conceptually, it resides between a communication device, such as a mobile phone, a computer, or a remotely controlled machine, and the CN, providing connectivity to the CN. The CN is the core part of the communication network, providing various services to customers interconnected via the RAN. More specifically, it transmits communication streams through the communication network and, possibly, other networks.
[0058] Furthermore, in this disclosure, the terms "base station" (BS) and "network" are used synonymously. This means, for example, that when a "network" is said to perform a particular operation, the operation is performed by a CN function of a wireless communications network or by one or more base stations that are part of such a wireless communications network, and vice versa. It also means that some of the functions are performed by a CN function of a wireless communications network and some of the functions are performed by a base station.
[0059] Furthermore, the term "metaverse" is understood to refer to a persistent, shared set of interactive spaces in which users interact with each other with mutually perceived virtual features (i.e., augmented reality (AR)) or where those spaces are composed entirely of virtual features (i.e., virtual reality (VR)). VR and AR are commonly referred to as "mixed reality" (MR).
[0060] Furthermore, the term "data" is understood to refer to a representation in a known or agreed-upon format of information that is subject to storage, transmission, or other processing. The information may comprise, among other things, one or more channels of synchronized audio, video, image, haptic, motion, or other forms of multimedia information. Such multimedia information may be obtained from sensors (e.g., microphones, cameras, motion detectors, etc.) or may be partially or wholly synthesized (e.g., live actors in front of synthetic backgrounds).
[0061] The term "data object" refers to a set of one or more data according to the definition above, optionally accompanied by one or more data descriptors that provide additional semantic information about the data that influences how the data is processed at the sender and receiver. Data descriptors are used, for example, to describe how the data is classified by a sender and rendered by a receiver. For example, data representing an image or video sequence may be divided into a set of data objects that collectively describe the complete image or video, and each may be processed (e.g., compressed) in a manner that is optimal for the object and its semantic context, substantially independent of other data objects. As a further example, a content program (described below in some embodiments) may also be understood as a data object (e.g., a compressed semantic object).
[0062] Furthermore, the terms "data object classification" or "data object identification" are understood to refer to a process in which data is divided or segmented into multiple data objects, in other words, a process in which (semantic) data objects are identified. For example, an image may be divided into multiple parts, such as a forest in the background and a person in the foreground (e.g., as illustrated below in connection with FIG. 8). Data object classification criteria are used to classify the data objects. In the present disclosure, such criteria include at least one of a measure of the semantic content of the data object, the context of the data object, the type of data object, the class of compression technique optimal for preserving sufficient semantic content for a given context, etc. For example, an AI / ML (artificial intelligence / machine learning) model is used to determine all data objects representing data objects in a drawing, e.g., cats.
[0063] Furthermore, "compression techniques" are understood to refer to methods that reduce the size of data to make its transmission or storage more efficient, for example by removing redundant data or data that is deemed semantically irrelevant to an end user, and efficiently encoding the remaining data so that a faithful, or nearly semantically faithful, representation of the original data can be reconstructed.
[0064] Furthermore, a "compression or reconstruction model" is understood to refer to a repository of tools and data objects that can be used to assist in the compression and reconstruction of data. For example, the model comprises algorithms used to analyze and compress data objects, or comprises data objects that can be used as the basis for a generative compression technique. Advantageously, the model is shared or owned by the sender and receiver, and / or is updated or optimized depending on the semantic content of the data being transferred.
[0065] It should be noted that throughout this disclosure, the accompanying drawings only show blocks, components, and / or devices related to the proposed data distribution functionality. Other blocks have been omitted for the sake of brevity. Furthermore, blocks designated by the same reference numerals are intended to have the same or at least similar functionality, and therefore, their functionality will not be described again later.
[0066] 1A and 1B are schematic diagrams of a network architecture (e.g., IEEE P1918.1 architecture) that is being considered for implementing the metaverse. The architecture includes an Actuator Gateway Interface (AG), an Actuator Node (AN), a Controller Node (CN), a Control Plane Entity (CPE), a Gateway Node (GN, GNC corresponds to GN and CN), a Human System Interface Node (HN), a Network Controller (NC), a Sensor / Actuator (S / A), a Computing and Storage Entity (SE), a Sensor Gateway (SG), a Sensor Node (SN), a Haptic Device (TD), a Haptic Edge (TE), a Haptic Service Manager (TSM), a User Plane Entity (UPE), an Access Interface (A), a first Haptic Interface Ta (for inter-TD communication), a second Haptic Interface Tb (for TD-GNC communication), an Open Interface (O), a Service Interface (S), a Network Side (N), a Network Domain (ND), a Bidirectional Information Exchange (BIE), an External Application Service Provider (EASP), and a Dedicated Low Latency Network (LLNW).
[0067] The architectures in Figures 1A and 1B provide a generalized overall communication architecture that can operate over any network, including 5G. It covers various modes of interconnection network domains between two TEs (TE A, TE B). Each TE is composed of one or more TDs. The TD in TE A communicates information such as tactile / haptic information with the TD in TE B through an ND to meet the requirements of a given TI use case. The ND can be a shared wireless network (e.g., a 5G radio access and core network), a shared wired network (e.g., an Internet core network), a dedicated wireless network (e.g., a point-to-point microwave or millimeter wave link), or a dedicated wired network (e.g., a point-to-point leased line or optical fiber link). Each TD can support one or more functions of sensing, actuation, haptic feedback, or control via one or more corresponding entities. The S or A entities refer to devices that perform sensing or actuation functions, respectively, without a network module. SN or AN refer to devices that perform sensing or actuation functions, respectively, using air interface network connectivity modules. To connect S to SN or A to AN, a SG or AG entity must be used, respectively. These gateways provide a generic interface for connecting to third-party sensing and actuation devices and another interface for connecting to SN and AN. TD can function as either a HN that can translate human input into haptic output, or a CN that uses the necessary network connectivity modules to execute control algorithms to process the operation of the SN and AN systems.
[0068] The Network Network (GN) is an entity with enhanced network functions that resides at the interface between the TE and the Network Distributed Network (ND) and is primarily responsible for user-plane data transfer. The GN is accompanied by a Network Controller (NC) responsible for control plane processing, including admission and congestion control, service provisioning, resource management and optimization, and connection management intelligence, to achieve the required QoS for TI sessions. The GN and CN (collectively labeled GNC) can reside either on the TE side (shown in Figure 1A) or the ND side (shown in Figure 1B), depending on the network design and configuration. The GNC is a central node that facilitates interoperability with various possible network domain options and is essential for compatibility with other emerging standards, such as the 3GPP 5G NR specification. Residing the GNC within the ND under 5G, for example, is intended to support the option of absorbing the GNC's functions into existing management and orchestration functions. In Figures 1A and 1B, the ND is shown to consist of a wireless access point or base station logically connected to the Customer Presence Equipment (CPE) and User Presence Equipment (UPE) in the network core.
[0069] A user within a region of interest (ROI) is surrounded by a set of TDs linked to a TE. A TD may comprise rendering actuators and / or sensors. Rendering actuators have the task of creating a metaverse environment around the user, and examples include VR glasses, 3D televisions (TVs), and holographic devices. Sensors TDs are devices responsible for capturing the user's actions and / or the environment, and include video cameras, audio devices such as microphones, tactile sensors, etc. Generally, in a 5G system, a TD is a UE.
[0070] The TDs within the ROI are connected to the user's TEs, for example, wired or wirelessly. In the wireless case, the UEs are connected to base stations such as 5G gNBs or Wi-Fi access points. The TE's network infrastructure and computational resources are either co-located within the ROI or located at nearby edge servers (at a distance shorter than the maximum edge distance) to ensure fast response.
[0071] To support the implementation of the following embodiments, at least one of three communication functions is introduced. First, a delay-based flow synchronization (LBFS) function is provided, which is a function executed in a device within a TE. It can also be deployed in a receiving TD that can determine communication parameters with a transmitting TD (or TS) and synchronize communication flows based on those communication parameters, particularly the relative delay between the TDs (or TEs). Second, an edge application in a TE is configured to execute a delay-dependent configurable predictive model (LDCPM) of an environment / person within a metaverse session in another TE. Third, a model management and configuration function is provided that can register generic models of ROIs, devices, and / or people in the TE, store them in a database, and deploy reconfigured LDCPMs when determining communication parameters.
[0072] Another pioneering node in the architectures of Figures 1A and 1B is the SE, which provides both computing and storage resources to improve the performance of the TE and meet the latency and reliability requirements of E2E communication. The SE executes advanced algorithms, employing AI techniques, among others, to offload resource- and / or energy-intensive processing operations (e.g., haptic rendering, motion trajectory prediction, and sensory compensation) to the TD. The goal is to overcome challenges and uncertainties along the path between the source and destination TDs while using predictive analytics to become aware of real-time connectivity, dynamically estimate network load and speed changes over time to optimize resource utilization, and enable learning experiences about the environment to be shared between different TDs. Meanwhile, the SE also provides intelligent caching functionality, which is highly effective in reducing E2E traffic load and data transfer latency. The SE can reside locally within the TE to improve response speed to requests from the TD or GNC, and / or can reside remotely in the cloud, providing services to the TE and ND. Furthermore, the SE can be either centralized or distributed. Each of these options has advantages and disadvantages in terms of latency, reliability, capacity, cost, and practicality. Communication between two TEs can be unidirectional or bidirectional, can be based on a client-server model or a peer-to-peer model, and can belong to any of the above use cases with the corresponding reliability and latency requirements. To this end, the TSM plays a key role in defining the characteristics and requirements of the service between two TEs and distributing this information to key nodes in the TEs and NDs. The TSM also supports registration and authentication functions and provides an interface to the TI's EASP.
[0073] The A interface provides connectivity between the TE and the ND. It is the primary reference point for user and control plane information exchange between the ND and the TE. Depending on the architecture design, the A interface can be located between the TD and the ND or between the GNC and the ND. Furthermore, the T interface provides connectivity between entities within the TE. It is the primary reference point for user and control plane information exchange between entities in the TE. The T interface is divided into two subinterfaces, Ta and Tb, to support different modes of TD connectivity. Thus, the Ta interface is used for TD-to-TD communication, and the Tb interface is used for communication between the TD and the GNC when the GNC resides within the TE. Furthermore, the O interface provides connectivity between any architecture entity and the SE, and the S interface provides connectivity between the TSM and the GNC. The S interface carries control plane information. Finally, the N interface refers to any interface that provides internal connectivity between ND entities. This is typically covered as part of a network domain standard and can include subinterfaces for both user and control plane entities.
[0074] Tactile information is embodied in two broad categories, haptic or kinesthetic, which are combined. Haptic information refers to the perception of information by the various mechanoreceptors in human skin, such as surface texture, friction, and temperature. Kinesthetic information refers to information perceived by the human body's skeleton, muscles, and tendons, such as force, torque, position, and velocity.
[0075] A first differentiating aspect of TI and related standards compared to 5G Ultra-Reliable Low Latency Communications (URLLC, ITU-R M.2083) is that 1-millisecond propagation times require TI to be developed to support requirements over long distances of 150 km round trip (or 100 km for fiber). Such capabilities can be achieved through network-side support capabilities built into the TI architecture, as envisioned through the standardization work in IEEE 1918.1. These capabilities can, for example, use artificial intelligence (AI) techniques to model the remote environment and, in some cases, may reside partially or entirely in the TI end device (i.e., the TI / haptic client).
[0076] The second differentiating aspect is that TI creates applications with unique characteristics implied by the application, and the expectation is that the application can be deployed as an overlay network on top of a network or combination of networks, which is not intended to apply only in the context of 5G URLLC as the underlying communication medium.
[0077] 1A and 1B, data streams, e.g., haptic feedback, also need to be synchronized; users expect to "feel" or "experience" a visually depicted event when it occurs, regardless of whether the event is audible or not. Therefore, synchronization of audio, video, and haptic data is very important. Incidentally, this can be achieved by buffering at the receiver, thereby completely eliminating the challenges faced by communication networks in achieving the required delay (e.g., jitter).
[0078] To meet stringent end-to-end (E2E) QoS requirements, the architecture must also provide advanced operational and management capabilities, such as lightweight signaling protocols, caching with distributed computing and predictive analytics, intelligent adaptation based on load and network conditions, and integration with external application service providers (ASPs).
[0079] As a result, reliability, latency, and scalability are considered as key performance indicators (KPIs), and omnipresence (rapid "latching" of TDs to the TI infrastructure), ad hoc (minimal maintenance of the TI network domain), and hybrid (scalable yet minimal maintenance of TI rendezvous devices) are considered as the three main approaches for bootstrapping TI services and instantiating the architecture. The TI design is essentially built around the concept of mandatory operational configuration at the edge and E2E maintenance managed in the core. That is, TDs at the edge specify and communicate communication and operational parameters (e.g., expected latency and reliability) to the TI architecture, which allocates the necessary resources to meet those requirements, both for bootstrap setup and E2E communication.
[0080] FIG. 2 shows a schematic state diagram illustrating the operational states of the overall operational finite state machine for implementing the metaverse application.
[0081] A TD device begins with the registration phase (REG), which is defined as the operation of establishing communication with the TI architecture. In the ubiquitous TI paradigm, registration is performed with the GNC and may involve TI components such as the ND to the TSM. Selected applications may provide a user interface for the user / ROI to register with, for example, the TSM. Registration may be achieved through the application itself or through the TSM and involves registering the TDs that the user / ROI has. This involves registering the TD as part of the communication infrastructure, for example, as part of the 5G system (5GS), to gain access to features provided by 5GS, such as quality of service (QoS), low latency, edge computing, or synchronization of communication flows.
[0082] Upon registration, the TSM assigns TEs to users, ROIs, and / or TDs that are in close proximity or have suitable communication parameters. In some cases, a sensor TD generates an output that is fed to a rendering TD within the same ROI.
[0083] The "latch" point for a TD to initiate registration is called the TI anchor. At this stage, the TD is probing the TI architecture to invoke E2E communication and performs no other function than latching into the TI architecture. In both the ad-hoc and hybrid models, this step involves the TSM, and in the former case via the GNC to establish registration.
[0084] The next state depends on the type of TD. In the case of a subordinate SN / AN, the TD has a designated "parent" in the vicinity, and the TD must first associate (ASS) with the parent. This parent TI node then ensures reliable operation and assists with connection establishment and error recovery. This is an optional step if the TD device operates independently. Some mission-critical TDs, as well as new TDs, must be authenticated (AUT) without a parent (Ap) before being allowed to join / initiate a TI session (SS).
[0085] The other phase is an optional state in which the TD (NATD) communicates with an authentication agent within the TI infrastructure to perform authentication. The TSM is an entity that can perform this task, with assistance from the SE if necessary and handling large amounts of traffic. The TD then initiates E2E control synchronization (Ctrl Sync) to probe and establish links to the end TEs. In this state, the TD cannot communicate operational data but instead focuses on relaying connection setup and maintenance parameters. This includes configuring the parameters of interfaces along the E2E path, which allows the ND to select the optimal path across the network to deliver the requested connection parameters. This state encompasses the path establishment and route selection phase of TI operation. Typically, multiple layers of the TI architecture are involved, communicating to ensure that a path that meets the minimum requirements established in the "setup" message is actually available and reserved.
[0086] If the TD participating in a TI session is a Haptic Node (HN) intended for haptic communication, the next state encompasses the specific communication and establishment of haptic-specific information prior to the actual data communication. This state involves determining the codec, session parameters, and messaging format specific to the current TI session. While different use cases require different haptic exchange frequencies, it is expected that all haptic communication will begin in the Haptic Synchronization state (H-Sync) to establish initial parameters. Future changes to codec and other haptic parameters will be handled as data communication in the "Operation" state (OP). This ensures that the initial settings are enforced for all haptic communication, regardless of future updates to parameters contained in the operational data payload.
[0087] All TD components then transition to the operational state, in which the E2E path is established, all connection setup requirements are met, and the TEs are ready to exchange TI information. If, while operating in this state, one TD detects an intermittent network error (ERR), the TD transitions to a "recovery" mode (REC) in an attempt to re-establish reliable communications, where a specified protocol takes over error checking and potential correction mechanisms. If the error is found to be intermittent and is resolved, the TD transitions back to the operational state. If, for any reason, the error persists, the TD transitions back to control synchronization and re-detects whether an E2E path is actually available according to the operational requirements set by the edge user.
[0088] Finally, once the TI operation has completed successfully, the TD moves into the "Termination" phase (TERM), in which all resources previously dedicated to this TD are released to the TI management plane. If originally handled by the NC, the resources are returned to the NC. Typically, the TSM is involved in the provisioning of TI resources.
[0089] Figure 3 shows the P A and P B 1 illustrates a schematic diagram of an exemplary metaverse scenario in which two people, P and P, intend to interact within the metaverse. To this end, the two people have corresponding rendering devices RD A and RD B, e.g., virtual reality (VR) devices, and corresponding sensor devices SD A and SD B. A and P B are separated by a distance d. Due to the requirement for high-quality data representation, sensor devices SD A and SD B need to sample high-quality representations of the person to be sent to other person rendering devices. Therefore, high-quality communication with reduced communication overhead is required.
[0090] In the following, embodiments are presented that allow for high quality communication or interaction while achieving communication overhead through compression.
[0091] FIG. 4 shows a schematic block diagram of a network architecture for implementing some of the embodiments.
[0092] A transmitter (Tx) 22 is understood to be a device (e.g., a 5G UE) that senses or generates data to be compressed. The data or compressed data is transferred to a network (NW) 20 via an access link (AL), e.g., a 5G New Radio (NR) radio link. Furthermore, a receiver (Rx) 28 (e.g., a 5G UE) is understood to be a device that renders the data or compressed data. The data or compressed data is transferred from the network 20 via an access link (AL), e.g., a 5G NR radio link.
[0093] Further provided is a network (NW) 20, which is understood to be any type of arrangement or entity used to transport, store, process, and / or otherwise handle data or compressed data. The network 20 comprises multiple logical components distributed across multiple physical devices. In an embodiment, network edge servers 24, 29 and a core network (CN) 21 are provided.
[0094] Network edge servers 24, 29 are understood to be devices that are physically close to a radio access network (not shown) and provide data processing and storage services to user devices (e.g., UEs) involved in an interaction or communication. Physical proximity ensures very low communication latency with the user devices. In an exemplary application, a transmit edge server (Tx ES) 24 and a receive edge server (Rx ES) 29 are provided at the transmitter 22 and receiver 28, respectively, and are configured to provide data storage and compression / decompression databases, compression / decompression functions, and negotiation functions for negotiating with peer devices.
[0095] Furthermore, the core network 21 is understood to comprise the remainder of the network 20 used to transfer data between the transmitter 22 and the receiver 28, possibly via respective edge servers 24, 29.
[0096] Additionally, shared storage (S) 23 is provided as a virtual device representing memory shared by both the sender 22 and receiver 28 and / or the compression and decompression functions of the edge servers 24, 29. It can be physically located in one or more locations, with multiple copies synchronized as needed.
[0097] Communication parameters are understood to be parameters that affect the performance of a communication or that impose requirements on the communication. These include at least one of delay, QoS, distance between communicating parties, computational requirements for processing the communication, computational power for processing the communication, memory requirements for processing the communication, memory power for processing the communication, available bit rate, number of communicating parties, etc. Some of these parameters are interrelated. For example, the delay between two user devices depends on the distance between the devices, but also on other aspects such as computational requirements and / or power for processing the communication. In particular, if the communication involves a predictive model, the communication delay is affected by the available / required computational power of both devices.
[0098] In some embodiments, without loss of generality, only some of the above communication parameters are mentioned, for example, when an embodiment refers only to delay, this should be understood as delay or other communication parameters, in particular other communication parameters that affect the delay of a communication link.
[0099] In embodiments, data compression can be used to provide sufficient image quality (e.g., realistic depictions of people) at the receiving end. Such data compression includes conventional data compression schemes. Furthermore, the following embodiments describe specific data compression schemes and corresponding devices or systems for compressing and decompressing data. It should be noted that while these embodiments are beneficial for the specific applications mentioned herein, they may also be implemented independently in other contexts and applications outside of the metaverse, for example.
[0100] The embodiment relates to a first type of system in which a compression device aims to take input data and efficiently store it in a storage unit. This is useful in cloud-based environments where several pieces of data need to be stored efficiently. Here, a compression device or encoder is used as a device (software product) that performs data compression. For example, the device can receive data, decompose (classify) the data into one or more data objects according to appropriate criteria, then re-compress the data objects according to appropriate criteria, and store the compressed data objects together with a compression model and other metadata necessary to later reconstruct the source data. Suitable criteria include considerations such as the required (semantic) accuracy of the reconstruction, the storage space and processing resources available for the reconstruction, etc. For purposes of this disclosure, the criteria also include consideration of the semantic content of the data and data objects, and the compression model used.
[0101] Optionally, a decompression device (or decoder) is provided, which is understood to be a device (software product) that retrieves the compressed data from storage and, using an appropriate compression (decompression) model, reconstructs and renders the original data with minimal semantic loss. To this end, a compression model repository is used, which is understood to be a database comprising the tools and data objects used for data compression, advantageously available to both the compression and decompression devices. A subset of the tools and / or data objects of the repository are combined to form a compression model that is optimized in some way for a given sample or type of data.
[0102] A storage unit is understood to be a device accessible by a compression device for storage and by a decompression device for retrieval, that stores compressed data and associated compression models and metadata, and also holds a compression model repository. Storage units can take many forms, such as a web-based server or a physical device such as a memory stick, hard drive, or optical disk.
[0103] Other embodiments relate to a second type of system in which a compressing / transmitting device aims to exchange data with a second decompressing / receiving device in an efficient manner. This is useful in settings where a transmitting device (e.g., a streaming service cloud) wants to efficiently share data with a receiving device (e.g., a television). In some cases, the transmitting device (e.g., transmitter 22 in FIG. 4 ) has sensors capable of sensing / capturing data, such as VR glasses, a mobile terminal, or a user device. In some cases, the receiving device (e.g., receiver 28 in FIG. 4 ) has sensors capable of playing / rendering data, such as VR glasses, a mobile terminal, or a user device. In some cases, an edge server (e.g., edge servers 24, 29 in FIG. 4 ) is associated with one of the transmitting / receiving devices and takes over and / or shares some of the functions. In some cases, some of the transmitting / receiving devices are part of a telecommunications system, such as a 5G system. In relation to other embodiments of the second type of system, a compressed sending device is generally understood to be a compression device that compresses data in a substantially streaming manner and delivers the compressed data to a transmission channel, taking into account delay, computational overhead, or communication overhead at either the sending device / receiving device as part of its compression criteria. Furthermore, a decompressing receiving device is understood to be a decompressing device that decompresses data arriving at the transmission channel, typically rendering it in real time, and typically taking into account delay, computational overhead, or communication overhead as part of its rendering criteria.
[0104] An (optional) edge server may be provided adjacent to the compressed sending device to assist the compressed sending device with compression (post)processing, providing or updating compression models on the fly, ensuring timely delivery of compressed data, etc. Similarly, an (optional) edge server may be provided adjacent to the decompression receiving device to assist the decompression receiving device with decompression (pre)processing, providing or updating compression models on the fly, ensuring timely rendering of decompressed data, etc.
[0105] Additionally, (optional) sensors (e.g., audio, video) are provided in the compression transmission device, typically configured as a device or array of devices capturing some aspect of the scene, although not necessarily in real time. Examples include cameras, microphones, motion sensors, etc. Some devices capture stimuli outside the range of human sensations (e.g., infrared cameras, ultrasonic microphones) and "down-convert" them to a form that humans can perceive. Some devices comprise an array of sensor elements that provide an extended or more detailed impression of the environment (e.g., multiple cameras capturing a 360-degree perspective, multiple microphones capturing a stereo or surround sound field). Sensors with different modalities are used together (e.g., sound and video). In such cases, the different data streams need to be synchronized. The compression transmission device equipped with sensors is the VR / AR glasses, or simply the UE.
[0106] Additionally, an (optional) rendering device (e.g., audio, video) is provided at the uncompressed receiving device; it is typically a device or array of devices that renders some aspect of the scene in real time. Examples include a video display or projector, headphones, loudspeakers, haptic transducers, etc. Some rendering devices comprise an array of rendering elements that provide an enhanced or more detailed impression of the captured scene (e.g., multiple video monitors, loudspeaker arrays for rendering stereo or surround sound audio). Rendering devices with different modalities are used together (e.g., sound and video). In these cases, the rendering subsystem must ensure that all stimulus channels are rendered synchronously.
[0107] Furthermore, the second type of system is provided with an (optional) communication manager, which is a centralized or decentralized entity that manages the communication. The goal of the communication manager is to optimize the communication in terms of delay, overhead, etc. The communication manager may be an entity within the communication network, such as 5GS, or an external entity, such as an application function.
[0108] As a further (optional) element of the second type of system, a condensed model repository contains information such as the data, the machine learning (ML) models used to derive the data, and serves to reconstruct the data based on prompts, for example.
[0109] In the following embodiments (at least some of which are combined to further improve performance), data compression and reconstruction are based on prompts from a text-to-image model (such as a latent diffusion model) that can learn on the fly to represent previously unseen objects without having to retrain the entire reconstruction model. This technique ("text inversion") can be performed quickly and iteratively, as described in "An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion" by Rinon Gal et al. (available from https: / / textual-inversion.github.io / ).
[0110] Furthermore, in embodiments, data compression and reconstruction are based on a system in which text-inversion based generative compression is guided by the input image and the reconstruction maintains a similarity to that image, as described in "EXTREME GENERATIVE IMAGE COMPRESSION BY LEARNING TEXT EMBEDDING FROM DIFFUSION MODELS" by Zhihong Pan et al., available at https: / / arxiv.org / pdf / 2211.07793.pdf.
[0111] Furthermore, in embodiments, image classification is performed quickly even on low-capability devices, as described in "Image Classification on IoT Edge Devices: Profiling and Modeling" by Salma Abdel Magid et al., available at https: / / arxiv.org / pdf / 1902.11119.pdf.
[0112] Diffusion models are typically good at recalling items that form part of the training dataset, but poor at recalling items that are not. Because the training dataset is obtained from the public internet and retraining a diffusion model is costly, they have problems reproducing inputs that are not known in advance.
[0113] For example, a diffusion model can conjure up a general image of a person from a prompt such as "young man," but cannot conjure up an image of a specific person (with some exceptions, such as celebrities).
[0114] This loss of realism between the reproduced output and the original observation is called "semantic loss" and differs in several ways from the distortions introduced by traditional codecs, notably in that it depends heavily on the original subject being observed.
[0115] In addition to semantic loss, spatial and temporal stability may also be an issue. In a related technique (Neural Radial Fields, or NERF), recent research has explored techniques to improve spatial and temporal stability.
[0116] Recent techniques (e.g., those introduced above) have attempted to address the problem of semantic loss. The state-of-the-art in this area is represented by text inversion, i.e., dynamically learning embeddings that represent previously unseen objects without retraining the entire diffusion model. Thus, a guide image is used to ensure that the learned embeddings adequately represent the observed reality (represented by the guide image).
[0117] In the following embodiments, such learned embeddings are referred to as "learned prompts."
[0118] Furthermore, inpainting is the process of replacing one object with another in the reconstruction generated by the diffusion model. Recent techniques have demonstrated exemplar-based image editing, which allows a given image (an exemplar) to be inpainted onto an existing reconstructed image without introducing fusion artifacts.
[0119] Apart from the above, instance segmentation is a machine learning workload. For example, image segmentation can be used to identify and classify objects in images. Such techniques can be performed quickly on modest hardware.
[0120] Additionally, the saliency of objects in an image can be calculated to consider the semantics of objects in an image. In the context of visual processing, saliency refers to the intrinsic characteristics of an image (e.g., pixels, resolution, etc.). These intrinsic characteristics describe visually appealing parts of an image. A saliency map is their topographical representation. Saliency typically arises from the contrast between an item and its surrounding area. It can be represented, for example, by a red dot surrounded by white dots, a blinking message indicator on an answering machine, or a loud noise in a quiet environment.
[0121] According to an embodiment, a device such as a UE uses generative compression based on the device's local generation of inputs ("learned prompts") to a shared reconstruction model and comparison with an observed scene segmented for each instance, to achieve a controllable degree of semantic loss within bandwidth, computation, and latency objectives (negotiated with network capabilities).
[0122] The sending device observes a scene using some form of sensing (which may include audio, video, image capture, etc.) and then segments its observations to generate instances related to identified semantic objects. It then classifies these objects according to their ease of reconstruction by the shared reconstruction model, deriving portions of the input that are likely to be properly reconstructed (i.e., because they are well represented in the training dataset) and portions that are not. For portions of the scene that can be properly reconstructed, the sending device directly generates the appropriate learned prompts.
[0123] For portions of a given scene that cannot be easily reconstructed by the shared reconstruction model, the sending device performs at least one of several passes. For example, (i) sending learned best-guess prompts, (ii) using more sophisticated and demanding techniques (such as guided images), or (iii) reverting to sending only "physical" input (i.e., unprompted, compressed data compressed using conventional techniques), e.g., in descending order of semantic loss. The receiving device then uses inpainting (or a similar technique) to piece together a reconstructed scene from a combination of the reconstructed, learned prompts and any physical input. Additional consistency data generated by the sending device helps to preserve spatial, temporal, and other forms of consistency in the reconstruction. Furthermore, the reconstruction serves as a predictor for at least part of the image, and a conventional video encoder is used to compress the residual image, whereby rate-distortion optimization is guided by a saliency map, allocating more bits to salient regions (e.g., faces) and leaving realistic but inaccurate data for non-salient image regions (e.g., background). Furthermore, the reconstruction serves as one of the reference frames used to encode the frame by conventional video coders, again with the help of a saliency map to guide the rate-distortion optimization.
[0124] Therefore, a system for generative compression based on learned prompts is proposed. In this system, segmentation is performed by classifying the input according to its "ease of reconstruction" using a reconstruction model applied before transmission, and an estimated semantic loss is derived. This is performed for multiple semantic regions and / or features in the input, deriving an estimated semantic loss for each region (e.g., by instance segmentation in an image). The "ease of reconstruction" calculation is dynamically updated with the dialogue, taking into account new learned prompts available on both the sending and receiving devices.
[0125] Furthermore, semantic loss negotiation is applied by having the transmitting device (encoder) decide which input portions to encode / transmit using prompts and which input portions to encode / transmit conventionally (or otherwise) based on negotiation with an external party (e.g., network functionality). The negotiation is bidirectional (i.e., the transmitting device requests the bandwidth necessary to achieve an acceptable semantic loss and feeds back requests based on the observed scene, since this tradeoff changes dynamically).
[0126] Based on a given rate-distortion equation, L(k;p) = kR(p) + D(p), where R(p) is the bitrate for a given set of parameters p, D is the distortion measured by comparing the reconstructed signal with an uncompressed reference signal, and k controls the importance of bitrate vs. quality, the loss function L(k;p) is minimized.
[0127] For example, based on configured policy, the negotiation is one-way.
[0128] Finally, the learned prompts are compared to reality before being sent. To allow for selection, reality is observed before sending the prompts. The comparison is made instead of or in addition to a predicted version of reality, and optional modifications to the pre-computed prompts are sent after the actual reality is observed.
[0129] Spatial or temporal stability can be achieved by generating additional data at the transmitting device based on knowledge of the reconstruction model and observations of the scene, improving the spatial or temporal consistency of the reconstruction.
[0130] The contributions of the diffusion codec and the conventional video codec to the overall video coding bitrate and quality are gradually changed to hide strategy switches, such as prompt changes, by smoothly changing parameters. For example, given the saliency or rate allocation maps of two intra-frames, the maps of the surrounding inter-frames are interpolated, and these maps then guide the conventional video codec.
[0131] Furthermore, to limit the output of the diffusion model for the dependent frame, the reference frame (output by the decoder) is downscaled and input to the diffusion model along with the prompt. In a hierarchical group of pictures (GOP) structure, this process is repeated for each temporal layer.
[0132] As discussed in relation to the first type of system above, rather than transmitting the compressed data to the receiving device, a further option is for the compressed data to instead be stored in a local database on the sending device.
[0133] In the following, the above options of the proposed compression scheme are explained in more detail based on different embodiments.
[0134] In a first embodiment, a transmitting device (e.g., a UE device) uses (generative) compression on some input data (e.g., audio, images, video) based on a partial comparison of the device's local input generation with a shared reconstruction model of some input data (e.g., an observed scene) to achieve a controllable degree of semantic loss within bandwidth, computation, and delay objectives negotiated with a communications manager (e.g., a network capability in a 5G system).
[0135] The sending device first classifies the data (e.g., the observed scene) and uses the shared reconstruction model knowledge to derive which parts are likely to be properly reconstructed (i.e., because they are well represented in the training dataset) and which parts are not. Because data (e.g., images) classification can be performed quickly even on low-grade hardware, this initial data map can be generated very quickly. For parts of the data that can be properly reconstructed, the device can generate appropriate prompts using text inversion techniques, such as those described in "An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion" by Rinon Gal et al.
[0136] For data portions that cannot be easily reconstructed, the sending device either: (i) sends a best-guess prompt, (ii) uses more advanced and demanding techniques (e.g., "EXTREME GENERATIVE IMAGE COMPRESSION BY LEARNING TEXT EMBEDDING FROM DIFFUSION MODELS" by Zhihong Pan et al.), and / or (iii) falls back to sending just "physical" data (i.e., compressed using traditional compression techniques), e.g., in order of decreasing semantic loss.
[0137] By simply partitioning the scene in this way, semantic loss is controlled in an end-to-end manner, tailored to bandwidth and computational demands.
[0138] In a second embodiment (targeted at negotiation between the network and a user device (e.g., UE)), the amount of acceptable semantic loss is jointly set between a transmitting device with input data (e.g., a view of the actual scene) and a communications manager (e.g., 5G network functionality) that seeks to optimize throughput, delay, or overhead. The transmitting device seeks to maximize the similarity of the reconstructed dataset to reality, while the communications manager implements a higher level of compression, accepting semantic loss in exchange for lower latency and bandwidth. This, for example, leads to the negotiation of compression parameters applied by the user device.
[0139] In a third embodiment (directed at joint learning of new prompts and their semantic losses), as communication or interaction continues between the sending or receiving device, a compression / reconstruction model shared between the two users (or their associated sending / receiving devices) is dynamically updated and / or extended to previously unseen entities by sending feedback on how well newly generated prompts match the observed scene. This score is then fed into an initial data object classification (e.g., image segmentation) process at each end, for example. The model is also extended by tracking visible entities at the sender; when an unseen entity is detected, the newly generated prompt for that unseen entity and the unseen entity are added to the sending device's local model and sent to the receiving device, thereby also extending the receiving device's model.
[0140] In a fourth embodiment (directed at predictive operation), the process of the above or other embodiments is performed predictively. For example, the transmitting device predictively generates prompts based on locally predicted motion in the scene or based on communication parameters (e.g., delay) with the receiving device. These are sent in advance, and then, when an actual change in the scene is observed, it simply sends a short command on which prompt to use, adding, for example, a small correction factor (such as that described in "EXTREME GENERATIVE IMAGE COMPRESSION BY LEARNING TEXT EMBEDDING FROM DIFFUSION MODELS" by Zhihong Pan et al.) to correct for differences between the observed and predicted scene. This allows for near-zero additional delay, even for very high-bandwidth content.
[0141] In a fifth embodiment relating to the use of the proposed compression scheme in a telecommunications network (e.g., a 5G network), two users use devices (e.g., UEs) connected to an edge server to participate in telepresence, where virtual reality technology is used, for example, to remotely control a machine or to appear to participate in a distant event. Here, the edge server can be used to optimize bandwidth and latency on the 5G network. In a specific example, a transmitting UE is configured to observe a scene with a camera and forward the data to a transmitting edge server. The transmitting edge server classifies the data into data objects with high or low reconstructibility, and generates and transmits model prompts for easy-to-reconstruct data objects up to a threshold of a shared generative reconstruction model executed at a receiving edge server according to desired "bit rate" feedback from the network. For data objects that are difficult to reconstruct, the transmitting edge server further adjusts the bandwidth, computational requirements, and semantic loss according to network feedback. The transmitting edge server then sends the compressed data to a receiving edge server, which reconstructs the data from the prompts and forwards the reconstructed data to the UE. Additionally, the transmitting edge server uses the UE input to optionally predictively execute the above system as described in relation to the third embodiment above.
[0142] In a sixth embodiment relating to multimedia data, the prompt is linked to metadata including a number of parameters, for example as described below in a variation of the sixth embodiment, which serves to facilitate reconstruction of the compressed data, particularly when the data relates to multimedia data or video.
[0143] In a first variation of the sixth embodiment, prompts can be linked to a time range (e.g., to enable video or audio) so that only a single prompt needs to be sent in a given period. To this end, prompts can be linked to metadata that includes parameters such as an estimated decay time (e.g., number of frames in a video) that they are expected to be or are known to be effective. Optionally, prompts can be used beyond this decay time, but with increased semantic loss.
[0144] In a second variation of the sixth embodiment, the animation can be linked to an initial prompt, an ending prompt, a time range, and an animation pattern, and the receiver is configured to use a reconstruction model to reconstruct the motion from the prompts and metadata.
[0145] A third variation of the sixth embodiment involves synchronizing prompts or compressed data associated with different data types (e.g., audio and video). For example, a video prompt includes metadata that includes a “link” to the audio prompt. This third variation is particularly important when generative compression is applied to both audio and video, which requires linking the prompts. This can be done in various ways, such as matching the “decay times” of the audio and video prompts (as described above), using a trained reconstruction model that reconstructs both audio and video from a shared latent space, designing prompts for that latent space, and / or using generative compression only for video and separate for audio (in which case it is only necessary to temporally link the audio to the video frames). For example, compressed data for various data types is linked to metadata that determines the time frame in which the data is rendered. For example, if an image in a video is associated with the prompt "Alice walking in the street" with metadata "Time:[0.00'',5.00'']" and the audio in the video is associated with the prompt "Alice says: 'Hello, darling'" with metadata "Time:[2.00''~3.50'']", synchronization will require the audio associated with the audio prompt to play starting at 2 seconds and ending at 1.50 seconds. The fact that the audio prompt indicates that Alice is speaking will also affect the video rendering, and Alice will need to be rendered as saying "Hello, darling" between 2 seconds and 3.50 seconds.
[0146] In a fourth variant of the sixth embodiment (e.g., relating to audio and / or video prompts), the prompt metadata includes parameters that determine how to mix the reconstructed data, e.g., audio. For example, which voice is louder when two people are speaking at the same time, or who is leading when two people are walking nearby. In general, this is a relevant feature of any generative audio / image / video compression algorithm that uses text reversal. The trained prompts need to take into account overlapping data objects, e.g., overlapping voices. In some cases, it is more efficient to have multiple trained prompts (voice A, voice B, and degree of overlap) in such cases.
[0147] In a fifth variation of the sixth embodiment, for example in a metaverse scenario, an image of a person is linked to a prompt that is linked to an existing avatar. For example, a user's photo is linked to prompt "S," which is linked to avatar Y. An avatar is a digital representation of a user (participant), and this digital representation is exchanged (with other media, e.g., audio) with one or more users as a mobile metaverse medium.
[0148] In a fifth example variant, an avatar call is established, which resembles a video call in that it is visual, interactive, and provides participants with live feedback on emotions, attention, and other social information. Once an avatar call is established, the communicating parties provide information to the network in the uplink direction. The terminal device (e.g., UE) captures the call parties' facial information and locally determines an encoding (e.g., consisting of data points, colors, and other metadata) that captures the facial information. This encoded information is transmitted in the form of a media uplink and provided to the other participants in the avatar call by an IP Multimedia Subsystem (IMS). Once the media is received by the terminal device (e.g., UE) or an edge server responsible for decompressing / rendering the participants' data, the media is decompressed / rendered, for example, as a three-dimensional (or 3D) digital representation.
[0149] In this use case, processing is performed on data acquired by the terminal device to generate an avatar codec. The acquired data (e.g., video data from multiple cameras) could be transmitted over the uplink, and the avatar codec could be rendered by the 5G network. However, from a service perspective, it is advantageous to support this capability in the terminal device. First, the uplink data requirements are significantly reduced. Second, the confidentiality of the captured data may make the user unwilling to expose it to the network. Third, if the avatar is a "software-generated" avatar (such as from a game or other application), the avatar is not based on sensor data at all; in this case, there is no sensor data to send over the uplink for rendering.
[0150] If participants in an avatar call cannot be captured due to insufficient camera support in their terminal devices, they instead use text-based avatar media. This media allows participants to express what they want their avatar to say, which can include (through standardized conventions) speech pauses (e.g., "..." for a pause), emphasis (e.g., "SORRY I AM GETTING LOUD, BUT I HAVE TO SPEAK MY MIND" for louder speech and stronger gestures), and emotions (e.g., ":" for a smiling avatar). The text-based avatar media is transported to a point where it is rendered as a 3D avatar media codec. The rendering of text-based avatar media to 3D avatar media can occur at any point in the system: the caller's terminal device, the network, or the called party's terminal device. The called party's terminal device can display an avatar version of the caller and hear the caller's voice (e.g., convert text to speech). As long as the avatar configuration and speech generation configuration are properly associated with the caller, the callee can hear and see the caller speaking even if the caller provides only text as input to the conversation.
[0151] To implement the fifth variant, a network (e.g., a 5G system) supports means for a terminal device (e.g., a UE) to generate 3D avatar media information in the uplink direction and receive such avatar media information from the downlink direction, the transmission of which requires a significantly lower data rate than video.
[0152] Furthermore, the network (e.g., a 5G system) supports a means for generating 3D avatar media information running on the terminal device (e.g., UE) to support confidentiality of data used to generate the 3D avatar (e.g., data from the terminal device's camera, etc.).
[0153] The network (e.g., a 5G system) is further configured to support means for providing continuity of service to parties in an IMS video call even when communication performance for one or more parties degrades to the point where video quality is unsatisfactory or video is no longer possible. In this case, an avatar call between the same parties can be used as a fallback in lieu of the video call. Later, when communication performance improves and the video call becomes possible again, the avatar call is replaced by the video call.
[0154] Additionally, the network (e.g., a 5G system) is configured to support means for transporting and processing user-supplied standardized text-based avatar media as 3D avatar media, where the text-based encoding includes standardized expressions for indicating emphasis, speech pauses, and emotions.
[0155] In a sixth variant of the sixth embodiment, compression techniques are applied to audio information. For example, consider a (metaverse / movie) scenario where two people, one French and one Spanish, are talking in English in the middle of New York City traffic, and someone yells, and the content of the conversation is, for example, about politics. Then, interchangeable meanings are: traffic, woman yelling, English, Person 1: Spanish accent / male, Person 2: French accent / female, Person 1 says: "Mr. X is my favorite politician," and Person 2 angrily replies: "Are you crazy? We can't hang out anymore!" These meanings can also be passed to a predictive model to generate audio / utterances. In particular, the meaning of "Mr. X is my favorite politician" can be expanded to a one-minute monologue in which the person gives reasons for this opinion and explains why Mr. X is their favorite politician.
[0156] In a related variant, the audio prompt itself indicates the meaning of the message and the length of the utterance, but the content is generated locally or by other means, such as a language model, for example the Generative Pre-Trained Transformer (GPT) family of language models such as ChatGPT (https: / / en.wikipedia.org / wiki / Generative_pre-trained_transformer). The generated text is then text-to-speech converted to fit the required time interval.
[0157] In a seventh embodiment relating to negotiation between a local (user) device (e.g., UE) and a network with split rendering, the local device interacts with a local edge server that communicates with a remote edge server to which a remote device is connected. In such cases, functionality is split between the local device and the edge server. For example, the local device sends initial data (e.g., the entire image) to the local edge server, which calculates prompts based on a "full" shared reconstruction model. The prompts and the object and the partial "reconstruction model" are then fed back to the local device, which can then generate prompts associated with the object based on the "partial" (lighter) reconstruction model. In some cases, there is only one edge server.
[0158] Typically, generation and reconstruction are split between the local device and a local edge server (conceptually similar to split rendering). In some cases, existing mechanisms that enable split rendering can help determine which parts of the reconstructed model run where, e.g., on the device, on the edge server, or in the cloud.
[0159] In a related embodiment, an edge server can be used to take over some of the UE's computational load, e.g., in the context of split rendering. For example, the UE can perform certain computationally less intensive tasks locally, such as rendering the background of an image, while objects that are more computationally intensive without loss of semantic data are rendered at the edge server. In this embodiment, the transmitting edge server (UE) indicates which data objects are preferably reconstructed where, or indicates the required reconstruction capabilities, allowing the receiving edge server (UE) to decide which entity does what, i.e., to split the decoding / decompression / rendering capabilities, e.g., based on local policy.
[0160] In an eighth embodiment, a shared reconstruction model is learned during the interaction. For example, when a sending device (or encoder) first observes an image or view of a target to derive a prompt, the sending device transmits both the prompt and the image or view of the target (e.g., using conventional compression techniques). The image is then added to the local and remote reconstruction models.
[0161] In general, an "effective holistic model" can be learned during an interaction. For example, an "effective holistic model" combines a pre-trained diffusion model (which remains static) with dynamically learned prompts generated during the interaction. The sending device learns the prompts and transmits images of relevant objects. The system can then dynamically update the "effective holistic model" during use by taking into account new shared prompts and images present on both the sender and receiver. In other words, even if the baseline diffusion model remains unchanged, relevant objects are marked as "easy" to reconstruct. For example, if a person in an image (i.e., a subject in the data) turns sideways, the "dynamic" learned prompts for that person no longer work, and the sender must derive new learned prompts, which can be updated in turn. In some cases, it is advantageous for both the sending device (encoder) and the receiving device (decoder) to store all these learned prompts and effective holistic models, even after the communication has ended.
[0162] In a ninth embodiment relating to receiver-side prediction, the receiving device (decoder) can calculate a prediction (e.g., a motion prediction) based on the compressed data already received and generate a prompt accordingly. It then returns its best guess prompt to the sending device (encoder), which compares it with the actual one and selects which prompt to use and sends a small correction image if necessary.
[0163] In general, the decision of where to generate the motion estimate depends on at least one of the application, delay requirements, desired reconstruction fidelity, and number of communicating parties. This is configurable by a policy deployed by the communications manager and deployed to the sending device (encoder) or receiving device (decoder). This is also configured / indicated in the metadata transmitted along with the compressed data. Note that this includes not only the motion estimate, but also the measured motion vectors at a given time, e.g., at the time of transmission by the sending device. For example, the measured motion vectors, which refer to the velocity of the semantic object at time t, are part of the transmitted prompt. The receiving device takes the received prompt into account and measures the motion vectors to perform predictive decompression / rendering.
[0164] In a tenth embodiment relating to a structured vocabulary and generation model for prompts, the description language can be a structured language / vocabulary, a human-readable language, or pseudo-words, such as those described in "An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion" by Rinon Gal et al.
[0165] In an eleventh embodiment related to the predictive properties of the proposed compression scheme, if both the sending device and the receiving device have the same generative model, the receiving device will also generate the same prompts. That is, the sending device only needs to select one of the common prompts and send a correction factor. In such an embodiment, the system can operate in one of two ways: sending a prompt and then sending a correction factor, or sending an indication of which prompt to use and a correction factor.
[0166] Once the prompt is generated, there is no significant difference in bandwidth requirements between just sending the prompt itself or sending instructions on which prompt to use (because the prompt is just a short text string).
[0167] In a first variation of the eleventh embodiment, the appropriate configuration depends on how quickly the prompt can be generated from the "instructions" at the receiving device compared to the delay in sending the already generated prompt. This tradeoff is determined by the sending device and may be encoded in the transmitted metadata or configured by policy. The tradeoff may vary depending on time, data subject, etc.
[0168] In a second variation of the eleventh embodiment, it is advantageous to select a prompt at the receiving device before the (predicted) event occurs and receive the correction factor as soon as possible thereafter. This is because the receiver first uses the prompt to generate a (lossier) output and then uses the correction factor to "smooth" it to the desired output. In many cases, it is desirable for at least some of the output to have near-zero delay.
[0169] In a twelfth embodiment, the proposed compression scheme is applied to a 5G media streaming framework that includes concepts such as reliable media capabilities, including aspects such as adaptive bitrate encoders (see 3GPP® specification TS26.501). The transmitting device accesses the reliable media capabilities, for example, via a media control interface. In this scenario, a "semantic loss optimizer" is defined as an additional reliable media capability. TS26.117 also defines speech and audio profiles for 5G media streaming, which are adapted accordingly. Similar to the bitrate negotiation mechanism in Voice over LTE (VoLTE), the proposed compression scheme implies a semantic loss negotiation mechanism. The part that needs to be adapted and defined involves negotiation between the transmitting device or edge server and the network capabilities to optimize the semantic loss for a given bandwidth. The reasons are as follows. -Semantic loss is much more scene-dependent than traditional lossy compression algorithms, so simply having the network provide the bandwidth objective and the transmitting device (e.g., UE) compress the video to reach that objective is not optimal. In some cases, the transmitting device can achieve high quality at very low bandwidth (e.g., very common scenes like forests), while in other cases, the same bandwidth can lead to near-total semantic loss. Given the proposal to dynamically update the entire shared model during interaction, it is difficult / impossible for network features to accurately predict semantic loss and justify device feedback. The predicted or determined delay or bandwidth target may be received some time in advance from the network by the sending device or edge server. - Shared reconstruction model structures can be standardized. -A central database of already learned prompts can be provided for download / pre-caching.
[0170] In a thirteenth embodiment, a video codec is provided that uses a rate-distortion optimization algorithm to provide optimal reconstruction quality for a given target bit rate or minimum bit rate for a predetermined quality level. Rate-distortion optimization is used to improve video coding efficiency and aims to find an optimal trade-off between reconstructed video quality and coding rate. Conventional video coding standards, such as the new H.266 / Versatile Video Coding (VVC), H.265 / High Efficiency Video Coding (HEVC), and H.264 / Advanced Video Coding (AVC), use the sum of squared errors (SSE) as a distortion metric because SSE can efficiently represent image fidelity. In the context of the proposed compression scheme, rate-distortion optimization is achieved by, for example, generating text prompts using text inversion, performing text-to-image transformation based on the text prompts, and calculating color transformation vectors that modify the colors in the spatial domain to better fit the actual data. Additionally or alternatively, texture transformation (texture transformation vectors) can also be used to achieve the same. In general, a transformation can be created and applied to a given feature that modifies the region within the given region to better fit the data.
[0171] In a fourteenth embodiment, a shared reconstruction model is used that derives data portions that reference at least one of data objects, bounding boxes of classified regions, and masks obtained from instance segmentation, where the reconstruction is a policy that determines how portions of data need to be reconstructed in various situations / contexts.
[0172] In the above embodiments, it is not clear how to handle spatial boundaries between regions and objects that appear / disappear over time. For example, most text-to-image diffusion models do not clearly define the location of objects, but simply place the main subject in the center of the image.
[0173] Thus, a fifteenth embodiment, combined with other embodiments or used independently, includes the ability to spatially or temporally transition between prompt-based synthesis and conventional video codecs.
[0174] In a first example of the fifteenth embodiment, in combination with other examples, a video atlas is used to group spatial regions that would otherwise need to be coded using traditional video coding into a single video atlas. These regions may be generated by instance segmentation. The renderer first runs a prompt-based synthesizer and then overplots the instance segmentation texture from the video atlas during a second pass.
[0175] In a second example of the fifteenth embodiment, in combination with the other examples, the portion to be coded (target) is input to a denoising process using a conventional video codec, which requires training a second diffusion model that synthesizes an image using the text and partially defined image data as input.
[0176] In a third example of the fifteenth embodiment, in combination with other examples, spatial consistency is another form of semantic loss that can be accepted or traded off against bandwidth and computational requirements.
[0177] In a fourth example of the fifteenth embodiment, in combination with other examples, image-guided inpainting combined with instance segmentation is used.
[0178] In another example, similar capabilities can be applied to temporal transitions between prompt-based synthesis and conventional video codecs. For example, an object in (initial) time interval t0 is synthesized using a conventional video codec (e.g., because the object is not yet part of the model at the receiver). During this time interval, the model at the receiver is expanded to include the prompt associated with this object. Then, in time interval t1, the same object is synthesized using a prompt-based approach. Therefore, the synthesized object between t0 and t1 needs to be temporally transitioned, similar to, for example, the spatial transition example above.
[0179] The above embodiment does not explain how to composite frames when the video / image / sound is composed of two or more objects that cannot be rendered from a single text prompt.
[0180] Thus, in a sixteenth embodiment, which may be combined with other embodiments or used independently, a diffusion model for image (and other data types) synthesis is trained from a partially masked image, where the masked regions correspond to portions predicted by a first text prompt and the unmasked regions correspond to portions predicted by a second text prompt. The network can then be recursively applied during inference time to fill in the entire image using multiple text prompts.
[0181] The above embodiments do not fully explain how exactly to process the video information.
[0182] Thus, in a seventeenth embodiment, which may be combined with other embodiments or used independently, camera or independent object motion parameters are included during training to provide the network with information on how to generate temporally consistent video.
[0183] The difference between video and images is that in addition to the spatial structure found in images, video also has a temporal structure. Because video is simply a collection of images operating at a particular temporal resolution, i.e., frames per second, the information in the video is not only encoded spatially (i.e., into objects or people in the video), but also sequentially according to a particular order (e.g., catching a ball vs. throwing a ball, salsa dancing vs. hugging, etc.). This additional information makes video classification both very interesting and difficult.
[0184] In a first variant of the seventeenth embodiment, a new generative prompt is predicted for each frame, at the expense of computation time. Generative compression techniques are applied, such as processing the video by operating on keyframes and infilling. Increasing the number of keyframes (up to 100% as described above) to reduce the infilling error can make this more robust in situations with large relative motion. Similarly, techniques based on neural video compression are also used, using optical flow concepts to train a "motion-aware" branch that is used to predict errors in the interpolation branch.
[0185] Optical flow is a powerful idea that has been used to significantly improve accuracy and reduce computational cost when classifying videos. It is a pixel-by-pixel prediction based on an assumed brightness constancy, which means that we attempt to estimate how the brightness of a pixel on a screen will change over time. This assumes that the pixel characteristics (e.g., RGB values) at time t are the same as those at a later time t + Δt, but at a different position (denoted by Δx and Δy, for example). The change in position is what is predicted by the flow field. As an example of this, assume that at t = 1 second, we have an RGB value of (255, 255, 255) and it is at an x,y position in the frame (e.g., 10, 10). Optical flow assumes that at t = 2 seconds, the same RGB value (255, 255, 255) is still present in the screen, but with motion, it will be at a different position in the frame (e.g., 15, 19). Therefore, the optical flow displacement vector for this motion is [9, 5]. This means that once the original pixel position is taken and the displacement vector is applied, a new image can be predicted.
[0186] In a further related variation of the seventeenth embodiment, generation prompts are used to generate objects (moving / static) over multiple frames.
[0187] In an eighteenth embodiment, which is an extension of the second embodiment, the reconstruction serves as a predictor for at least part of the data (e.g., image), and a conventional data (e.g., image, video) encoder is used to compress the residual data (e.g., image). This means that enhanced video coding can be performed using a diffused reference layer and a conventional enhancement layer. Additionally or alternatively, the reconstruction serves as a reference frame used by the conventional video encoder to encode the frame.
[0188] In the above embodiment, without mitigation, it is visible when the codec switches strategies between diffusion-based neural coding and traditional video coding, as the tradeoffs are very different (e.g., where to place bits).
[0189] Thus, in a nineteenth embodiment, which may be combined with other embodiments or used independently, The contribution of diffusion to the overall video coding is gradually changed (at some cost) to hide strategy switches. There are hybrid solutions (e.g., as in the 18th embodiment), in which some traditional video codec is involved to hide the "diffusion weirdness"; - a "smoothing" function is applied, for example by using a "smoothing" technique on the inpainted image; and / or For example, the variant of embodiment 15 applies.
[0190] In a twentieth embodiment, the (e.g., video) codec is primarily conventional, and text labels and diffusion are used to improve in-loop or post-processing filters in the encoder and / or decoder. This means that the codec is defined and operates at the object level. This allows temporal features, such as motion (in video) or changes in tone of voice (in audio), to be encoded using fewer bits, since the codec operates at the object level rather than the pixel / block level. The text labels may or may not be known (object database), and the decoder must learn the object appearance on the fly. Note that "in-loop filter" implies that the encoder and decoder run the same filter, and in this context also implies that the same learning step is performed for unknown labels.
[0191] In a 21st embodiment, there is a side channel for a subject database to efficiently transmit videos of a particular genre. To achieve this, the decoder fetches one or more databases if there are labels corresponding to the video. This can be done once, periodically, or on-demand, and the decoder keeps a local copy. Alternatively or additionally, there is a general subject database that is part of the standard, and all decoders that conform to that standard recognize all these labels. The labels have a namespace that indicates which database they belong to. A video uses multiple databases.
[0192] A twenty-second embodiment relates to a further class of embodiments (combined with other embodiments or used independently) in which multiple receiving devices receive compressed data presented by a transmitting device. The receiving devices render the compressed data at different quality levels depending on their capabilities, link bandwidth, user requirements, etc. The transmitting device may generate, for example, multiple compressed streams of different levels of quality and / or semantic loss / compression, a single layered stream from which streams of desired quality levels are extracted, or a single high-quality stream that can be further compressed by components in the network as needed.
[0193] In a first example of the twenty-second embodiment, a sending device generates multiple streams and forwards them to a local edge server. Receiving devices granted access to the streams negotiate quality levels with the local edge server, and the local edge server negotiates with the sending edge server for delivery of streams at appropriate quality levels.
[0194] In a second example of the twenty-second embodiment, a sending device generates a hierarchical stream and distributes it to a local edge server, which extracts the appropriate quality level stream for distribution to a requesting receiving edge server.
[0195] In a third example of the twenty-second embodiment, the sending device generates a single high-quality stream, either uncompressed or compressed with conventional techniques, and delivers it to a local edge server, which performs image segmentation and compression according to the quality level requested by the requesting receiving edge server.
[0196] In a first variation of the first to third examples above, the sending edge server distributes all data to the receiving edge server, which then derives the compressed data stream upon request by the receiving device.
[0197] In a second variation of the first to third examples above, the sending edge server collects feedback from the receiving devices and optionally distributes it to the sending device.
[0198] In a further variation, the involved devices (e.g., receiving device and receiving edge) perform negotiation and / or signaling of the required data stream to use, e.g., based on the current network capacity (e.g., bandwidth), device capabilities / capacities (e.g., CPU). For example, if the receiving device has available CPU (the receiving edge receives an indication from the receiving device) and the network is not saturated, the receiving edge decides to provide / send a higher quality data stream (requiring more bandwidth / more CPU).
[0199] In a further variation, the participating devices explicitly / implicitly announce or decide which data streams they require, for example, a device may decide which data streams are preferred based on the screen size of the receiving device.
[0200] In a twenty-third embodiment, as a variation of the twenty-second embodiment, combined with other embodiments or used independently, data from a transmitting device is recorded for later playback. The recording comprises a single high-quality data stream, a hierarchically compressed data stream, or one or more compressed streams of different quality / semantic compression characteristics. The recording further comprises auxiliary data for compression purposes, such as a compression / decompression database, associated prompts, etc.
[0201] The recordings are stored either on the sending device, on a network device (eg, a local edge server at the sender or a local edge server at the receiver), or on a physical carrier (eg, a memory card, a disk drive, an optical disk, etc.).
[0202] Delivery of the recording to a requesting receiving device may be in a streaming format, or the recording may be delivered as a single file or other data object. The receiving device negotiates the compression quality level in the same manner as in the previous embodiment.
[0203] This embodiment (and other embodiments) is advantageous when applied to streaming services such as video streaming services or video conferencing services where meetings are stored.
[0204] In a twenty-fourth embodiment, generated compression is used for two-way or multi-way communication between multiple users. This is particularly suitable when, for example, screenshots of peer users are rendered as thumbnails during a presentation. In this configuration, a common compression / decompression database can be used to save resources.
[0205] As mentioned above, the proposed data compression method enables various multimedia data to be compressed with high efficiency.
[0206] Generative compression based on diffusion models offers potentially significant bandwidth savings. The process is based on generating small inputs, usually short text strings, which can be used by a reconstruction model to recreate an approximate version of the original observed data. These inputs can then be stored in a database or transmitted over a network with a much smaller footprint than the original data, thereby outperforming known compression codecs. This is relevant for images and video, but also applies to audio compression or other data types.
[0207] In a further embodiment, combined with other embodiments or used independently, a user has an application, such as an application running on a smartphone or computer, that can create and edit content, e.g., audiovisual material such as video, using or supported by at least a generative model and an interpretive programming language. The user, who is the content owner, uses prompts to create audio, video, images, etc. from the model that constitutes the subject content, e.g., audiovisual material, even if other types of data / content are included. The user then stores the content and / or transmits the content to at least a (receiving) user device.
[0208] In further related embodiments, the generated content is considered to be or referred to as synthetic data specified by a "content program" taken as input by an interpreter of an "interpreted programming language" that relies on a generative model for interpretation.
[0209] In further related embodiments, the generative model provides the user with several options when trying out a prompt, so the user selects one of the model's outputs. Because multiple generative models are involved, the user includes a model identifier and / or version to ensure that the receiving party can reproduce the same content. Because the generated audio / image / video, etc. does not fully satisfy the user, the user fine-tunes the output (during the content creation process) and includes changes to the data. During the content creation process, the user takes audio / video / ... samples and assigns them to prompts to enhance the model based on user-defined inputs. The user creates content, e.g., videos, by using a programming language, where a "content program" (or program) looks like the following example: [New video, Duration 12 seconds, Resolution xyx] [Generative model ID xyz, version vxyz] [Background image: part in spring, sunny weather] [Background sound: happy piano music] [Prompt: “small white cat”; Prompt_output: #3; Start time: 1 second; Duration: 10 seconds; Action: “walks from left to right”] [Prompt: “big fat dog”; Prompt_output: #2; Start_time: 2 second; Duration: 9 seconds; action: “walks from right to left”; action: “smells cat and smiles”; action: “follows cat”]
[0210] In this "content program" example, new commands are given on new lines, enclosed in brackets for example, which means that a standard is needed to determine what the new commands are.
[0211] In this example, one or more (generative) models are shown, which may be indicated, for example, by name or URL. These generative models are then used by the interpreter to generate content specified by a "content program."
[0212] In this example, there are several keywords that help determine the type of content to generate, such as "new", "video", "duration", "resolution", "Generative model", "Background image", "Background sound", "Prompt", "Prompt_output", "Start_time", and "Duration".
[0213] In this example, there is a standard way to indicate which action is associated with a keyword, e.g., the sequence [Keyword:Action;] is used to indicate that "Keyword" starts a new keyword, the action associated with the keyword appears after the ":", and ends with a ";".
[0214] This information may be entered as text or through some other type of user interface, such as a graphical user interface.
[0215] The user then plays back the generated video and further edits it until the user is satisfied, at which point the user releases / publishes it.
[0216] To verify the user who published the content, the data used to generate the (audiovisual) content (or a hash of it) is signed by the user. The fingerprint is made available (e.g. attached to the content or made available in a public repository or on a blockchain), and the data is also made available, allowing other users to create further content based on it.
[0217] In a further embodiment, the program is written in an interpreted programming language, an interpreter runs, for example on a computer, for example at the sender or receiver, and one or more (generative) models are used by the interpreter to generate the content.
[0218] In a further embodiment, the compressed data obtained by the encoder has the same format as the "content program" and the compressed data is characterized by similar keywords / metadata, e.g., to facilitate decompression / reconstruction of the data.
[0219] In a further embodiment, the compressed data has compressed data fields for different data types (e.g., audio and video), and keywords / metadata are used to synchronize the different data types during decompression / reconstruction.
[0220] In a further embodiment, which can be combined with other embodiments or used independently, a sending user / company / device has an application for generating content, such as audiovisual content. For example, audiovisual content for a video streaming service such as Netflix. The sending user / company / device then distributes the content to one or more receiving user devices. In this embodiment, the models applied at the receiving end are personalized. For example, the sending user / company provides the receiving user with two or more (generative) models. For example, Generative Model A generates content with specific characteristic A (e.g., people are taller, have an accent when speaking, buildings have a certain style), and Generative Model B generates content with specific characteristic B. The receiving user then selects / uses / is encouraged to use one of Model A or Model B. This embodiment has the advantage that the receiving user can better choose the appearance of the generated output, thereby being more satisfied with the generated output and the service provided by the sending user / company.
[0221] In a related embodiment, the sending user / company / device uses a content generation application (as in the above embodiment) to create / specify (slightly) different versions of the content through the content program so that the decompressed / reconstructed content / data matches the preferences of the receiving user device. For example, the content program includes the word "PREFERRED" to indicate to the receiving device that background sound should be generated according to local preference configuration. For example, "PREFERRED piano music" means that the receiving device / user's preferred piano music should be used. For example, "Prompt: "Little white cat"; Prompt_output: PREFERRED;" indicates that the data generated when using "Little white cat" as the prompt is the preferred data according to local policy. For example, User A's local preference policy prefers long-haired cats, while User B's local preference policy preference configuration prefers short-haired cats. [New video, Duration 12 seconds, Resolution xyx] [Generative model ID xyz, version vxyz] [Background image: part in spring, sunny weather] [Background sound: PREFERRED piano music] [Prompt: “small white cat”; Prompt_output: PREFERRED; Start time: 1 second; Duration: 10 seconds; Action: “walks from left to right”] [Prompt: “big fat dog”; Prompt_output: #2; Start_time: 2 second; Duration: 9 seconds; action: “walks from right to left”; action: “smells cat and smiles”; action: “follows cat”]
[0222] Generally, these embodiments illustrate that personalized content / data is generated by, for example, a personalized (generative) model, a personalized content program, and / or preference policies.
[0223] In a related embodiment, the receiving user device and the sending user device negotiate a decompression / content generation model based on user preferences, network capabilities, and the capabilities of the receiving device (e.g., CPU, rendering device, etc.), where the negotiation involves zero or more interactions. The sending device creates a profile of the receiving device and assigns one or more corresponding models. If one or more interactions are required, the receiving device indicates its preferences to the sending device, and the sending device adjusts the provided models accordingly.
[0224] In a related embodiment, the sending user / company / device provides one or more models to the receiving user, and the receiving user uses one or more models. For example, many models are available for reconstructing a person's voice. For example, many models are available for reconstructing people / people in video data.
[0225] In a related embodiment, the sending user / company / device provides data (e.g., audiovisual data such as movies) containing "packages" of official models, allowing the end user to select the preferred one to decompress / render.
[0226] In a related embodiment, such data includes the approved models used, including official or third-party models, with authorized models signed by the sending user / company / device.
[0227] In a related embodiment, the sending user / company / device is also the data / content owner of the sent data / content, and the data / content is stored (e.g., on a cloud-based video streaming service such as Netflix) and the data / content is streamed on demand.
[0228] In a related embodiment, the sending entity has a generic model that is personalized for the receiving entity, for example, based on preferences given by the receiving entity, and then (only) the personalized model is provided to the receiving entity.
[0229] In related embodiments, aspects that are personalized or part of the model include clarity of speech (e.g., by selecting clearer voices), aiding visibility (e.g., by changing lighting or color rendering, or viewing angle), the tone of people's voices, a person's emotional state (happy, sad, etc.), the level of certain actions (e.g., violence; e.g., the same action: Man A hitting Man B may be interpreted differently, e.g., harder or softer, depending on the model's preferences).
[0230] In related embodiments, data, content, or signals refer to images, audio, video, etc. that are generated / transmitted.
[0231] In a related embodiment, the model is based on, uses, or is augmented by a generative adversarial network trained to adjust the "style" of one data / content / signal to resemble the style of another data / content / signal.
[0232] In a related embodiment, the decompressed output based on use of the generative model with the received prompts is passed to a second model, e.g., a generative adversarial network, to adjust the decompressed output / signal to conform / resemble the style of a "target / preferred" signal, where the target / preferred signal is determined according to the preferences of the receiving user. In some examples, implementations of the above-described embodiments of systems and methods in a cellular network environment (e.g., 5G) are used, for example, as described below or elsewhere in this application.
[0233] In the context of scene observations, "scene" is understood to mean the true observed reality observed by some device that uses sensing to generate the "scene observation" to be compressed. For example, a device that observes a visible scene using a camera generates a scene observation in the form of image or video data. Instances derived from a segmentation process operating on the scene observation data are presumed to represent semantically linked objects, items, or people. For example, in an image of a person in a forest, the "forest" is an object, just like the "person." The loss of realism between the regenerated output and the original scene observation is a semantic loss. This differs in several ways from the distortions introduced by traditional codecs, particularly in that it is highly dependent on the original object being observed. Thus, there are multiple types of semantic loss, some of which only apply to specific modalities (e.g., images or video). Examples include object loss (e.g., category loss due to reconstruction of an object different from what was observed, e.g., a generic person rather than a specific person), color loss (e.g., the correct object is generated with the wrong color), focus loss (e.g., an out-of-focus object in the scene observation is reconstructed in focus), segment loss (e.g., the correct object is reconstructed in the wrong location), and motion loss (e.g., loss of realism over time in a video, such as occurs with moving objects). The input (i.e., "prompt") to the reconstruction model (e.g., a diffusion model) is trained or created to describe a given object in more detail compared to a simple prompt. Typically, a trained prompt is necessary to enable the reconstruction of objects that did not form part of the reconstruction model's training dataset. An example is learned embeddings (e.g., those described above) generated by text inversion techniques. Reconstruction refers to a reconstructed version of a scene observation (including some residual errors or losses, such as semantic loss and traditional distortions).Finally, "ease of reconstruction" is understood to mean a measure that takes into account both the resulting semantic loss and the computational requirements to generate a reconstruction of a given subject using a given model from a given prompt.
[0234] FIG. 5 shows a schematic block diagram of the different layers involved in a compression system, according to various embodiments.
[0235] IO layer and network (NW) layer (L IO and L NM The 5G system includes a transmitting mobile terminal (Tx UE) 22. The transmitting mobile terminal (Tx UE) 22 is a device such as a smartphone or AR glasses that performs sensing (S) 12 (e.g., enabling generation of scene observations via a camera), a user interface (UI) 14 for user input and output, and general network and / or connectivity functions (e.g., Wi-Fi, 5G). Furthermore, a transmitting edge server (Tx ES) 24 is a local edge server of the Tx UE 22 equipped with high computational resources and network connectivity. The Tx UE 22 is network-connected to the Tx ES 24. Furthermore, a network function (NF) 26 may be executed remotely from the Tx UE 22 and the Tx ES 24 and is a function that provides bandwidth and throughput information, possibly predictively. The network function may be implemented as a trusted media function 26 accessed via the 5G system, for example, as defined in 3GPP TS 26.501. Finally, a receiving mobile terminal (Rx UE) 28 and a receiving edge server (Rx ES) 29 are provided, which are similar to Tx UE 22 and Tx ES 24, respectively, but are both somehow distant (spatially and / or temporally) from the scene and are therefore not configured to perform sensing 12 to generate scene observations.
[0236] The codec layer (L_CD) further includes an encoder (ENC) 32 capable of generating a bitstream (BS) from input scene observations. The encoder 32 is implemented as software and / or hardware by an appropriate device (such as an edge server, e.g., Tx ES 24). The encoder includes or accesses a common model suite (CMS) 40 (described below) and a scene comparator (SC) 324, a software / hardware module for comparing the reconstruction with scene observations and calculating a loss score. To calculate distortion, the scene comparator 324 implements a data comparison algorithm, such as a record checksum calculation algorithm, a matching algorithm, or a correlation function (which determines whether and to what extent two input data sets (e.g., pixels of two images) are correlated), or a norm function (e.g., the L2 norm, which determines whether two input data sets (e.g., pixels of two images) are close to each other, e.g., using two norms). To calculate semantic loss, the scene comparator 324 must implement a more task-specific algorithm depending on the type of data being encoded. For example, for images containing people, the scene comparator 324 may implement face detection and recognition algorithms to determine how closely the reconstructed face matches the observed face, for example, by implementing algorithms that check whether the reconstructed / decompressed image / sound / data is semantically correct, e.g., whether all people have two hands with five fingers each, or whether cats are wind-free.
[0237] Additionally, the encoder includes an in-loop decoder (ILD) 326, which is a decoder for generating the reconstruction (described below).
[0238] The bitstream 34 generated by the encoder 32 includes at least one of a learned prompt (LP) 342, a physical input (PI) 344 (e.g., scene observations compressed using any type of non-prompted codec, e.g., a conventional codec (CC) 46 of the common model suite 40 described below), a saliency map (SM) 346 (e.g., a map of the scene observations indicating the locations (in space and / or time) of objects along with estimated saliency scores, e.g., generated by a saliency mapping module (SMM) 47 of the common model suite 40 described below), and consistency data (CD) 348 (e.g., additional data generated by the encoder 32 and used to increase the consistency of the reconstruction across the spatial and / or temporal dimensions, e.g., generated by a consistency data module (CDM) 45 of the common model suite 40 described below).
[0239] The codec layer further comprises a decoder (DEC) 36 that can receive the bitstream 34 as input and generate a reconstruction. The decoder 36 is implemented in software by a suitable device (such as an edge server, e.g., Rx ES29) and includes or has access to a common model suite 40, described below, and a reconstruction order (RO) module 364 that is configured to analyze the input bitstream 34 and determine an optimal order for creating reconstructions of various objects on the scene observation. To this end, the RO module 364 implements a set of heuristics applicable to each object and uses feedback from the hardware capabilities of the reconstruction device (e.g., Rx ES29).
[0240] Finally, the model layer (LMOD) comprises a common model suite 40 including at least one shared reconstruction model (SRM) 412 (e.g., a machine learning model (perhaps a diffusion model) that receives prompts (including learned prompts (LP) 414) as input and produces reconstructions as output). An example is the diffusion model mentioned above. After a dialogue has begun, a base version of the SRM 412 can be combined with any newly learned prompts 342 to form an overall effective model (OEM) 41.
[0241] Additionally, the common model suite 40 comprises a prompt generation model (PGM) 43 for generating learned prompts for given objects within the scene observation. This is achieved by using text reversal to generate learned embeddings, as described above, although similar techniques are also suitable.
[0242] Additionally, the common model suite 40 includes an instance segmentation (IS) model 44 configured, for example, to segment input scene observations, identify semantically linked objects, and / or estimate the ease of reconstruction of those objects. The type of instance segmentation model 44 depends on the type of input, but for example, for images, a simple fully convolutional model for real-time instance segmentation, such as that disclosed in "YOLACT - Real-time Instance Segmentation" by Daniel Bolya, ICCV2019 (available from https: / / openaccess.thecvf.com / content_ICCV_2019 / papers / Bolya_YOLACT_Real-Time_Instance_Segmentation_ICCV_2019_paper.pdf), is suitable. The model for determining ease of reconstruction may, for example, implement an in-loop decoder (e.g., in-loop decoder 326) or be based on a set of heuristics.
[0243] The saliency mapping module 47 is configured to generate a map that provides saliency scores for a given segmented input, e.g., both per object and, optionally, per region within an object, where a hierarchy of objects is applied. The type of model used by the saliency mapping module varies depending on the type of input, but for example, images, the model described in M. Ahmadi, M. Hajabdollahi, N. Karimi, and S. Samavi, "Context-Aware Saliency Map Generation Using Semantic Segmentation," Iranian Conference on Electrical Engineering (ICEE), Mashhad, Iran, 2018, pp. 616-620, doi:10.1109 / ICEE.2018.8472577 is suitable. An alternative implementation uses a set of saliency heuristics.
[0244] The consistency data model 45 is configured to generate consistency data for a given object and / or entire scene observation. Several algorithms are used for this purpose. For example, for per-object consistency data, techniques such as guided imagery by Zhihong Pan et al., "EXTREME GENERATIVE IMAGE COMPRESSION BY LEARNING TEXT EMBEDDING FROM DIFFUSION MODELS," November 14, 2022, are suitable. For overall scene observation consistency data, inpainting algorithms such as those described in "Paint by Example: Exemplar-based Image Editing with Diffusion Models" by Binxin Yang et al. are suitable for spatial consistency within images. Furthermore, keyframe techniques (which define a "decay period" during which keyframes are assumed to remain relevant), such as those described in "Generative Compression" by Shibani Santurkar et al., are suitable for temporal consistency within video (or other data types, such as audio). Alternatively, there are more advanced techniques such as those described in "Stable View Synthesis" by Gernot Riegler et al., CVPR 2021.
[0245] Conventional codecs 46 capable of compressing scene observations according to conventional algorithms include non-generative compression codecs (eg, JPEG, H.265, etc.).
[0246] The scene prediction model (SPM) 42 is provided as part of the common model suite 40 and is configured to predict changes in scene observation before they occur. For example, for video input, a next-frame prediction model is used, which is configured to predict what will happen next in the form of one or several images. This prediction is based on understanding the information contained in previous images captured so far. This refers to starting with a series of unlabeled video frames and building a network that can accurately generate subsequent frames. The network's input is the previous few frames, and the prediction is the next frame. These predictions can be not only human motion, but also the motion of any objects in the image and the background. Modeling content and dynamics from a video or image is the main task of next-frame prediction, which is different from motion prediction. Next-frame prediction predicts future images based on several previous images or video frames, while motion prediction refers to inferring dynamic information such as human motion and object trajectories from several previous images or video frames.
[0247] Additionally, the common model suite comprises or has access to at least one of a Semantic Loss Type Database (SLTD) 49 for storing different types of semantic losses and their associated weightings, and / or a Prompt Database (PD) 48 for persistently storing learned prompts (that persist across interaction times).
[0248] Additionally, there is a side channel for the target database to efficiently transmit videos of a particular genre, and if the video has such a label, the decoder fetches from such target database once.
[0249] Furthermore, there may be a general object database, e.g., as part of a standard (or system, technology, product), and all decoders conforming to that standard are configured to recognize labels in this general object database. These labels have a namespace that indicates which database they belong to. Video uses multiple namespaces.
[0250] Figure 6 shows a schematic flow diagram of the compression and decompression process according to various embodiments using the different layers of Figure 5. Note that not all steps are always required, the order of steps may be adjusted, and some steps may be performed multiple times. For example, for compression, it is not necessary to sample / sense the scene.
[0251] For example, at the start of the process, input I from the scene via sensing function 12 S and / or input from the user via the user interface 14 in the IO layer. u are converted into scene observation and UI data for Tx UE 22 in the network layer, and Tx UE 22 forwards the corresponding scene and UI data to Tx ES 24 and exchanges control information with Tx ES 24 .
[0252] Note that the above example applies when a user wants to send an image to another user. Alternatively, the user can directly enter a prompt to send to the other party. For example, the user tries the prompt "Sweet cat", the local model generates four cats, the user selects one of them (e.g., #2), and then the user decides to send the prompt ["Sweet cat", #2] to the other party. In this case, there is no need to further compress the input in the UI.
[0253] The Tx ES 24 implements an encoder 32 at the codec layer, which accesses the common model suite 40 to obtain instance segmentations, saliency maps, and learned prompts, and forwards them to the in-loop decoder 326. The in-loop decoder 326 generates a reconstruction that can be compared with actual scene observations by the scene comparator 324. This loop is repeated until a desired semantic loss is reached. Based on this, the encoder 32 generates (GEN) a bitstream 34 that is used as input (I) for the decoder 36 at the receiving end, which accesses the common model suite 40 to obtain partial reconstructions that are fed to the RO module 364 to generate reconstruction instructions (RO).
[0254] At the network layer, the Tx ES 24 transmits the generated bitstream 34 to the network function 26. The network function 26 optimizes the bitstream (BS) 34 and forwards the optimized bitstream 34 to the Rx ES 29. The Rx ES 29 implements a decoder 36 at the codec layer and obtains a reconstruction (REC) output from the decoder 36. Furthermore, the Tx ES 24 and the network function 26 exchange feedback (FB) and negotiation (NEG) messages, for example, to control the bitstream optimization process. The Rx ES 29 forwards the reconstruction to the Rx UE 28 at the network layer, which then forwards the corresponding output (O) to allow the user to terminate the process. u ), provides a reconstruction (REC) to the user interface 14 in the IO layer.
[0255] At the model layer, the common model suite 40 receives input (INP) from the Tx ES 24 or the Rx ES 29, or any predicted input (PI) from the scene prediction model 42. Based on this, the instance segmentation module 44 generates segmented input (SI) that is forwarded to the saliency mapping module 47, which generates a saliency map (SM) that is supplied to the conventional codec 46, the prompt generation model 43, and the consistency data model 45. Based on the received saliency map, the conventional codec 46 generates physical input (PI), the prompt generation model 43 generates learned prompts 414 to update (UD) the overall validation model 41 in the prompt database 48, and the consistency data model 45 generates consistency data (CD). The overall validation model 41 uses the generated learned prompts as input (I) to generate and output a reconstruction (REC) of the Rx ES 29. The semantic loss type database 49 is used to output a loss type (LT) to the scene comparator 324 of the encoder 32 at the codec layer.
[0256] 7 shows a schematic flow diagram of the various processes involved in a compression and decompression process (i.e., a main process (MP), an encoding process (ENC), and a decoding process (DEC)) according to various embodiments, with two devices performing the generated compression across a network using feedback from a network function. Note that not all steps are always required, the order of steps may be adjusted, and some steps may be performed multiple times.
[0257] Specifically, the various processes in Figure 7 include a network-related part (i.e., the main process) and two parts related to the proposed codec functionality (i.e., the encoder 32, bitstream 34, and decoder 36 in Figures 5 and 6).
[0258] End Device In some optional embodiments related to the first type of system described above, the data is compressed and stored locally on the first device (e.g., Tx UE 22). In this case, the encoder 32, bitstream 34, and decoder 36 of Figure 5 are substantially the same or are located on the same device (the first device), and in the main process, the first device simply stores the bitstream rather than transmitting it.
[0259] In step S1.1 of the main process, scene observations are generated via sensing (either local to the transmitting device (eg, Tx UE 22) or via an external device).
[0260] Scenes are also generated from prompts. Consider a filmmaker who uses prompts and existing models to create scenes by writing text commands. As an example, an app running on a mobile phone allows users to create videos / music from prompts and models, and then users can share those videos / music with other users.
[0261] Optionally, user input is collected via a UI. At least some aspects of the scene observation data are forwarded to a transmitting edge server (e.g., Tx ES24). Whether the entire scene observation or only some aspects thereof (e.g., specific objects) is forwarded depends on several factors, including the hardware available at the transmitting device, e.g., whether learned prompts can be generated locally, whether the transmitting edge server is available with low latency, etc. These considerations are similar to those related to split rendering, as discussed, for example, in the 3GPP® TR26.803 study on 5G media streaming extensions for edge processing.
[0262] In step S1.2, the sending edge server implements an encoder to generate a bitstream representing the scene observation. The sending edge server negotiates with a remote party (e.g., a network function) to derive a bandwidth objective for the generated bitstream, which is passed to the encoder as an additional input. The network function, for example, provides the available bandwidth for transmitting the bitstream. The sending edge server instructs the encoder to generate a bitstream that matches this objective, using an in-loop decoder and a scene comparator to estimate its semantic loss by comparing the reconstruction with the original scene observation. The encoder generates further bitstreams using different settings and similarly calculates their semantic loss. The main difference between these bitstreams is that they use learned prompts for more or fewer objects (the remaining objects are represented by physical inputs). This results in a smaller or larger (i.e., more compressed or less compressed) bitstream, respectively.
[0263] For example, if a bitstream slightly larger than the originally set bandwidth limit can obtain a much lower semantic loss, the encoder flags this to the sending edge server, which then negotiates with the network capabilities to be temporarily allocated slightly more bandwidth. This occurs, for example, when the input contains many rare or unique objects, which are associated with a higher semantic loss.
[0264] Conversely, if the input contains many common objects, the encoder can achieve very low semantic loss even at very low bandwidth, and therefore flag this to the network function, freeing up unnecessary bandwidth resources for other processes. The amount of additional bandwidth that is acceptable to apply to achieve a given reduction in semantic loss depends on user input or operator policy.
[0265] In step S1.3, the bitstream is stored in a database and / or transmitted over a network to a receiving edge server (e.g., Rx ES 29). The receiving edge server implements a decoder (see below) to decode the bitstream. The decoder generates a reconstruction in a format usable by the receiving edge server and the receiving device (e.g., Rx UE 28).
[0266] In step S1.4, the receiving edge server forwards the decoded (decompressed) bitstream to the receiving device for generating the user output.
[0267] Based on the generated learned prompts (available to both the sending and receiving edge servers), the global effective model is updated in step S1.5 and used by both the encoder and decoder instead of the base version of the shared reconstructed model.
[0268] The decoder stores the learned prompts it receives from the bitstream in a prompt database for future use. This is especially true if the encoder includes consistency data (e.g., the guide images mentioned above) to aid in the reconstruction of a given learned prompt. In that case, the reconstruction of that object generated using the consistency data can be stored along with the learned prompt. The next time the same object is sent, the sending edge server only needs to generate and transmit the learned prompt and little or no consistency data, thereby saving computing and bandwidth resources.
[0269] The prompt database is optionally persisted beyond the time frame of a single interaction, allowing learned prompts (and any necessary consistency data) and their semantic descriptions to be stored for future interactions in which the same object is observed. Storage can occur locally on the end-receiving device or in a local / proximate edge server, such as the edge server of a content delivery network or video streaming service. This introduces several optional network-side additions to the method, including a side channel for the prompt database to efficiently transmit specific types of input (e.g., videos of a specific genre) and decoders that fetch from the associated prompt database once input of a matching genre is detected; a general object database that is part of the standard such that all decoders conforming to that standard will recognize all these learned prompts in a known way; or namespaces provided to indicate which database a learned prompt belongs to, with certain inputs (e.g., videos) requiring the use of multiple namespaces.
[0270] In step S2.1 of the encoder-related process, an encoder (e.g., encoder 32) receives scene observations as input. The encoder uses an instance segmentation model to classify the scene observations into bounded instances representing linked objects. Instance segmentation is the process of detecting connected regions in an image and assigning a category to each connected region. Two or more regions may receive the same category label, but they constitute different instances of the same object category (e.g., multiple people) and are therefore individually identifiable. The instance segmentation model classifies segmented objects according to their ease of reconstruction via a basic version of a shared reconstruction model.
[0271] Calculating ease of reconstruction can occur simply (i.e., by using an in-loop decoder to create a reconstruction of the object in question, comparing it to observations of the scene, and calculating a semantic loss) or via a set of heuristics (e.g., general objects are estimated to be easier to reconstruct than specific people).
[0272] In step S2.2, the encoder further uses the saliency mapping to estimate the saliency of the segmented objects. For example, the saliency of objects in an image varies depending on their type and location (e.g., a human face in the foreground is the most salient). In addition, saliency is also calculated within objects (e.g., facial keypoints). The obtained saliency data is stored in a spatially or temporally resolved format to match the format of the scene observation, in order to generate a saliency map.
[0273] In step S2.3, for selected observations (e.g., those with high ease of reconstruction by the shared reconstruction model and / or lower salience), the encoder uses the prompt generation model to locally generate potentially appropriate learned prompts (LPs) representing those objects and place them in the bitstream. The encoder also generates global consistency data (CD). In this way, easy and / or unimportant objects are placed in the bitstream with minimal size (i.e., no per-object consistency data, only learned prompts). What levels count as "lower" salience and "higher" ease of reconstruction depends on the bandwidth objectives the encoder is directed to achieve and / or user, system, or operator policies.
[0274] Global consistency data can take several forms, such as data to ensure spatial consistency of reconstructed objects within an image. It could be a segmentation mask indicating where objects should be reconstructed. If two or more objects are reconstructed from different learned prompts, the consistency data includes a partially masked input (e.g., a masked image). This indicates that the masked regions correspond to the parts predicted by the first (text) prompt, while the unmasked regions correspond to the parts predicted by the second (text) prompt. To fill in the entire image using multiple prompts, a decoder (see below) can be applied recursively.
[0275] Instead of dealing with multiple objects, an alternative is to use instance segmentation on the intermediate images resulting from the prompt generation output. Assume that a given category instance is segmented in the source image. A prompt capable of synthesizing this category instance can be generated alone or in an appropriate context that matches the context of the source image. An intermediate image is now synthesized for this single instance only. Object instance segmentation is then applied again, this time to the intermediate synthesized image, to isolate the segment that will be placed in the final image. The segment needs to be cropped, transformed, and scaled to best fit the segment in the source image. To determine these parameters, the set-theoretic union (IoU) metric between the source image segment and the intermediate image segment is maximized. The cropping, transformation, and scaling parameters need to be sent along with the (text) prompt. In this scenario, instance segmentation also needs to be performed on the decoder side.
[0276] Another form of global consistency data is data to ensure the temporal consistency of video generated from prompts, including a picture of the relative motion between the scene and the camera. For example, this could take the form of an estimated decay time (number of frames) for which the prompt is expected or known to be effective. The prompt can optionally be used past this decay time, but this tends to increase semantic loss. This is similar to working with keyframes and infills between them, but with the addition of an estimated decay time.
[0277] Another related approach is to use the aforementioned optical flow concepts to train a "motion-aware" branch of the reconstruction model, which is then used to predict errors in the interpolation branch. In this case, the consistency data consists of the motion-aware branch.
[0278] If the objects in the scene only follow the camera motion, and the camera only rotates or zooms, no parallax occurs, and the (text) prompt remains constant over time, global camera motion parameters are added to the bitstream to compensate for the movement of the image parts synthesized from the text prompt. Temporal color / light changes can also be transmitted. This technique is applicable to specific background categories, such as mountains in the background, streets with houses in the background, forests in the background, etc.
[0279] A further form of global consistency data is data to ensure agreement between different modes of scene observations regenerated from prompts (e.g., audio matching video frames). This may involve matching the "decay times" of audio and video prompts (as described above). In some scenarios, it is possible to train a reconstruction model that reconstructs both audio and video from a shared latent space, and then design prompts for that latent space to handle temporal matching natively. Optionally, only one mode (e.g., video only) uses generative compression, with consistency data forcing the audio to use a conventional codec, and then providing a simple timeline for video prompts to match the reconstructed audio.
[0280] Another form of global consistency data is data for resolving ambiguities in audio input. For example, a learned prompt may be associated with two overlapping voices. In this case, it is more efficient to have multiple learned prompts (e.g., Voice A, Voice B) and specify the degree of overlap via consistency data.
[0281] In step S2.4, for the remaining segmented observations, the encoder is configured to perform at least one of generating learned prompts (tolerating higher expected semantic loss), using a conventional codec to generate physical inputs (PIs) representing those objects (tolerating higher bandwidth usage), and generating prior object consistency data that serves as a correction to the learned prompts.
[0282] The consistency data per object can take several forms. For example, it uses the guided image technique described above, which acts as a correction for the reproduction of the image from learned prompts. To ensure color consistency in the image, the consistency data includes calculated color transformation vectors that modify the color of a given object or spatial region of an object to better match the scene observation. The same is done using texture transformations to render complex textures. Since the color and texture transformation parameters are specified per object, the rate-distortion formula becomes: L total ≡L conventional +Σ i L i where the sum is over all regions i generated by the text prompt. The object / region i generated by the text prompt is L i (k;p)=k i R i (p)+D i (p). Parameter k i Note that the rate / distortion balance specified by can be chosen differently depending on the object category. The color and texture transformation parameters provide options to make the synthesized image more realistic by applying a per-object color transformation or adding / modifying spatial texture. For example, the appearance of a "wooden chair" generated from a text prompt can be made more like a real chair by increasing / reducing the spatial frequency of the synthesized texture in the object. A simple parameterized high-pass or low-pass filter is used to achieve these effects. The distortion term D for object / region i is i (p) can be divided into color terms and texture terms. D i (p)≡D i,colour (p colour )+D i,texture (p texture ), where p is a vector of combined color and texture parameters. The color distortion term D i,colour (p colour ) can be calculated using the distance function between the distributions. The texture distortion term D i,texture (p texture) can be calculated by comparing the spatial frequencies between the synthesized texture and the observed image of object / region i.
[0283] The reconstruction from the learned prompt in step S2.3 serves as a predictor for at least part of the input, and a conventional codec is used to compress the residual input (a conventional codec operates on the entire residual input rather than a specific target). This is similar to the extension of the coding using a reference layer based on a diffusion model and a conventional extension layer.
[0284] The reconstruction using the learned prompts in step S2.3 serves as a reference frame used by conventional codecs to encode other frames (particularly applicable to video).
[0285] In step S2.5, the encoder optionally finishes, producing a bitstream containing the saliency map, the learned prompts, any necessary physical input or consistency data, and optionally an estimated loss score.
[0286] In step S2.6, the encoder uses an in-loop decoder that accesses the shared reconstruction model to generate an initial reconstruction using the initial version of the saliency map, the learned prompts, the physical input, and the consistency data. This initial reconstruction is passed to the scene comparator. In subsequent runs of the encoder, this step uses the global valid model rather than the shared reconstruction model to take into account the learned prompts that have already been sent.
[0287] In step S2.7, the scene comparator compares the initial reconstruction with the scene observations and generates or updates loss scores (both semantic loss and traditional compression loss if physical inputs are used). For semantic loss, the scene comparator aims to identify objects that contributed significantly to the semantic loss score. For such objects, the encoder attempts to generate consistency data (as described above) to lower the loss score.
[0288] In step S2.8, the encoder repeats, generating new learned prompts, consistency data, and physical inputs until the desired loss score is achieved.
[0289] Finally, in step S2.9, the encoder places, for example, learned prompts, consistency data, physical inputs, saliency maps, or other encoded data into the bitstream.
[0290] The decoder process begins at step S3.1, where a decoder (e.g., decoder 36) receives the bitstream as input. The decoder uses a reconstruction order module to calculate the optimal reconstruction order to generate reconstructions from learned prompts and physical inputs. This takes into account both the estimated speed of reconstructions and the saliency from the saliency map. The RO module implements several heuristics to calculate the reconstruction order.
[0291] For example, a heuristic may state that, using a given receiving edge server, the physical input can be reconstructed faster than the learned prompts, and therefore the learned prompts should be executed first to speed up the overall reconstruction.
[0292] A second example is to reconstruct more salient objects (from the saliency map) first (e.g., reconstructing the foreground before the background).
[0293] A third example is to initially use only the learned prompts and then add consistency data, thereby reconstructing the subject with both the learned prompts and their associated consistency data, which results in a faster initial reconstruction (although potentially with greater semantic loss).
[0294] If necessary, known stitching and smoothing techniques are used to hide transitions between different versions of the reconstructed object.
[0295] In step S3.2, the decoder uses a shared reconstruction model (or, in later iterations, a globally valid model) to reconstruct objects from the learned prompts received according to the reconstruction order. For example, one or more shared reconstruction models are used depending on the type of data. The shared reconstruction model used is based on the preferences or profile of the receiving device. The reconstruction model refers to one or more chained reconstruction models. For example, a first reconstruction model reconstructs / generates general data such as audio / video from the received prompts, and a second reconstruction model uses the general data as input to generate data such as personalized audio / video. A conventional codec is used to generate the output for the received physical input. A saliency map is used to ensure correct placement of the reconstructed objects to generate a global reconstruction of the scene observation. Additional global consistency data is used (for example) to ensure the temporal stability of the reconstructed output.
[0296] Global coherence data may consist of several elements depending on the format of the scene observation, such as a simple set of masks for images showing the placement of objects (to ensure spatial coherence), a temporal version for audio / video of the above, and / or information needed to increase the coherence of the video reconstructed from the prompts (e.g., a "decay time" during which the prompts are expected or known to be effective, after which continued use of the prompts increases semantic loss).
[0297] In step S3.3, the decoder passes the reconstructed output in a usable format to a downstream function (eg, a receiving device).
[0298] According to an alternative prediction embodiment, in addition to comparing the initial reconstruction to the scene observation, the encoder also, or instead, compares it to a predicted version of the scene observation generated by the scene prediction model 42 of Figure 6. In such a case, the above process of Figure 7 is modified as follows.
[0299] In step S2.1, in addition to or instead of receiving scene observations as input, the encoder receives a predicted scene from a scene prediction model for a short distance into the future. The time period into which the prediction is made is set using at least one of data from the scene prediction model (which outputs an estimated time period for which the prediction is expected to be valid), a required delay input from the network function (a lower required delay implies predicting scene observations for a longer period into the future), an acceptable semantic loss (a lower loss implies a shorter prediction period), and user, system, or operator policy. Additionally, the time period for which the prediction is expected to remain accurate (i.e., the "decay time") is added to the consistency data.
[0300] In step S2.6, the scene comparator compares the initial reconstruction with the scene observations in addition to, or instead of, comparing it with a predicted scene some time in advance.
[0301] In step S2.9, the encoder generates the prompts, physical inputs, consistency data, and initial reconstruction of this predicted scene as outlined in the main method and passes it to the sending edge server, which transmits it to the receiving edge server, which can now create a reconstruction of the scene observation with zero or negative delay by ensuring that the receiving device displays the appropriate predicted reconstruction at the time of its predicted occurrence.
[0302] Optionally, the sending edge server waits until the actual scene observations for the predictively generated time period are available. It then instructs the encoder to use a scene comparator to compare the actual scene observations with the predicted reconstruction generated from the predicted bitstream. If there are large differences between the predicted reconstruction and the actual scene observations (i.e., high semantic loss scores), the sending edge server instructs the encoder to generate additional consistency data that serves as a correction factor for those differences. This consistency data can be per object or global. The sending edge server sends the consistency data to the receiving edge server. The receiving edge server applies this correction data to update the reconstructed objects and, optionally, uses smoothing techniques to mask transitions.
[0303] Alternatively, the receiving edge server may implement a scene prediction model and make predictions based on already received reconstructed scene observations. In that case, the receiving edge server may optionally choose to send the predictions back to the sending edge server for comparison with actual scene observations as they become available. The sending edge server then generates and sends prediction consistency data appropriate for prediction reconstruction at the receiving edge server.
[0304] FIG. 8 shows a schematic of the processing steps and output of an exemplary embodiment.
[0305] In step S161, a transmitting device (e.g., UE) initially observes a scene (the top large image in FIG. 8) that includes an easy-to-reconstruct forest area (F) and an unknown person (UP) that cannot be reconstructed. Therefore, the forest area is classified (CLASS) as "easy to reconstruct." Then, in step S162, the transmitting device generates prompts for both the easy and difficult objects, transmitting the easy object (e.g., "forest") first, as shown in the top right small image in FIG. 8.
[0306] In step S163, the sending device uses network feedback to handle difficult subjects with increasing levels of semantic loss and optimize bandwidth (OPT-BW). To achieve this, it first generates a generic person (GP), as shown in the small image in the middle of Figure 8. Then, it obtains a corrected image (CORR-IM) with low to moderate loss and high computational load by comparing the best prompt with a physical image (PHY-IM) of the actual observed scene without semantic loss (the small image in the bottom of Figure 8).
[0307] In step S164, the sending device sends the best prompt along with a small physical correction image to the receiving device.
[0308] Similar steps can be performed at the receiving end to obtain prompts and reconstruction models to reconstruct the content, in this case, an image. Different recipients are configured with different reconstruction models so that the reconstructed content better matches the recipient's preferences. For example, the original forest is reconstructed as a jungle for a first person living in Brazil, as a pine forest for a second person living in Norway, and as a beach for a third person living along the coast. For example, some aspects of the reconstructed person (e.g., skin color, eye shape, mouth shape, etc.) are also reconstructed to better match the recipient's preferences. This also implies that one or more correction images may be available and / or may need to be transmitted.
[0309] Through the above-described embodiments, applications such as metaverse, video conferencing, or video streaming that utilize the communication infrastructure can configure and use the communication infrastructure for optimized performance. This configuration and use is performed through a Touch Service Manager (TSM), which coordinates communication with the underlying network. The 5G (or 6G) TSM resides in the 5GS (or 6GS). The 6G TSM interacts with the 5G TSM. Furthermore, configuration is controlled by a policy that includes configuration items for each haptic device (TD) in each haptic edge (TE). Each time a new TE (TD) joins a (new) (metaverse) communication session, the application adds an entry to the policy corresponding to the new TD or TE. The TSM and / or application also adjust the preferences of the sending / receiving entities and adapt the coding / communication parameters, e.g., the model used, accordingly. For example, the TSM distributes the policy of the new TD (TE) to all existing TDs (TEs) already involved in the (metaverse) communication session. Additionally, the TSM distributes a policy to the new TD (TE) that contains entries for all existing TDs (TEs) already involved in the (metaverse) communication session. The configuration can be a one-time configuration or a metaverse session configuration for a metaverse session between multiple TEs (e.g., multiple users (A, B, ..., i, ...)).
[0310] Furthermore, the configuration includes policies specifying QoS targets depending on, for example, the number of users, relative delays, the need for continuous monitoring of delays between TEs, as well as parameter update rates as described in other embodiments, the delay requirements of each of the TDs within the TE, the need for QoS equalization, etc., and, if applicable, the compression scheme or model and / or prediction model for each TD within the TE is adapted accordingly or the TSM is able to deploy models or compression models to other TDs / TEs in the communication session.
[0311] Similarly, the communications infrastructure informs and / or configures (metaverse) applications of communication parameters.
[0312] Note that the transmitting and receiving devices may be separate or nearby, and therefore the transmitting and receiving TEs are co-located. TEs are not necessarily required.
[0313] A unicast communication flow requires maintaining a unicast flow for each sensing TD (N devices) for each actuator / rendering TD (M devices). As N and M increase, efficiency may decrease. A more efficient approach is a multicast approach, where each sensing TD multicasts a flow and distributes it to each of the subscribed rendering TDs. Although this involves N multicast flows, it is still important to consider that multicast flows arrive at different rendering TDs / TEs at different moments, and TDs / TEs receiving the multicast flow earlier may, for example, use a compressed model of the sending TD / TE, while TDs / TEs receiving the multicast flow later may, for example, require a less compressed model.
[0314] Furthermore, the proposed compression system architecture may be enhanced or used to enhance, for example, next-generation real-time communications or multicast and broadcast services. For example, 3GPP® specification TR23.700-87 v1.0.0 describes 5G system architecture enhancements for next-generation real-time communications, including IP Multimedia Subsystem (IMS) network architecture enhancements necessary to support AR telephony for various types of AR-capable UEs. IMS procedures, including signaling and media processing, must be modified to support AR telephony. Solutions #8 and #9 in the TR23.700-87 specification address these architectural enhancements. TR23.700-87 concludes that the data channel architecture is used as the baseline for supporting AR telephony. If the UE requires network support for media rendering, the architecture and procedures specified in Solution #9 are used. Otherwise, if the UE can perform media rendering without network support, the procedures specified in Solution #8 are adopted as the baseline for the terminal rendering process. The IMS architecture is enhanced accordingly, as described in Annex AC.9 of TS 23.288. In particular, steps 2 and 3 describe an AR media rendering negotiation procedure, in which in step 2, UE-A requests network media rendering based on the status of power, signal, computing capability, internal storage, etc., and in step 3, UE-A completes the AR media rendering negotiation with the AR AS. In the subsequent step 9, UE-A transmits AR data to the MF, and the MF also receives instructions from the AR AS, based on which the MF performs AR media rendering according to the negotiation result in step 3.
[0315] In one embodiment, the system and functionality described in Solution #8 in TR 23.700-87 are extended to support at least some of the aforementioned embodiments. Figure 6.8.2-1 in TR 23.700-87 describes a communication flow between two UEs, including three procedures: (1) an IMS multimedia telephone call, (2) establishing a bootstrap data channel (DC), and (3) establishing an application DC. In a further embodiment, the system and functionality of Solution #9 in TR 23.700-87 (leading to Appendix AC.9 of TS 23.228) are extended to support at least some of the aforementioned embodiments. Figure 6.9.2.2-1 in TR 23.700-87 describes a communication flow between two UEs with a network rendering process. In this process, the AR media processing network function (ARMF) is responsible for transmitting AR communication media and media rendering functions. It includes the AR rendering logic function, which controls the application-based rendering logic of AR communication, and the AR media processing function, which includes a vision engine and a 3D rendering engine, which establish a spatial map and render the scene, virtual human model, and 3D object model according to the field of view, pose, position, etc. transmitted from the UE using a data channel. For example, referring to Appendix AC.9 of TS23.228, the UE-A has the capability of data compression and acts as a sender, while the MF has the capability of data decompression and acts as a receiver. These entities also negotiate with the AR AS for a general compression model, a personalized compression model, or a personalized policy. The compression policy determines the amount of semantic loss allowed, the desired compression ratio, the desired computational overhead, the desired storage overhead, and the desired communication overhead.
[0316] UE-A has a compression model that, given an image, determines whether a person, say person Y, appears in the image. The compression model then converts the image into a prompt person Y. The compression model also contains rendering information about person Y on the image, such as the location in the image where person Y will appear when the prompt is decompressed, the size of the person on the decompressed image, as well as the orientation of person Y.
[0317] UE-A has a compression model that, given an image, determines whether a person, say person Y, is in the image. The compression model then converts the image into a prompt person Y. The compression model also contains rendering information about person Y on the image, such as the position in the image where person Y will appear when the prompt is decompressed, the size of the person on the decompressed image, as well as positional information such as the orientation / rotation of person Y. This information can be used by the receiving party (MF) to obtain an image of person Y in a specified orientation and render it at a specified size in a specified location.
[0318] UE-A is also transmitting the movement of an object, e.g., person Y, as in the previous example. In this case, the compression model then converts the image into a prompt person Y. The compression model also includes rendering information about person Y on the image, such as the position in the image where person Y will appear at time t when the prompt is decompressed, the size of the person in the decompressed / rendered image, and position / movement information such as the orientation / rotation / movement direction / velocity of person Y in the decompressed / rendered image. This information can be used by the receiving party (MF) to obtain an image of person Y in a specified orientation and render it at a specified size at a specified location. If the MF and UE-A are also able to determine a given communication delay, for example, by using the protocol described in clause 4.4.4 of TS 26.522, which uses RTP header extensions for in-band end-to-end delay measurement, the MF determines that the received compressed data arrived with delay T, and therefore the received compressed data is predictively decompressed. This allows the receiving entity / decompressor / renderer to compensate for the communication delay by using the movement information (e.g., position(t) + velocity) instead of the position indicated as sent at time t. * This can mean rendering the person at their updated position at time t, taking into account the received velocity (T, where velocity includes the received direction of movement). Similarly, if person Y is approaching UE-A, this can also imply that the predictively decompressed data of person Y is rendered at a larger size than the size sent at time t.
[0319] In general, predictive decompression / predictive rendering refers to techniques that enable an entity (e.g., a receiving entity) to reconstruct various types of data, such as images, video, or audio, from a compressed data stream that contains semantic information about objects in a scene, such as their identity, location, size, orientation, and motion. The receiving entity uses a decompression model that can generate realistic data about the objects based on description prompts extracted from the compressed data stream. The decompression model also accounts for communication delays between the sender and receiver and adjusts the rendering of the objects according to the transformations that they will undergo upon display. These transformations include changes in location, size, orientation, or other aspects that affect the data. In this way, the receiving entity can generate a smooth and accurate representation of the scene without requiring high bandwidth or storage capacity.
[0320] The above embodiments are also applicable to architectures that use split rendering. Split rendering means that the heavy rendering process is performed by a computationally resource-intensive device (e.g., a haptic edge (TE), e.g., an edge server), and subsequent user-specific or device-specific light rendering is performed locally, e.g., in the haptic device (TD). Split rendering allows offloading computation to the TE while keeping the TD simple. When a split rendering architecture is used, predictive models (e.g., those described in embodiments related to model registration) may be executed in the TE.
[0321] One or more prediction models may be run per user. One of these prediction models may, for example, utilize a volumetric video (VV) representation of the user for prediction, allowing the user to be represented in a photorealistic manner. A TE may run multiple prediction models, e.g., one per user, requiring synchronization of data streams from multiple users / TEs with multiple tactile sensors in different locations. When a TE runs prediction models for multiple remote users, the TE renders a combined, time-synchronized, and time-predicted representation, e.g., a VV-rendered representation, of all users involved in the metaverse session. Time synchronization means that the generated data streams are aligned, i.e., follow a common clock. A received, time-synchronized data stream (from another remote TE) arrives delayed by a given time Delta compared to the local clock of the local TE. Thus, time-predicted means that the representation is predicted by a time Delta into the future to synchronize with the local clock of the local TE. This time Delta varies depending on the delay or communication parameters between each pair of remote TEs.
[0322] The TE consumes information about the local user, such as the local rendering device (e.g., TD) associated with the user in the local environment. For example, the TE uses the height, position, and orientation of the VR / AR glasses the user is wearing. Using this information, the TE can derive a TD-specific representation of the environment that can be used by the user's rendering device (TD). For example, this representation is a 2D representation of the volumetric video rendering at the edge server from the perspective of the user's rendering device (e.g., VR / AR glasses).
[0323] In this environment, the local TE requires the communication system to allocate communication resources so that TDs in the environment can continuously provide input related to their posture, etc. This should be done in a time-deterministic manner. For example, 5GS allocates H resource blocks every m milliseconds to transmit data related to head posture. When this is done, the delay when the TE receives the user's posture is T = T + m + T, where T is the processing delay from sensing the posture to being able to transmit the value, m is the delay due to discrete measurements, and T is the propagation time from the TD to the TE. This involves deterministic allocation of uplink communication resources, for example, through a mechanism similar to semi-persistent scheduling, so that the TD can continue to transmit input in a reliable and time-deterministic manner. This requires the TE to take T into account, for example, by running a posture prediction model that can provide past samples and predict the user's actual current posture.
[0324] Similarly, TE requires that the communication system allocates communication resources so that TDs in the environment can continuously receive TD-specific representation input generated at the TE. TE must also consider the transmission delay at the local TE, i.e., T = Trending + m + Tflight, where Trending is the time required to update the rendering at the local TD after receiving the data, m is the delay due to the discrete transmission time, and Tflight is the propagation time from the TE to the TD.
[0325] In this embodiment, the delay in uplink communication from the TD to the TE, including information about the TD (e.g., attitude), as well as the delay in downlink communication and local rendering, are some of the communication parameters taken into account when synchronizing data streams from other users at other locations and / or when applying predictive models.
[0326] One possible embodiment enabling a system with split rendering is as follows: An edge server (MF in Annex AC.9 of TS 23.228) receives a compressed data stream from a sending entity, such as a UE-A or a remote server, and performs partial decompression based on semantic information (e.g., prompts) in the data stream. For example, the edge server decompresses some of the semantic objects, such as faces, text, or symbols, that are more complex, require higher resolution, or require a more complex decompression model, and leaves the remaining objects in compressed form. The edge server also performs some preprocessing tasks, such as cropping, scaling, filtering, or enhancing the decompressed objects, according to the rendering information in the data stream. The edge server then transmits the partially decompressed data stream to an end device, such as another UE, which performs final rendering before presentation.
[0327] In another embodiment, the edge server performs predictive decompression based on semantic and rendering information in the data stream and proactively transmits the predicted decompressed data to the end device. Predictive decompression uses models or algorithms to predict the future state or movement of semantic objects, such as position, orientation, shape, color, or texture, based on the semantic objects' past or current state or movement, or other contextual information such as the user's gaze, head pose, gestures, or actions. Predictive decompression also takes into account the latency, bandwidth, or reliability of the communication channel and adjusts the accuracy, frequency, or granularity of the predictions accordingly. Predictive decompression aims to reduce perceived rendering delay or improve visual quality on the end device.
[0328] The edge server compares the predicted decompressed data with the actual decompressed data received from the sending entity at a later time and calculates the error or difference between them. The edge server then transmits correction data representing the error or difference to the end device. This correction data can be used by the end device to correct or update a previously rendered image or scene. The edge server has a policy or configuration that determines when and how often to send correction data depending on factors such as an error threshold, network conditions, user feedback, or system load. For example, the edge server will send correction data only if the error exceeds a certain value, if the network has sufficient capacity, if a user reports a low level of satisfaction, or if the system has spare resources.
[0329] The end device receives the predicted decompressed data from the edge server and renders it to the user's display or view. The end device also receives correction data from the edge server and applies it to previously rendered images or scenes to modify or improve visual quality or fidelity. The end device has policies or configurations that determine how to apply the correction data depending on factors such as the rendering mode, user preferences, device capabilities, or application requirements. For example, the end device may apply the correction data immediately, after a certain delay, only for specific semantic objects, only when the user is not looking, or only when the application allows it.
[0330] In a further embodiment, before the edge server and the end device perform partial decompression and final rendering, respectively, a negotiation phase occurs between the management entity, the edge server, and the local user equipment to determine under what conditions the operations performed at the edge server and the operations performed at the local user equipment will be performed. The negotiation phase involves exchanging information such as each entity's capabilities, resources, models, preferences, or policies, and agreeing on the appropriate allocation of tasks and parameters of the data compression / decompression process, e.g., policy configuration. For example, the management entity receives requests from the edge server and the local user equipment to access or provide specific semantic objects, decompression models, rendering information, or personalized content, and grants or denies the requests based on resource availability, security, privacy, or cost. The management entity also coordinates communication and synchronization between the edge server and the local user equipment and monitors service quality and user experience. The negotiation phase can be performed periodically, dynamically, or on-demand in response to changes in the scene, network, user, or system. To enable a consistent and coherent user experience across multiple user devices, it is necessary to negotiate settings that are common to a group of user devices, such as the placement, orientation, scale, or viewpoint of a scene or semantic object. For example, when several people interact at the same location and use different user devices to access the same scene or content, they may wish to have a shared view of the scene or content so that they can communicate and collaborate effectively. Alternatively, when multiple people interact from different locations and use different user devices to access the same scene or content, they may wish to have synchronized views of the scene or content to achieve a sense of presence and immersion.
[0331] In one embodiment, negotiation of settings between user equipment is facilitated by a management entity, such as a cloud server, edge server, or peer device, that acts as a mediator or coordinator for a group of user equipment. The management entity receives information from each user equipment regarding its capabilities, resources, model, preferences, or policies and uses this information to determine optimal or acceptable settings for the group of user equipment. For example, the management entity calculates the average, minimum, maximum, or median values of parameters related to settings, such as resolution, frame rate, latency, or bandwidth, and selects the setting that best matches or satisfies these values. Alternatively, the management entity uses voting, ranking, weighting, or negotiation mechanisms to determine the setting that is most preferred or agreed upon by a majority or all of the user equipment. The management entity also considers application requirements, network conditions, user feedback, or system performance when selecting settings. The management entity then transmits the selected settings to each user equipment and instructs each user equipment to adjust its operations, such as compression, decompression, rendering, or display, according to the selected settings. The management entity also monitors the user experience and quality of service and updates the configuration as needed.
[0332] In another embodiment, a split-rendering approach enables negotiation and / or application of settings between user devices, where a common or synchronized view of a scene or content is computed by an edge server and then distributed to the user devices. The edge server performs rendering-intensive tasks, such as calculating geometry, shading, lighting, or occlusion, to generate high-quality images or videos of the scene or content. The edge server also applies settings common to a group of user devices, such as camera position, orientation, or field of view, to create a consistent or coherent view of the scene or content. The edge server then sends the images or videos to the user devices and instructs each user device to perform lightweight rendering tasks, such as post-processing, filtering, or warping, to adapt the images or videos to the user devices' specific characteristics or preferences, such as screen size, resolution, aspect ratio, or color scheme. The edge server also receives feedback from the user devices and adjusts settings as needed.
[0333] One possible embodiment that allows the end device to control the (predictive) decompression process is as follows: The end device sends instructions to the edge server about the semantic objects it wants to be (predictively) decompressed by the edge server and the semantic objects it prefers to decompress locally (predictively). The instructions include parameters specifying the duration, condition, desired resolution, or priority of the decompression request. For example, the end device may indicate that it wants the edge server to decompress only semantic objects relevant to the user's focus, attention, or interaction, and that it can handle the decompression of background or peripheral objects. Alternatively, the end device may indicate that it wants the edge server to decompress resource-intensive semantic objects, such as high-resolution textures, animations, or effects, and that it can manage the decompression of simpler or lower-resolution objects. The edge server then performs partial decompression according to the instructions from the end device and sends the partially decompressed data stream to the end device. In a further embodiment, the end device receives the partially decompressed data stream from the edge server and uses its own decompression model to generate the remaining semantic objects from description prompts in the data stream. The end device also uses its own rendering engine to combine the decompressed objects with the rendering information and display the reconstructed scene on a screen or other output device. The end device takes into account the user's preferences, profile, or context to customize the rendering process and generate personalized content. For example, the end device adjusts the color, brightness, contrast, or sound of the scene depending on the user's settings or environment. The end device also modifies the appearance, behavior, or interaction of some of the semantic objects depending on the user's interests, goals, or feedback. For example, the end device changes the clothing, hairstyle, or facial expression of a virtual character, or adds or removes some elements or effects in the scene based on the user's input or response. Different embodiments can be combined with each other or used independently as needed to address requirements and / or missing capabilities.
[0334] In summary, an apparatus and method for data compression / decompression have been described, in which input observation data is classified into semantic object types according to one or more criteria to obtain compressed data, and compression techniques are applied to the object types based on the ease of reconstruction via a compression model. Classification is performed based on the performance of the compression model. For data objects belonging to a data object type that can be properly reconstructed, appropriate description prompts are generated, e.g., based on text reversal techniques. Furthermore, an apparatus and method for personalized data compression / construction / decompression / reconstruction are described, in which an encoder / decoder uses appropriate personalized description prompts (or programs), a personalized reconstruction model, and / or personalized policies to generate personalized content according to a user's preferences and / or profile. While the present invention has been described in the context of a virtual space such as the Metaverse, its application is not limited to such types of operations. Other systems, such as AR / VR, can also benefit from the present invention. Low-latency systems, such as industrial IoT systems, can also benefit from the teachings of the present invention and its embodiments.
[0335] Furthermore, the present invention can be applied to various types of UE or terminal devices, such as mobile phones, vital signs monitoring / telemetry devices, smart watches, detectors, vehicles (for vehicle-to-vehicle (V2V) communication or more general vehicle-to-everything (V2X) communication), V2X devices, Internet of Things (IoT) hubs, IoT devices including low-power medical sensors for health monitoring, medical (emergency) diagnostic and treatment devices for hospitals or emergency personnel, virtual reality (VR) headsets, etc.
[0336] Other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the singular form of an element does not exclude a plurality. A single process or other unit fulfills the functions of several items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage. The above description has detailed certain embodiments of the present invention. However, no matter how detailed the above description appears in the text, it will be understood that the invention can be practiced in various ways and is not limited to the disclosed embodiments. It should be noted that the use of a particular term in describing a particular feature or aspect of the present invention does not imply that the term be redefined herein to include any particular characteristic of the feature or aspect of the invention with which the term is associated. Furthermore, the phrase "at least one of A, B, and C" should be understood as disjunctive, i.e., "A, B, and / or C."
[0337] A single unit or device may fulfill the functions of several items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.
[0338] The described operations as shown in the above embodiments may be implemented as program code means of a computer program and / or as dedicated hardware of associated network devices or functions, respectively. The computer program may be stored and / or distributed on a suitable medium, such as an optical storage medium or a solid-state medium, provided together with or as part of other hardware, or distributed in other forms, such as over the Internet or other wired or wireless telecommunications systems.
Claims
1. 1. An apparatus for data compression, comprising: Identifying or segmenting the input data An apparatus that generates instances related to identified semantic objects according to one or more criteria and applies compression techniques to the identified semantic objects.
2. The apparatus of claim 1 , wherein the apparatus generates description prompts for semantic objects that belong to a class of data objects that can be represented or reconstructed by compression models and description prompts.
3. Through the compression model, the ease of reconstructing the identified semantic objects; semantic loss requirement, personalization requirements, and Compression Needs The apparatus of claim 1 or 2, wherein the apparatus applies the compression technique based on at least one of:
4. 3. The apparatus of claim 2, further comprising: synthesizing an image using the generated description prompt; identifying semantic objects or performing instance segmentation on the synthesized image; determining at least one of cropping parameters, transformation parameters, and scale parameters; and outputting the determined parameters together with the description prompt.
5. The apparatus of claim 2 , further comprising: performing temporal data association of the generated description prompts over time for frames of input observation data; and determining global motion model parameters for each instance associated with the generated description prompt.
6. a. generating inferential writing prompts suitable to carry sufficient semantic content for a given context; b. developing a writing prompt based on text reversal induced by said data subject; and c. Using conventional compression techniques 3. The apparatus of claim 2, wherein the data object is compressed by at least one of:
7. 7. Apparatus according to any one of claims 1 to 6, which labels compressed data objects according to the compression technique and / or reconstruction model used.
8. 8. The apparatus of claim 1, wherein a compression policy is configured by or negotiated with a communication manager, the compression policy determining at least one of an amount of semantic loss tolerated, a desired compression ratio, a desired computational overhead, a desired storage overhead, and a desired communication overhead.
9. 1. An apparatus for data decompression, receiving compressed semantic objects and applying a decompression technique based on the compression model and a type of the received compressed semantic data object to obtain decompressed data.
10. 10. The apparatus of claim 9, wherein the decompression technique relies on at least one of a general compression model, a personalized compression model, a personalized policy, and a personalized prompt to obtain personalized decompressed data.
11. 11. The device according to claim 9 or 10, wherein at least one of the compressed semantic objects is decompressed based on a shared reconstruction model and description prompts.
12. 12. An apparatus according to any one of claims 9 to 11, configured with a decompression policy negotiated with or configured by a communications manager.
13. The apparatus of claim 9 , wherein the apparatus determines and updates the compression model based on at least one of reconstruction performance, the semantic loss during reconstruction, network performance, and computational resources.
14. 14. An apparatus according to any one of claims 9 to 13, which uses received description prompts for predictive decompression.
15. 10. A transmitting device comprising the apparatus of claim 1, in communication with a receiving device and sharing the compression model with the receiving device, wherein the shared compression model is determined or updated based on at least one of reconstruction performance, transmitter preferences, receiver preferences, network connectivity, and computational capabilities.
16. 16. The sending device of claim 15, wherein the description prompts are generated predictively and decompressed data based on the predicted description prompts is compared to the input observation data to determine a correction factor.
17. 17. The transmitting device of claim 15 or 16, wherein a shared reconstruction model is retrained based on the correction factors.
18. 18. The transmitting device of claim 15, wherein semantic loss is determined based on an instance rate-distortion function, and total loss is calculated as a sum or weighted sum over multiple semantic objects identified in the input observation data, and object loss is composed of object rate and object distortion that depends on object color, object shape, and texture parameters.
19. Negotiating the use of a semantic decompression mechanism between the edge server and the management entity; sending an indication to the edge server indicating a need for split rendering of a particular semantic object; A receiving device that receives the decompressed and compressed data from the edge server based on the instructions.
20. 20. A receiving device according to claim 19, comprising an apparatus according to claim 9 for decompression.
21. 21. A receiving device according to claim 19 or 20, which receives the correction factor from the sending device that received the compressed semantic object, and uses the correction factor to correct the obtained decompressed data.
22. 22. A receiving device according to any one of claims 19 to 21, wherein a shared reconstruction model is retrained based on the correction coefficients.
23. 23. A receiving device according to any one of claims 19 to 22, wherein predicted decompressed data is compared with obtained decompressed data to determine a correction factor, the obtained decompressed data being made available by the transmitting device.
24. A system comprising a transmitting device according to any one of claims 15 to 18 and a receiving device according to any one of claims 19 to 23.
25. 25. The system of claim 24, wherein semantic loss is determined based on an instance rate-distortion function, and wherein total loss is calculated as a sum or weighted sum over multiple semantic objects identified in the input observation data, and object loss is composed of object rate and object distortion that depends on object color, object shape, and texture parameters.
26. identifying or segmenting the input observation data to generate instances related to the identified semantic objects according to one or more criteria; applying a compression technique to the identified semantic objects based on their ease of reconstruction via a compression model; 1. A method for data compression comprising:
27. receiving a compressed semantic object; applying a decompression technique based on a compression model and a type of the received compressed semantic objects to obtain decompressed data, wherein the compression model and / or the compressed semantic objects depend on at least one of a decompression performance requirement, a semantic loss requirement, and a user preference requirement; 10. A method for data decompression, comprising:
28. 28. A computer program comprising code means for generating the steps of the method according to claim 26 or 27 when the computer program is executed on a computing device.
29. 27. A bitstream produced by the method of claim 26, comprising at least one description prompt representing a compressed semantic object belonging to a class of data objects that can be represented or reconstructed by a compression model and a description prompt.