Tokenized compression with task-oriented plug-and-play tokens

WO2026148360A3PCT designated stage Publication Date: 2026-09-24FUTUREWEI TECHNOLOGIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/025544
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-07-16
Filing Date
2026-04-28
Publication Date
2026-09-24

AI Technical Summary

Technical Problem

Existing image and video compression techniques optimized for human visual system degrade machine vision task performance due to sensitivity to small pixel variations, lacking scalability and generalization across diverse tasks and models, and failing to leverage multimodal vision-language models.

Method used

A task-oriented compression framework using a base tokenized visual representation (TVR) with plug-and-play auxiliary tokens that are modular and adaptable across tasks, incorporating visual, audio, and textual inputs, enabling scalable and efficient compression for both human and machine vision tasks.

Benefits of technology

The framework maintains high-quality image and video data reconstruction for human perception while enhancing machine vision task performance, supporting seamless integration with advanced vision-language models and reducing transmission and computation costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2026025544_24092026_PF_FP_ABST
    Figure US2026025544_24092026_PF_FP_ABST
Patent Text Reader

Abstract

A method implemented by a sending device for performing task-oriented image or video compression, the method comprising encoding an input visual signal using a base tokenized visual representation (TVR) framework to generate a discrete indices string; selectively encoding the input visual signal in a quality-enhancement processing stream to generate a continuous latent string to improve at least one of a task performance or a reconstruction of the input visual signal; and transmitting the discrete indices string and, when generated, the continuous latent string to a receiving device.
Need to check novelty before this filing date? Find Prior Art

Description

Atty. Docket No. 4502-86601 (6000767PCT02)TOKENIZED COMPRESSION WITH TASK-ORIENTED PLUG-AND-PLAY TOKENSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U. S. Provisional Application Ser. No. 63 / 845,136 filed on 7 / 16 / 2025, the disclosure of which is incorporated herein by reference.TECHNICAL FIELD

[0002] Disclosed embodiments relate generally to image and video compression, and more specifically to a tokenized compression with task-oriented plug-and-play tokens.BACKGROUND

[0003] Most image and video compression techniques, such as JPEG, H.264, H.265, H.266, Learned Image Compression (LIC), and Learned Video Compression (LVC), are optimized for the human visual system rather than for machine vision tasks. The goal of compression for human perception is to ensure that the reconstructed images and videos are visually pleasing. In contrast, for machine vision tasks such as object detection, recognition, classification, and segmentation, the objective is to maintain acceptable task performance on the reconstructed images or videos. Machine vision models are typically highly sensitive to small pixel variations, which can significantly impact prediction accuracy. Compression methods tailored for the human visual system generally eliminate redundancy in the pixel domain, often resulting in severe degradation of task performance.SUMMARY

[0004] A first aspect relates to a method for task-oriented image or video compression implemented by a sending device. The method includes encoding an input visual signal using a base tokenized visual representation (TVR) framework to generate a discrete indices string, selectively encoding the input visual signal in a quality-enhancement processing stream to generate a continuous latent string to improve at least one of a task performance or a reconstruction of the input visual signal, and transmitting the discrete indices string and, when generated, the continuous latent string to a receiving device.

[0005] Optionally, in a first implementation according to the first aspect, the method further includes selectively generating one or more task-adaptive auxiliary tokens that are modular and independent of the base TVR framework, wherein the one or more task-adaptive auxiliary tokensAtty. Docket No. 4502-86601 (6000767PCT02)provide enhanced semantic representation for at least one specific downstream task, and transmitting the one or more task-adaptive auxiliary tokens along with the discrete indices string and, when generated, the continuous latent string to the receiving device.

[0006] Optionally, in a second implementation according to the first aspect or any implementation thereof, generating the one or more task-adaptive auxiliary tokens comprises receiving a text description of the input visual signal and generating the one or more task-adaptive auxiliary tokens based on the text description.

[0007] Optionally, in a third implementation according to the first aspect or any implementation thereof, generating the one or more task-adaptive auxiliary tokens comprises receiving an audio signal associated with the input visual signal and generating the one or more task-adaptive auxiliary tokens based on the audio signal.

[0008] Optionally, in a fourth implementation according to the first aspect or any implementation thereof, generating the one or more task-adaptive auxiliary tokens comprises receiving additional data related to the input visual signal that can assist or improve performance of the at least one specific downstream task and generating the one or more task-adaptive auxiliary tokens based on the additional data.

[0009] Optionally, in a fifth implementation according to the first aspect or any implementation thereof, generating the one or more task-adaptive auxiliary tokens comprises receiving an additional prompt input used for reconstructing the input visual signal and generating the one or more task-adaptive auxiliary tokens based on the additional prompt input.

[0010] Optionally, in a sixth implementation according to the first aspect or any implementation thereof, the base TVR framework is used for various tasks and does not require retraining when used with the one or more task-adaptive auxiliary tokens.

[0011] Optionally, in a seventh implementation according to the first aspect or any implementation thereof, encoding the input visual signal using the base TVR framework to generate the discrete indices string comprises patchifying the input visual signal into a sequence of blocks, generating for each block an embedded discrete latent feature tensor, mapping each embedded discrete latent feature tensor to a closest token in a pre-trained discrete codebook to produce discrete token indices, and entropy-encoding the discrete token indices to generate the discrete indices string.Atty. Docket No. 4502-86601 (6000767PCT02)

[0012] Optionally, in an eighth implementation according to the first aspect or any implementation thereof, selectively encoding the input visual signal in the quality-enhancement processing stream to generate the continuous latent string comprises patchifying the input visual signal into a sequence of blocks, generating for each block an embedded continuous latent feature tensor, computing a tokenized continuous latent representation from the embedded continuous latent feature tensor, and applying quantization and entropy encoding to the tokenized continuous latent representation to generate the continuous latent string.

[0013] A second aspect relates to a method for task-oriented image or video processing implemented by a receiving device. The method includes receiving from a sending device a discrete indices string of an input visual signal generated using a base TVR framework, receiving when generated by the sending device using a quality-enhancement processing stream a continuous latent string, decoding the discrete indices string to recover discrete latent features, and performing one or more downstream tasks and reconstruction of the input visual signal based on the discrete latent features and when available the continuous latent string.

[0014] Optionally, in a first implementation according to the second aspect, the method further includes receiving when generated by the sending device one or more task-adaptive auxiliary tokens that are modular and independent of the base TVR framework, wherein the one or more task-adaptive auxiliary tokens comprise one or more task-oriented prompt tokens that provide enhanced semantic representation for at least one specific downstream task, and performing the at least one specific downstream task using both the discrete latent features, the one or more task-oriented prompt tokens, and when available the continuous latent string.

[0015] Optionally, in a second implementation according to the second aspect or any implementation thereof, the one or more task-adaptive auxiliary tokens is based on a text description of the input visual signal.

[0016] Optionally, in a third implementation according to the second aspect or any implementation thereof, the one or more task-adaptive auxiliary tokens is based on an audio signal associated with the input visual signal.

[0017] Optionally, in a fourth implementation according to the second aspect or any implementation thereof, the one or more task-adaptive auxiliary tokens is based on additional data related to the input visual signal that assists or improves performance of the at least one specific downstream task.Atty. Docket No. 4502-86601 (6000767PCT02)

[0018] Optionally, in a fifth implementation according to the second aspect or any implementation thereof, the one or more task-adaptive auxiliary tokens comprise an additional prompt token that is further used in the reconstruction of the input visual signal.

[0019] Optionally, in a sixth implementation according to the second aspect or any implementation thereof, the base TVR framework is used for various tasks and does not require retraining when the one or more task-adaptive auxiliary tokens are used.

[0020] Optionally, in a seventh implementation according to the second aspect or any implementation thereof, decoding the discrete indices string comprises decoding the discrete indices string to recover discrete token indices and retrieving a recovered discrete latent feature tensor from a pre-trained discrete codebook for each recovered token index.

[0021] Optionally, in an eighth implementation according to the second aspect or any implementation thereof, when the continuous latent string is received, the method further comprises decoding the continuous latent string to recover a decoded tokenized continuous latent representation and providing the decoded tokenized continuous latent representation to at least one of a task-oriented tokenization module, a task execution module, or a reconstruction module.

[0022] A third aspect relates to an apparatus comprising a memory configured to store instructions; and one or more processors coupled to the memory and configured to execute the instructions to cause the apparatus to perform the method according to the first aspect or any implementation thereof.

[0023] A fourth aspect relates to an apparatus comprising a memory configured to store instructions; and one or more processors coupled to the memory and configured to execute the instructions to cause the apparatus to perform the method according to the second aspect or any implementation thereof.

[0024] A fifth aspect relates to a computer program product comprising computer-executable instructions stored on a non-transitory computer-readable storage medium and that, when executed by one or more processors of an apparatus, cause the apparatus to perform the method according to the first aspect or any implementation thereof or the second aspect or any implementation thereof.

[0025] For clarity, any one of the foregoing aspects may be combined with any one or more of the other foregoing aspects to create a new embodiment within the scope of the present disclosure.Atty. Docket No. 4502-86601 (6000767PCT02)

[0026] These and other features, and the advantages thereof, will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims.BRIEF DESCRIPTION OF DRAWINGS

[0027] For a more complete understanding of this disclosure, reference is now made to the following brief description, taken in connection with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.

[0028] FIG. 1 is a diagram illustrating an example a task-specific feature compression framework or pipeline for processing images and videos in accordance with one or more embodiments of the disclosure.

[0029] FIG. 2 is a diagram illustrating an example of a TVR framework or pipeline for processing images and videos in accordance with one or more embodiments of the disclosure.

[0030] FIG. 3A and 3B illustrate a dual-stream Learned Video compression (LVC) with TVR framework in accordance with one or more embodiments of the disclosure.

[0031] FIG. 4A and FIG. 4B illustrate a task-oriented TVR framework for image and video compression in accordance with one or more embodiments of the disclosure.

[0032] FIG. 5A and FIG. 5B illustrate the sender side and receiver side workflows of an embodiment of the task-oriented TVR framework with additional multi-modal inputs in accordance with one or more embodiments of the disclosure.

[0033] FIG. 6 is a flowchart illustrating a process implemented by a sending device for performing task-oriented image or video compression in accordance with one or more embodiments of the disclosure.

[0034] FIG. 7 is a flowchart illustrating a process implemented by a receiving device for task-oriented image or video processing in accordance with one or more embodiments of the disclosure.

[0035] FIG. 8 is a diagram illustrating an apparatus in accordance with one or more embodiments of the disclosure.DESCRIPTION OF EMBODIMENTS

[0036] It should be understood at the outset that, although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods may beAtty. Docket No. 4502-86601 (6000767PCT02)implemented using any number of techniques, whether currently known or in existence. The disclosure should in no way be limited to the illustrative implementations, drawings, and techniques illustrated below, including the exemplary designs and implementations illustrated and described herein, but may be modified within the scope of the appended claims along with their full scope of equivalents.

[0037] The present disclosure provides a framework, architecture, systems, and methods that use TVR for task-oriented image and video compression. TVR refers to the compact latent representation of an input image or video obtained by an input tokenizer. In general, a tokenizer encodes raw pixels into a lower-dimensional latent feature map, splits the map into small patches (spatial for images, spatiotemporal for video), maps each patch to either a discrete token (e.g., an integer index that points to one entry in a learned codebook) or a continuous token (e.g., a floatingpoint vector that encodes the patch in a continuous latent space), and outputs a sequence of tokens (i.e., the TVR). Thus, each token represents one patch or unit of the original image / video after it has been encoded and quantized. There are many tokenizers specially designed for either images (for example, Vector Quantized Generative Adversarial Network (VQGAN)) or videos (for example, like Masked Generative Video Transformer (MAGVIT)). Task-oriented image and video compression means compression is being optimized for artificial intelligence (AI) / machine-vision tasks while also supporting human visual perception. For example, instead of trying to make the reconstructed picture look sharp and natural for human perception like traditional codecs, task-oriented image and video compression is designed to maximize accuracy for specific Al tasks (e g., facial recognition, object detection, autonomous driving, robotics, etc.).

[0038] Prior compression approaches for machine vision suffer from several fundamental limitations including poor generality, poor scalability, and unsustainable design. First, these methods are tightly coupled with specific downstream models, resulting in poor flexibility / generality. For example, for each task model, a separate compression scheme is typically optimized, but this tight binding severely limits generalization. For instance, when a task model is updated or replaced with a different architecture, the corresponding compression model often requires retraining because the intermediate features change significantly. Second, these methods are also task-specific, meaning a compression strategy tailored for one task, such as object detection, cannot be readily applied to another, such as visual caption, due to substantial differences in intermediate representations. Third, by focusing solely on compressing intermediateAtty. Docket No. 4502-86601 (6000767PCT02)visual features, these methods struggle to leverage recent advances in multimodal vision -language models (VLMs), which are increasingly used to enhance downstream task performance. This narrow focus prevents integration with broader, more powerful representations that combine visual and linguistic signals, limiting the potential for cross-modal reasoning and transferability across tasks. With the rapid proliferation of artificial intelligence (Al) applications, the number of tasks and associated models has become vast. This design paradigm relying on task-model-specific compression of intermediate features has thus become increasingly inefficient and difficult to scale.

[0039] In addition, machine vision tasks typically place high demands on transmission, storage, and computational efficiency. Prior methods often fall short in meeting these requirements, making them impractical for deployment in real-world, resource-constrained scenarios. Compared to conventional coding paradigms, transferring token indices offers several key advantages for machine vision tasks such as robustness across heterogeneous platforms, the ability to operate at extremely low bitrates due to the flexibility of expanding latent feature dimensions without increasing the bitstream size, and the resilience of high-quality visual tokens even when inputs are of low quality. However, directly applying tokenized compression to machine vision faces the same scalability and generalization challenges as earlier feature-based methods such as having to learn a separate tokenizer for each specific task, which is highly impractical. Accordingly, the disclosed embodiments provide a flexible, scalable image and video compression framework that retains the benefits of tokenized representations while generalizing across diverse tasks and models. In particular, embodiments of the present disclosure provide a general framework for task-oriented image and video compression that utilizes TVR. The visual tokens carry rich semantic information that enables reconstruction of high-quality image and video data suitable for the human visual system as well as facilitation of machine vision tasks.

[0040] To ensure scalability and generalization across diverse tasks, models, and modalities while maintaining low bitrates, the framework incorporates three key mechanisms:

[0041] a) Plug-and-play auxiliary tokens: Built upon a base tokenized image or video compression framework, the system incorporates task-adaptive auxiliary tokens that are modular and are selectively appended to the latent token stream. These auxiliary tokens do not alter the base compression model. This configuration enables scalable deployment across multiple tasks orAtty. Docket No. 4502-86601 (6000767PCT02)task models. The auxiliary tokens are activated only when required for a specific task, thereby avoiding unnecessary bitrate overhead and enabling task-specific bitrate optimization.

[0042] b) Token reuse across tasks: The framework supports reuse of encoded tokens for multiple tasks during inference without re-encoding or retransmission. This approach significantly reduces transmission and computation costs in multi-task scenarios.

[0043] c) Unified multimodal token integration: The framework is configured to incorporate tokens from various modalities, such as visual, audio, and textual inputs. This enables generalization across tasks and seamless integration with advanced vision-language models for improved task performance.

[0044] Additionally, the framework supports tasks directed to both human visual perception and machine vision. In particular, the framework can be adapted to enhance perceptual quality based on human-centric criteria (e.g., visual fidelity or aesthetics) or to optimize performance on machine vision tasks based on task-specific objectives (e.g., accuracy in detection, segmentation, or classification). This dual capability provides the framework with versatility across a wide range of applications.

[0045] FIG. 1 is a diagram illustrating an example of a task-specific feature compression framework or pipeline for processing images and videos in accordance with one or more embodiments of the disclosure. The task-specific feature compression process includes an Encoder 102, Feature Coding module 104, a Decoder 106, and a Task Model 108. The Encoder 102 receives as input a visual signals. The visual signals may be images and videos. The Encoder 102 processes the input visual signal x through multiple neural network layers to produce a latent representation zx. The latent representation zxis a compact, lower-dimensional feature vector or tensor that encodes the essential information of the input visual signal x. In some embodiments, the Encoder 102 may extract task-relevant features from x. For example, the Encoder 102 may focus on semantic features that are most useful for the target machine vision task, such as object detection, facial or object recognition, image classification, or semantic segmentation. In some embodiments, the Encoder 102 may employ a convolutional neural network (CNN), a Vision Transformer (ViT), or other deep feature extraction architectures to map the high-dimensional visual input into the compact latent representation zxwhile preserving critical task-specific information.Atty. Docket No. 4502-86601 (6000767PCT02)

[0046] The Feature Coding module 104 receives the latent representation zx. The Feature Coding module 104 performs feature coding on zx. The Feature Coding module 104 produces a compressed representation zxto reduce the bit rate required for transmission or storage. For example, the Feature Coding module 104 may apply quantization or entropy coding to zxto generate the compressed representation zx. Quantization converts the continuous floating-point values in zxinto a set of discrete integer values. This quantization step significantly reduces the number of bits required to represent each element while introducing controlled information loss. In some embodiments, the Feature Coding module 104 then applies entropy coding to the quantized values. Entropy coding removes statistical redundancy present in the quantized data by assigning shorter codewords to more probable symbols. The Feature Coding module 104 may use arithmetic coding, Huffman coding, or other entropy coding techniques.

[0047] The Decoder 106 receives the compressed representation zxfrom the Feature Coding module 104. The Decoder 106 reconstructs an approximate image or video x of the original visual signal x. In some embodiments, the Decoder 106 first applies dequantization to convert the discrete values in zxback into a continuous latent representation. The Decoder 106 then reshapes or projects this latent representation into a multi-dimensional feature map that matches the spatial and temporal structure of the original visual signal. The Decoder 106 progressively upsamples and refines this feature map through multiple neural network layers until the original dimensions of the visual signal x are restored. For example, the Decoder 106 may utilize transposed convolutional layers, residual blocks, attention mechanisms, normalization layers, pixel-shuffle layers, or non-linear activation functions. These operations enable the Decoder 106 to generate a high-quality approximation x of the original visual signal while preserving sufficient task-critical information for the downstream machine vision task. In some embodiments, the Decoder 106 is trained jointly with the Encoder 102 and the Feature Coding module 104 to achieve optimal task performance.

[0048] The Task Model 108 receives the reconstructed image or video x from the Decoder 106. The Task Model 108 performs the target machine vision task on x. For example, the Task Model 108 may execute object detection, object recognition, image classification, semantic segmentation, or instance segmentation. In some embodiments, the Task Model 108 is a pretrained or fine-tuned deep neural network that is optimized specifically for the intended task. TheAtty. Docket No. 4502-86601 (6000767PCT02)performance of the Task Model 108 on the reconstructed signal x may serve as a metric for evaluating the overall effectiveness of the task-specific feature compression framework.

[0049] FIG. 2 is a diagram illustrating an example of a TVR framework or pipeline for processing images and videos in accordance with one or more embodiments of the disclosure. The TVR process includes a Patchify module 202, an Input Tokenizer 204, and a Decoder 206. The Patchify module 202 receives an input image or video x.

[0050] The Patchify module 202 divides the input visual signal x into a sequence of overlapping or non-overlapping two-dimensional (2D) or three-dimensional (3D) patches xp. In some embodiments, for a 2D image input characterized by height H, width W, and C color channels, the module may partition the image into fixed-size patches of resolution P × P pixels, where P is a predetermined patch size (for example, 8, 16, or 32). These patches may be extracted either non-overlappingly (for example, using a stride equal to the patch size P, producing exactly (H / P) x (W / P) patches arranged in raster-scan order) or overlappingly (for example, using a stride smaller than P to provide denser sampling and richer contextual overlap between adjacent patches). In some embodiments, each extracted patch may be flattened into a one-dimensional vector of length P × P × C. For video inputs that include an additional temporal dimension T (i.e., T frames), the Patchify module 202 may, in some embodiments, perform 3D patchification by extracting spatio-temporal patches of size Tp× P × P, where Tpis the temporal patch size (such as 1, 2, or 4). The resulting ordered sequence of patches may preserve localized visual information and may be directly forwarded as input to the Input Tokenizer 204 for conversion into latent representations.

[0051] The Input Tokenizer 204 receives as input the sequence of patches xp. In some embodiments, the Input Tokenizer 204 is a continuous tokenizer and is configured to model the probability distributions of the visual space and transforms each patch into a latent feature. A latent feature is a compact vector of continuous floating-point numbers (real values) that encodes the essential characteristics of one image / video patch. Alternatively, in some embodiments, the Input Tokenizer 204 is a discrete tokenizer and instead of outputting the latent feature itself, the Input Tokenizer 204 can output a latent code, which is a discrete integer index value that points to one specific entry in a learned lookup table called a codebook that contains the closest matching vector to the latent feature. This process produces either a discretized sequence of latent codes (for example, by using VQGAN) or a continuous latent space (for example, by using a Variational Autoencoder (VAE)). The resulting sequence of tokens (latent codes or latent features) is thenAtty. Docket No. 4502-86601 (6000767PCT02)concatenated along the sequence dimension to form an embedding zx(when transformed to continuous latent features) or zxq(when transformed to discrete latent codes).

[0052] The Decoder 206 receives as input the embedding zxor zxq. In some embodiments, when the embedding consists of discrete latent codes zxq, the Decoder 206 may first map each latent code to its corresponding continuous vector from a learned codebook to recover a continuous latent representation. The Decoder 206 may then reshape or project the sequence of latent representations back into a multi-dimensional feature map that corresponds to the spatial (and temporal for video) structure of the original patches. The Decoder 206 then progressively transforms this feature map back to the pixel space through a series of upsampling and refinement layers. These layers may consist of transposed convolutional layers, residual blocks, attention mechanisms, normalization layers, and non-linear activations that gradually increase the spatial and temporal resolution until the original dimensions of the input visual signal x are restored, thereby reconstructing the output image or video x. In some embodiments, a final output convolutional layer generates the reconstructed pixels or frames. The Decoder 206 may be implemented as a deep neural network, such as a convolutional decoder or a transformer-based decoder. In some embodiments, the Decoder 206 is jointly optimized end-to-end with the Input Tokenizer 204 to balance the efficiency of latent representation and the reconstruction quality.

[0053] FIG. 3A and 3B illustrate a dual-stream Learned Video compression (LVC) with TVR framework in accordance with one or more embodiments of the disclosure. LVC is a class of video compression methods that replace traditional video codecs (such as High Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC)) with deep neural networks (NNs) trained end-to-end that learns the entire compression pipeline directly from data. Research has shown that LVC consistently achieves better rate-distortion performance (higher visual quality at the same bitrate, or lower bitrate for the same quality) than traditional standards like HEVC and VVC.

[0054] FIG. 3A illustrates the sender side and includes a Continuous Tokenizer 302, a Latent Encoding module 304, a Discrete Tokenizer 306, a Masking module 308, and an Indices Encoding module 310. In a first stream on the sender side, the Continuous Tokenizer 302 receives the input frame xt. The input frame xtmay be one frame from a Group of Pictures (GoP) consisting of n frames in a video segment. The Continuous Tokenizer 302 computes a continuous token zxc. The continuous token zxcis a compact continuous latent representation of the input frame xtin a continuous latent space. In some embodiments, the Continuous Tokenizer 302 may transform theAtty. Docket No. 4502-86601 (6000767PCT02)input frame xtinto a continuous latent representation by modelling the probability distributions of the visual space (for example, using a neural network such as a VAE). The Latent Encoding module 304 encodes the continuous token zxcand produces a continuous token data string sxc. The continuous token data string sxcis the compressed data string obtained from the continuous token through quantization and arithmetic coding for efficient storage and transmission. In some embodiments, the Latent Encoding module 304 may apply quantization and arithmetic coding.

[0055] In a second stream on the sender side, the Discrete Tokenizer 306 receives the input frame xt. The Discrete Tokenizer 306 computes a discrete token z The discrete token z^is a sequence of discrete integer indices selected from a learned codebook that represents the input frame xt. The Discrete Tokenizer 306 may transform the input frame xtinto these discrete latent codes, for example using a neural network such as a VQGAN. The Masking module 308 receives the discrete token zxd. The Masking module 308 applies masking to the discrete token zxd. The Masking module 308 produces a masked discrete token zxd,mask. The masked discrete token zxd,maskhas part of the discrete tokens masked out to reduce the number of bits transmitted. The Indices Encoding module 310 receives the masked discrete token zxd,maskThe Indices Encoding module 310 computes a discrete token data string sxd. The discrete token data string sxdis the compressed data string obtained from the masked discrete token through lossless integer arithmetic coding for efficient transmission. In some embodiments, the Indices Encoding module 310 may apply lossless integer arithmetic coding.

[0056] FIG. 3B illustrates the receiver side and includes a Latent Decoding module 312, an Indices Decoding module 314, a Discrete Token Prediction module 316, and a Pixel Decoding module 318. In the first stream on the receiver side, the Latent Decoding module 312 receives the continuous token data string sxcfrom the sender side. The Latent Decoding module 312 computes a decoded continuous token zxc. The decoded continuous token zxcis the recovered continuous latent representation of the original input frame xtafter decoding the received continuous token data string. In some embodiments, the Latent Decoding module 312 may apply arithmetic decoding and dequantization.

[0057] In the second stream on the receiver side, the Indices Decoding module 314 receives the discrete token data string sxdfrom the sender side. The Indices Decoding module 314 computes the masked discrete token zxd,mask. The masked discrete token zxd,maskis the version ofAtty. Docket No. 4502-86601 (6000767PCT02)the discrete token in which selected tokens have been masked out (not transmitted) to reduce bitrate. The Indices Decoding module 314 performs lossless integer arithmetic decoding on the received discrete token data string s^to recover this masked discrete token z^t’mask. The Discrete Token Prediction module 316 receives the masked discrete token z^t’maskand the decoded continuous token z^t. The Discrete Token Prediction module 316 reconstructs the masked-out tokens using information from the remaining tokens and the decoded continuous token zxcto recover the decoded discrete token zxd.

[0058] The Pixel Decoding module 318 receives the decoded continuous token zxcand the decoded discrete token zxd. The Pixel Decoding module 318 computes the reconstructed frame x̂tby jointly processing both the decoded continuous token zxcand the fully recovered decoded discrete token zxdIn some embodiments, the Pixel Decoding module 318 may first concatenate or fuse the continuous latent features from zxcwith the discrete token embeddings from zxd. The Pixel Decoding module 318 then passes the combined representation through a series of neural network layers. These layers may include transposed convolutional layers, residual blocks, attention mechanisms, normalization layers, or pixel-shuffle operations. Through this joint processing, the Pixel Decoding module 318 progressively upsamples and refines the fused features until the original spatial resolution is restored. The Pixel Decoding module 318 thereby generates the final reconstructed frame denoted as xt. The reconstructed frame xtis the high-quality approximation of the original input frame xt.

[0059] FIG. 4A and FIG. 4B illustrate a task-oriented TVR framework for image and video compression in accordance with one or more embodiments of the disclosure. FIG. 4A illustrates the sender side of the task-oriented TVR framework, which includes a base TVR compression processing stream (i.e., a base TVR framework) and an optional quality-enhancement processing stream. In the base TVR compression processing stream, a Discrete Embedding (EMBd) module 402 receives an input visual signal xtof length T > 1. In an embodiment, the input visual signal xtis a general 4D tensor with shape T * h x M’ XC, where T is the length, h is the height, w is the width, and c is the number of channels. For example, c=3 for color videos, c=1 for spectral videos, or c=4 for RGB-D (color and depth) videos. The Discrete Embedding module 402 encodes the input xtinto an embedded discrete latent feature edIn an embodiment, the Discrete Embedding module 402 patchifies the input visual signal xtinto xtpconsisting of a sequence of blocks. ThenAtty. Docket No. 4502-86601 (6000767PCT02)for each block in the patchified xtp, the Discrete Embedding module 402 computes a discrete latent feature tensor with cddimensions. All the discrete latent feature tensors of all the blocks in the patchified xtpform the embedded discrete latent feature exdIn some embodiments, the Discrete Embedding module 402 may use various neural networks such as, but not limited to, a ViT. In some embodiments, when T = 1 (e.g., when the input visual signal xtis an image), xtis patchified along the width and height dimensions into nwand nhblocks, with each block having size i / nwx h / nhx c. In some embodiments, when T > 1 (e.g., the input visual signal xtis GoP, xtcan be patchified frame-by-frame first and then blocks from individual frames are stacked into xtp. Alternatively, in some embodiments, when T >1, xtcan be patchified as a whole unit by partitioning along the width, height, and temporal dimensions into nw, nh, and nTblocks, with each block having size T / nTx w / nwx h / nhx c. That is, edtconsists nwx nhx nTlatent features of cddimension, nT≥ 1 (nT= 1 when T=1).

[0060] A Discrete Tokenization module 404 receives the embedded discrete latent feature edt. The Discrete Tokenization module 404 computes a discrete token indices zdtbased on edt. The discrete token indices zdtis a sequence of integers corresponding to the indices of the mapped tokens in a pre-trained discrete codebook. In an embodiment, a pre-trained discrete tokenizer (such as Nvidia’s Cosmos discrete video tokenizer) is used, which comprises of a list of video tokens Tknd, each token being a latent feature vector of cddimension. Each latent feature vector in edtcorresponding to each block in x is mapped to a token in Tkndthat is closest to the latent feature vector of the block measured by a distance metric (such as LI or L2 norm). That is, the entire zdthas nwx nhx nTintegers corresponding to the indices of the mapped tokens. The Indices Encoding module 406 receives the discrete token indices zdt. The Indices Encoding module 406 computes a discrete indices string sd. The discrete indices string sdtis a compressed representation of the discrete token indices zdtand is transmitted to the receiver side. In some embodiments, the Indices Encoding module 406 may apply lossless integer entropy coding methods to further reduce the bit consumption of representing the token indices zd

[0061] While the base TVR framework is extremely bitrate-efficient, visual fidelity can be limited due to codebook quantization. Thus, the sender side of the task-oriented TVR framework may selectively include a quality-enhancement processing stream (indicated as optional by theAtty. Docket No. 4502-86601 (6000767PCT02)dash lines in the drawings) that generates and transmits a continuous latent representation that provides finer-grained, higher-fidelity features. The receiver uses the continuous latent representation together with the discrete tokens generated by the base TVR framework to produce a reconstructed image / video having better visual fidelity and fewer artifacts. This design provides a controllable trade-off between bitrate and perceptual quality. For example, for tasks that can be performed using lower fidelity images, only the base TVR framework is needed. For other tasks (e.g., when a task performance using just the discrete indices generated by the base TVR framework is below a certain threshold) or to improve human-visual-perception, the qualityenhancement processing stream may be selectively executed to improve the reconstructed image / video without touching or retraining the base TVR framework.

[0062] In the depicted embodiment, in the quality-enhancement processing stream on the sender side, a Continuous Embedding module 408 EMB1) also receives the input visual signal xtand encodes the input xtinto an embedded continuous latent feature eXlt. In some embodiments, the Continuous Embedding module 408 patchifies the input visual signal xtinto patchified (either in the same way as in the Discrete Embedding module 402 or differently from the Discrete Embedding module 402). The Continuous Embedding module 408 computes a continuous latent feature tensor with cldimensions for each block in the patchified x. All the continuous feature tensors of all the blocks in the patchified x form the embedded continuous reference latent feature eXt. Various neural networks can be used as the Continuous Embedding module 408 such as a ViT. The quality-enhancement processing stream also includes a Continuous Tokenization module 410 configured to receive the embedded continuous reference latent feature exland compute a tokenized continuous latent zxl. The tokenized continuous latent zxlis a compact continuous representation of the input visual signal xt. In some embodiments, the Continuous Tokenization module 410 may use a pre-trained continuous video tokenizer such as Nvidia’s Cosmos continuous video tokenizer based on an Autoencoder network. The qualityenhancement processing stream further includes a Latent Encoding module 412 configured to receive the tokenized continuous latent zxl. The Latent Encoding module 412 computes a continuous latent string sxl. The continuous latent string sxlis a compressed representation of the tokenized continuous latent zxland is transmitted to the receiver side. In some embodiments, theAtty. Docket No. 4502-86601 (6000767PCT02)Latent Encoding module 412 may apply quantization and entropy encoding so that the generated continuous latent string sxlconsumes less bits than the tokenized reference latent zxl.

[0063] FIG. 4B illustrates the receiver side of the task-oriented TVR framework, which includes a base TVR processing stream and a quality-enhancement processing stream corresponding to the sender side. In the base TVR processing stream, an Indices Decoding module 414 receives the discrete indices string sxdfrom a sender side. The Indices Decoding module 414 computes the discrete token indices zxd. In some embodiments, the Indices Decoding module 414 may apply lossless integer entropy decoding corresponding to the lossless integer entropy encoding methods in the Indices Encoding module 406, so that zxdis fully recovered. A Discrete Latent Retrieval module 416 receives the discrete token indices zxd. The Discrete Latent Retrieval module 416 computes a recovered discrete latent feature ẽxdusing the same pre-trained discrete tokenizer as the sender side to retrieve the mapped video token for each token index in zxdfrom the list of video tokens Tknd. That is, ẽxdhas the same shape as the original exdon the sender side, where each latent feature in exdis replaced by the mapped feature vector of the token in the pre-trained tokenizer.

[0064] In the quality-enhancement processing stream on the receiver side, a Latent Decoding module 424 receives the continuous latent string sxl. The Latent Decoding module 424 computes a decoded tokenized continuous latent representation z̃xl. In some embodiments, the Latent Decoding module 424 may apply entropy decoding and dequantization processes corresponding to the quantization and entropy encoding processes used in the Latent Encoding module 412 on the sender side.

[0065] To support a set of M tasks, where M is an integer equal to or greater than 1, for the i-th task, the recovered discrete latent feature ẽxdand optionally the decoded tokenized continuous latent representation z̃xlare fed into a Task-Oriented Tokenization i module 418, which computes an auxiliary task-oriented latent feature zxaux,i. The auxiliary task-oriented latent feature zxaux,iis a task-specific, supplementary latent representation generated for the i-th task in a multi-task system that uses shared latent spaces. The recovered discrete latent feature eXt, the auxiliary task-oriented latent feature zx^x’1, and optionally the decoded tokenized continuous latent representation zxare fed into a Task z Execution module 420 to perform the z-th task. The Task-Atty. Docket No. 4502-86601 (6000767PCT02)Oriented Tokenization i module 418 may employ a pre-trained task-oriented tokenizer, either discrete or continuous, which can have similar network structures as the Discrete Tokenization module 404 or the Continuous Tokenization module 410. A Reconstruction module 422 receives the recovered discrete latent feature ẽxdand optionally the decoded tokenized continuous latent representation z̃xl. The Reconstruction module 422 computes the reconstructed output xt. In some embodiments, the Reconstruction module 422 may be a convolutional UNet or a denoising diffusion model.

[0066] FIG. 5A and FIG. 5B illustrate the sender side and receiver side workflows of an embodiment of the task-oriented TVR framework with additional multi-modal inputs in accordance with one or more embodiments of the disclosure. In FIG. 5A, on the sender side, in addition to the base TVR framework and the quality-enhancement processing stream described in FIG. 4A, the system can generate one or more task-adaptive auxiliary tokens that are modular and independent of the base TVR framework based on one or more additional multi-modal inputs including an additional prompt Yx. Thus, the task-adaptive auxiliary tokens are task-oriented plug-and-play tokens that do not alter the base compression model, enabling scalable deployment across multiple tasks or task models. For example, Yxmay be a text description of the content of xt, an audio signal associated with xt, or any additional data related to xtthat can assist in or improve reconstruction the visual signal. In some embodiments, the additional prompt Yxcan be provided as direct inputs, or by a model from xt, such as a vision-language model (VLM) to generate a text description Yxfrom xt, or an audio generation model to obtain audio inputs from xt, etc. An Additional Prompt Tokenizer 502 processes the additional prompt Yxto compute additional prompt tokens zxyIn some embodiments, the Additional Prompt Tokenizer 502 can be a discrete tokenizer or a continuous tokenizer, e.g., similar to the Discrete Tokenization module 404 or the Continuous Tokenization module 410 of generating zxdand zxlusing the visual information in FIG. 4A. Then the additional prompt tokens zxyare encoded by an Additional Prompt Token Encoding module 504 to compute an additional prompt string sxywhich is sent to the receiver side. In some embodiments, the Additional Prompt Token Encoding module 504 employs a lossless integer entropy encoding process when the prompt tokens zxyconsist of discrete token indices, or employs a quantization and entropy encoding process when zxyconsists of continuous latent features.Atty. Docket No. 4502-86601 (6000767PCT02)

[0067] In some embodiments, for an i-th task, a task-oriented additional prompt input Yxaux,iis provided to the system. The task-oriented additional prompt input Yxaux,iis an extra input signal that is specifically tailored to enhance performance of one particular task i (e.g., object detection, segmentation, classification, captioning, etc.). Non-limiting examples of the task-oriented additional prompt input Y^x’1may include, but are not limited to, a task-specific text instruction or query (e.g., “detect all vehicles and report their bounding boxes”), or any other conditioning prompt generated or provided for task i. The task-oriented additional prompt input Y^lx’1is processed by a Task-Oriented Prompt Tokenizer z module 512 to compute a task-oriented auxiliary prompt tokens z““,t. Similarly, the Task-Oriented Prompt Tokenizer i module 512 can be a discrete tokenizer or a continuous tokenizer. Then the task-oriented auxiliary prompt tokens zx‘lare processed by a Task-Oriented Prompt Token Encoding module 514 to compute a task-oriented auxiliary prompt string s““x,t, which is sent to the receiver side. In some embodiments, the Task-Oriented Prompt Token Encoding module 514 may employ a lossless integer entropy encoding process when the task-oriented auxiliary prompt tokens z““x,tconsist of discrete token indices, or employs a quantization and entropy encoding process when z““x,tconsists of continuous latent features.

[0068] FIG. 5B illustrates the receiver side workflow of an embodiment of the proposed method when multi-modal inputs other than visual information are used. In addition to the base TVR compression processing stream and the task-adaptive auxiliary tokens processing stream of FIG. 4B, on the receiver side, an Additional Prompt Token Decoding module 516 computes a decoded additional prompt latent feature z̃xyusing the received additional prompt string sxy. Similar to the visual counterparts, when the prompt tokens zxyconsist of discrete token indices, the Additional Prompt Token Decoding module 516 may employ a lossless integer entropy decoding method to recover the prompt token indices zxyand then retrieves the tokenized latent feature from the discrete tokenizer used by the Additional Prompt Tokenizer module 502 to obtain z̃xy. Alternatively, when the prompt tokens zxyconsist of continuous latent features, the Additional Prompt Token Decoding module 516 employs an entropy decoding and dequantization process to compute z̃xy. Then the decoded additional prompt latent feature z̃xyis fed into the ReconstructionAtty. Docket No. 4502-86601 (6000767PCT02)module 422 along with the recovered discrete latent feature exand optionally the decoded tokenized continuous latent representation zXltto compute the reconstructed output xt.

[0069] In addition, on the receiver side, for the / -th task, a Task-Oriented Prompt Token Decoding module 506 computes a decoded task-oriented prompt latent feature z“tux,tusing the task-oriented auxiliary prompt string sx^lx'1received from the receiving side. Similar to the visual counterparts, when the task-oriented auxiliary prompt tokens zx™x’1consists of discrete token indices, the Task-Oriented Prompt Token Decoding 506 may employ a lossless integer entropy decoding method to recover the task-oriented auxiliary prompt tokens zx™x, Land then retrieves the tokenized latent feature from the discrete tokenizer used by the Task-Oriented Prompt Tokenizer 506 to obtain the decoded task-oriented prompt latent feature zxx'1. Alternatively, when the task-oriented auxiliary prompt tokens z““,tconsists of continuous latent features, the Task-Oriented Prompt Token Decoding module 506 may employ an entropy decoding and dequantization process to compute zxx'1. Then the decoded task-oriented prompt latent feature z““,tis fed into the Task i Execution module 420 along with the recovered discrete latent feature eXtand optionally the decoded tokenized continuous latent representation zxlto perform the z-th task. Thus, the framework in FIG. 5A and FIG. 5B enables fine-grained, task-specific multimodal conditioning inside the unified tokenized framework, while keeping everything modular and plug-and-play.

[0070] FIG. 6 is a flowchart illustrating a process 600 implemented by a sending device for performing task-oriented image or video compression in accordance with one or more embodiments of the disclosure. A non-limiting example of a sending device for implementing the process 600 is the sending device shown in FIG. 5A. The process 600 begins, at step 602, by encoding an input visual signal using a base TVR framework to generate a discrete indices string. At step 604, the sending device selectively encodes the input visual signal in a qualityenhancement processing stream to generate a continuous latent string to improve at least one of task performance or reconstruction of the input visual signal. That is, the sending device determines whether or not to encode the input visual signal in a quality-enhancement processing stream to generate a continuous latent string to improve at least one of task performance or reconstruction of the input visual signal. The determination of whether or not to encode the input signal is based on [the downstream task requirements. For example, for frame-level classificationAtty. Docket No. 4502-86601 (6000767PCT02)task, a relatively low-quality reconstruction is enough to categorize the frames. Fortasks requiring fine details such as segmentation or depth estimation, relatively high-quality reconstruction may be necessary to ensure task performance, and quality-enhancement processing stream may be needed. At step 606, the sending device selectively generates one or more task-adaptive auxiliary tokens that are modular and independent of the base TVR framework, wherein the auxiliary tokens provide enhanced semantic representation for at least one specific downstream task. That is, the sending device determines whether or not to generates one or more task-adaptive auxiliary tokens that are modular and independent of the base TVR framework. The determination of whether or not to generate the one or more task-adaptive auxiliary tokens is based on the downstream task requirements. For example, for tasks related to semantic understanding, such as object detection and recognition, the sender can obtain a more accurate semantic text description of the semantic content of the frames, and such text description can help the receiver to perform detection and recognition with reconstructed compressed frames. At step 608, the sending device transmits the discrete indices string and, when generated, the continuous latent string and the one or more task-adaptive auxiliary tokens to a receiving device.

[0071] FIG. 7 is a flowchart illustrating a process 700 implemented by a receiving device for task-oriented image or video processing in accordance with one or more embodiments of the disclosure. A non-limiting example of a receiving device for implementing the process 700 is the receiver shown in FIG. 5B. The process 700 begins, at step 702, by receiving a discrete indices string of an input visual signal generated by a sending device using a base TVR framework. At step 704, the receiving device receives, when generated by the sending device, a continuous latent string. At step 706, the receiving device receives, when generated by the sending device, one or more task-adaptive auxiliary tokens that are modular and independent of the base TVR framework. At step 708, the receiving device decodes the discrete indices string to recover discrete latent features, decodes the continuous latent string (when received) to recover a decoded tokenized continuous latent representation, and decodes the one or more task-adaptive auxiliary tokens (when received) to recover the corresponding auxiliary latent features. Finally, at step 710, the receiving device performs one or more downstream tasks and / or reconstruction of the input visual signal using the recovered discrete latent features, the decoded task-adaptive auxiliary tokens when received, and the decoded tokenized continuous latent representation when received.Atty. Docket No. 4502-86601 (6000767PCT02)

[0072] The present disclosure does not place any restrictions on how the disclosed embodiments described herein are trained. As a first non-limiting example of a training process, the sender side of FIG. 4A and the receiver side of FIG. 4B may be trained in several stages as follows:

[0073] Stage 1: The Discrete Embedding (EMBd) module 402, the Discrete Tokenization module 404, the Discrete Latent Retrieval module 416, and the Reconstruction module 422 are trained end-to-end by mainly optimizing the distortion loss between the reconstructed xtand the original input xtand the codebook loss, without the optional continuous stream processes.

[0074] Stage 2: The Discrete Embedding (EMBd) module 402, the Discrete Tokenization module 404, and the Discrete Latent Retrieval module 416 are fixed. The optional Continuous Embedding module 408 (EMB1), the Continuous Tokenization module 410, the Latent Encoding module 412, and the Latent Decoding module 424 are trained where the Reconstruction module 422 is also finetuned, by mainly optimizing the Rate-Distortion loss between the reconstructed xtand the original input xt.

[0075] Stage 3: All modules in stage 1 and 2 are fixed. The Task-Oriented Tokenization i module 418 and the Task / Execution module 420 are trained for each task z, by optimizing the corresponding loss for the task. Note that such models for each task can be trained individually or models for several similar tasks can be trained together.

[0076] As a second non-limiting example of a training process, the sender side of FIG. 5A and the receiver side of FIG. 5B may be trained in several stages as follows:

[0077] Stage 1: The Discrete Embedding (EMBd) module 402, the Discrete Tokenization module 404, the Discrete Latent Retrieval module 416, and the Reconstruction module 422 are trained end-to-end by mainly optimizing the distortion loss between the reconstructed xtand the original input xtand the codebook loss, without the optional continuous stream processes.

[0078] Stage 2: The Discrete Embedding (EMBd) module 402, the Discrete Tokenization module 404, and the Discrete Latent Retrieval module 416 are fixed. The optional Continuous Embedding module 408 (EMB1), the Continuous Tokenization module 410, the Latent Encoding module 412, and the Latent Decoding module 424 are trained where the Reconstruction module 422 is also finetuned, by mainly optimizing the Rate-Distortion loss between the reconstructed xtand the original input xt.Atty. Docket No. 4502-86601 (6000767PCT02)

[0079] Stage 3: Modules in Stage 1 and 2 are fixed, except for the Reconstruction module 422, which is finetuned at this stage. The Additional Prompt Tokenizer module 502, the Additional Prompt Token Encoding module 504, and the Additional Prompt Token Decoding module 516 are trained, by minimizing the Rate-Distortion loss between the reconstructed xtand the original input xt

[0080] Stage 4: All modules in stage 1, 2 and 3 are fixed. The Task-Oriented Tokenization z module 418, the Task-Oriented Prompt Tokenizer z module 512, the Task-Oriented Prompt Token Encoding module 514, and the Task i Execution module 420 are trained for each task z, by optimizing the corresponding loss for the task, as well as the bitrate loss. In some embodiments, models for each task may be trained individually. Alternatively, models for several similar tasks may be trained together.

[0081] FIG. 8 is a diagram illustrating an apparatus 800 according to an embodiment of the present disclosure. The apparatus 800 can be used to implement embodiments or various components of the present disclosure. The apparatus 800 includes receiver units (RX) 820 or receiving means for receiving data via ingress ports 810. The apparatus 800 also includes transmitter units (TX) 840 or transmitting means for transmitting via data egress ports 850.

[0082] The apparatus 800 includes a memory 860 or data storing means for storing the instructions and various data. The memory 860 can be any type of, or combination of, memory components capable of storing data and / or instructions. For example, the memory 860 can include volatile and / or non-volatile memory such as read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM). The memory 860 can also include one or more disks, tape drives, and solid-state drives. In some embodiments, the memory 860 can be used as an over-flow data storage device to store programs when such programs are selected for execution, and to store instructions and data that are read during program execution. In some embodiments, the memory 860 can be memory that is integrated with the processor 830.

[0083] The apparatus 800 has one or more processors 830 or other processing means (e.g., central processing unit (CPU)) to process instructions. The one or more processors 830 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), field-programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). The one or more processors 830 are communicatively coupled via aAtty. Docket No. 4502-86601 (6000767PCT02)system bus with the ingress ports 810, RX 820, TX 840, egress ports 850, Input / Output (I / O) (I / O means) 880, and memory 860. I / O 880 provides the communication interfaces for enabling the apparatus 800 to receive input (e.g., from a keyboard, mouse, or touchscreen) and output information (e.g., to a display or printer).

[0084] The one or more processors 830 can be configured to execute instructions stored in the memory 860. As an example, in one embodiment, the memory 860 stores a TVR with task-oriented plug-and-play tokens module 870. The TVR with task-oriented plug-and-play tokens module 870 includes data, executable instructions, and / or one or more sub-modules for implementing the disclosed embodiments. Thus, the one or more processors 830 provide a means for performing any computational, comparison, determination, initiation, configuration, or any other action corresponding to the claims when the appropriate instruction is executed by the processor 830. Thus, the inclusion of the TVR with task-oriented plug-and-play tokens module 870 substantially improves the functionality of the apparatus 800.

[0085] As described above, the framework described in the embodiments of the present disclosure provides the following novel features:

[0086] 1. A highly efficient, flexible and scalable tokenized compression framework. The disclosed embodiments employ a plug-and-play design with task-specific auxiliary tokens appended modularly to the main latent stream. These tokens are activated only when a specific task requires them, avoiding unnecessary bitrate overhead. This allows task-specific bitrate allocation, performance optimization, and flexible scaling to a wide range of applications.

[0087] 2. A unified system to support both human vision system and machine vision tasks. For human vision, the framework functions as a reconstruction-oriented codec, where auxiliary tokens can be progressively activated based on perceptual quality criteria. For machine vision, the framework serves as a machine-oriented coding pipeline, enhancing performance in tasks like detection or segmentation without compromising efficiency.

[0088] 3 A unified framework to support multimodal token integration. The framework supports seamless incorporation of tokens from multiple modalities, enabling generalization across diverse tasks. The framework is also compatible with advanced multimodal models (e.g., visionlanguage models), facilitating improved task performance through integrated multimodal learning, the framework.Atty. Docket No. 4502-86601 (6000767PCT02)

[0089] The present disclosure may be a system, a method, and / or a computer program product at any possible technical detail level of integration. The computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.

[0090] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0091] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0092] Computer readable program instructions for carrying out operations of the present disclosure may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data,Atty. Docket No. 4502-86601 (6000767PCT02)configuration data for integrated circuitry, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++, or the like, and procedural programming languages, such as the “C” programming language or similar programming languages. The computer readable program instructions may execute entirely on the user’ s computer, partly on the user’s computer, as a standalone software package, partly on the user’s computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user’s computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0093] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.

[0094] These computer readable program instructions may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.Atty. Docket No. 4502-86601 (6000767PCT02)

[0095] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0096] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the Figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.

[0097] While several embodiments have been provided in the present disclosure, it may be understood that the disclosed systems and methods might be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples are to be considered as illustrative and not restrictive, and the disclosure is not to be limited to the details given herein. For example, the various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.

[0098] In addition, techniques, systems, subsystems, and methods described and illustrated in the various embodiments as discrete or separate may be combined or integrated with other systems, modules, techniques, or methods without departing from the scope of the present disclosure. Other items shown or discussed as coupled or directly coupled or communicating with each other may be indirectly coupled or communicating through some interface, device, or intermediate component whether electrically, mechanically, or otherwise. Other examples of changes,Atty. Docket No. 4502-86601 (6000767PCT02)substitutions, and alterations are ascertainable by one skilled in the art and may be made without departing from the spirit and scope disclosed herein.

Claims

1. Atty. Docket No. 4502-86601 (6000767PCT02)CLAIMSWhat is claimed is:

1. A method implemented by a sending device for performing task-oriented image or video compression, the method comprising:encoding an input visual signal using a base tokenized visual representation (TVR) framework to generate a discrete indices string;selectively encoding, based on a quality requirement of a reconstructed input visual signal, the input visual signal in a quality-enhancement processing stream to generate a continuous latent string to improve at least one of a task performance or a reconstruction of the input visual signal, wherein the discrete indices string and the continuous latent string are applicable to all downstream tasks; andtransmitting the discrete indices string and, when generated, the continuous latent string to a receiving device.

2. The method of claim 1, further comprising:generating, based on task-specific requirements, one or more task-adaptive auxiliary tokens that are modular and independent of the base TVR framework, wherein the one or more task-adaptive auxiliary tokens provide enhanced semantic representation for at least one specific downstream task; andtransmitting the one or more task-adaptive auxiliary tokens along with the discrete indices string and, when generated, the continuous latent string to the receiving device.

3. The method of claim 2, wherein generating the one or more task-adaptive auxiliary tokens comprises:receiving a text description of the input visual signal; andgenerating the one or more task-adaptive auxiliary tokens based on the text description.

4. The method according to any of claims 2-3, wherein generating the one or more task-adaptive auxiliary tokens comprises:receiving an audio signal associated with the input visual signal; andgenerating the one or more task-adaptive auxiliary tokens based on the audio signal.Atty. Docket No. 4502-86601 (6000767PCT02)5. The method according to any of claims 2-4, wherein generating the one or more task-adaptive auxiliary tokens comprises:receiving additional data related to the input visual signal that can assist or improve performance of the at least one specific downstream task; andgenerating the one or more task-adaptive auxiliary tokens based on the additional data.

6. The method according to any of claims 2-5, wherein generating the one or more task-adaptive auxiliary tokens comprises:receiving an additional prompt input used for reconstructing the input visual signal; and generating the one or more task-adaptive auxiliary tokens based on the additional prompt input.

7. The method according to any of claims 1-6, wherein the base TVR framework is used for various tasks and does not require retraining when used with the one or more task-adaptive auxiliary tokens.

8. The method according to any of claims 1-7, wherein encoding the input visual signal using the base TVR framework to generate the discrete indices string comprises:patchifying the input visual signal into a sequence of blocks;generating, for each block, an embedded discrete latent feature tensor;mapping each embedded discrete latent feature tensor to a closest token in a pre-trained discrete codebook to produce discrete token indices; andentropy-encoding the discrete token indices to generate the discrete indices string.

9. The method according to any of claims 1-8, wherein selectively encoding the input visual signal in the quality-enhancement processing stream to generate the continuous latent string comprises:patchifying the input visual signal into a sequence of blocks;generating, for each block, an embedded continuous latent feature tensor; computing a tokenized continuous latent representation from the embedded continuous latent feature tensor; andapplying quantization and entropy encoding to the tokenized continuous latent representation to generate the continuous latent string.Atty. Docket No. 4502-86601 (6000767PCT02)10. The method according to any of claims 1-9, wherein selectively encoding the input visual signal in the quality-enhancement processing stream is selected when the task performance is below a threshold.

11. A method implemented by a receiving device for task-oriented image or video processing, the method comprising:receiving, from a sending device, a discrete indices string of an input visual signal generated by sending device using a base tokenized visual representation (TVR) framework; receiving, when generated by the sending device using a quality-enhancement processing stream, a continuous latent string wherein the discrete indices string and the continuous latent string are applicable to all downstream tasks;decoding the discrete indices string to recover discrete latent features; and performing one or more downstream tasks and reconstruction of the input visual signal based on the discrete latent features and when available the continuous latent string.

12. The method of claim 11, further comprising:receiving, when generated by the sending device, one or more task-adaptive auxiliary tokens that are modular and independent of the base TVR framework, wherein the one or more task-adaptive auxiliary tokens comprise one or more task-oriented prompt tokens that provide enhanced semantic representation for at least one specific downstream task; andperforming the at least one specific downstream task using both the discrete latent features, the one or more task-oriented prompt tokens, and when available the continuous latent string.

13. The method of claim 12, wherein the one or more task-adaptive auxiliary tokens is based on a text description of the input visual signal.

14. The method according to any of claims 12-13, wherein the one or more task-adaptive auxiliary tokens is based on an audio signal associated with the input visual signal.

15. The method according to any of claims 12-14, wherein the one or more task-adaptive auxiliary tokens is based on additional data related to the input visual signal that assists or improves performance of the at least one specific downstream task.Atty. Docket No. 4502-86601 (6000767PCT02)16. The method according to any of claims 12-15, wherein the one or more task-adaptive auxiliary tokens comprise an additional prompt token that is further used in the reconstruction of the input visual signal.

17. The method according to any of claims 11-16, wherein the base TVR framework is used for various tasks and does not require retraining when the one or more task-adaptive auxiliary tokens are used.

18. The method according to any of claims 11-17, wherein decoding the discrete indices string comprises:decoding the discrete indices string to recover discrete token indices; andretrieving a recovered discrete latent feature tensor from a pre-trained discrete codebook for each recovered token index.

19. The method according to any of claims 11-18, wherein, when the continuous latent string is received, the method further comprises:decoding the continuous latent string to recover a decoded tokenized continuous latent representation; andproviding the decoded tokenized continuous latent representation to at least one of a task-oriented tokenization module, a task execution module, or a reconstruction module.

20. The method according to any of claims 12-19, further comprising:computing, when the one or more task-oriented prompt tokens are received, an auxiliary task-oriented latent feature; andperforming the at least one specific downstream task based on the auxiliary task-oriented latent feature, the discrete latent features and, when received, the continuous latent string.

21. An apparatus, comprising:one or more processors or processing means; anda memory or storage means storing instructions that, when executed by the one or more processors or processing means, cause the apparatus to perform any one of the methods in claims 1-10.Atty. Docket No. 4502-86601 (6000767PCT02)22. An apparatus, comprising:one or more processors or processing means; anda memory or storage means storing instructions that, when executed by the one or more processors or processing means, cause the apparatus to perform any one of the methods in claims 11-20.

23. A computer program product comprising computer-executable instructions stored on a non-transitory computer-readable storage medium and that, when executed by one or more processors or processing means of an apparatus, cause the apparatus to perform any one of the methods in claims 1-10.

24. A computer program product comprising computer-executable instructions stored on a non-transitory computer-readable storage medium and that, when executed by one or more processors or processing means of an apparatus, cause the apparatus to perform any one of the methods in claims 11-20.

25. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause an encoder to perform operations according to any of claims 1-10.

26. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause a decoder to perform operations according to any of claims 11-20.

27. An apparatus configured to perform task-oriented image or video compression, the apparatus comprising:an encoder or encoding means configured to:encode an input visual signal using a base tokenized visual representation (TVR) framework to generate a discrete indices string; andselectively encode, based on a quality requirement of a reconstructed input visual signal, the input visual signal in a quality-enhancement processing stream to generate a continuous latent string to improve at least one of a task performance or a reconstruction of the input visual signal, wherein the discrete indices string and the continuous latent string are applicable to all downstream tasks; andAtty. Docket No. 4502-86601 (6000767PCT02)a transmitter or transmitting means configured to transmit the discrete indices string and, when generated, the continuous latent string to a receiving device.

28. The apparatus of claim 27, wherein the encoder or the encoding means is further configured to generate, based on task-specific requirements, one or more task-adaptive auxiliary tokens that are modular and independent of the base TVR framework, wherein the one or more task-adaptive auxiliary tokens provide enhanced semantic representation for at least one specific downstream task; andwherein the transmitter or the transmitting means is configured to transmit the one or more task-adaptive auxiliary tokens along with the discrete indices string and, when generated, the continuous latent string to the receiving device.

29. The apparatus according to any of claims 27-28, wherein the encoder or the encoding means is configured to encode the input visual signal in the quality-enhancement processing stream when the task performance is below a threshold.

30. An apparatus configured to perform task-oriented image or video processing, the apparatus comprising:a receiver or receiving means configured to:receive, from a sending device, a discrete indices string of an input visual signal generated by sending device using a base tokenized visual representation (TVR) framework;receive, when generated by the sending device using a quality-enhancement processing stream, a continuous latent string wherein the discrete indices string and the continuous latent string are applicable to all downstream tasks;a decoder or decoding means configured to decode the discrete indices string to recover discrete latent features; andone or more processors or processing means configured to perform one or more downstream tasks and reconstruction of the input visual signal based on the discrete latent features and when available the continuous latent string.

31. The apparatus of claim 30, wherein the receiver or the receiving means is further configured to receive, when generated by the sending device, one or more task-adaptive auxiliaryAtty. Docket No. 4502-86601 (6000767PCT02)tokens that are modular and independent of the base TVR framework, wherein the one or more task-adaptive auxiliary tokens comprise one or more task-oriented prompt tokens that provide enhanced semantic representation for at least one specific downstream task; andwherein the one or more processors or processing means is further configured to perform the at least one specific downstream task using both the discrete latent features, the one or more task-oriented prompt tokens, and when available the continuous latent string.

32. The apparatus according to any of claims 30-31, wherein, when the continuous latent string is received, the decoder or the decoding means is further configured to:decode the continuous latent string to recover a decoded tokenized continuous latent representation; andprovide the decoded tokenized continuous latent representation to at least one of a task-oriented tokenization module, a task execution module, or a reconstruction module.

32. The apparatus according to any of claims 30-32, wherein the one or more processors or processing means is further configured to:compute, when the one or more task-oriented prompt tokens are received, an auxiliary task-oriented latent feature; andperform the at least one specific downstream task based on the auxiliary task-oriented latent feature, the discrete latent features and, when received, the continuous latent string.