Systems and methods for intermediate feature compression with activation signal preservation

WO2026206895A1PCT designated stage Publication Date: 2026-10-01OP SOLUTIONS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/020465
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-24
Filing Date
2026-03-24
Publication Date
2026-10-01

Smart Images

  • Figure US2026020465_01102026_PF_FP_ABST
    Figure US2026020465_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods are provided for encoding and decoding signals in a Feature Coding for Machines (FCM) video system with activation preserving processing. Important pixels can be identified as active feature pixels which have a pixel intensity that departs significantly higher or lower from the mean pixel values of a frame. An activation‑preserving preprocessing module is provided to selectively reduce coding complexity of non‑important pixels while preserving active feature pixels according to the activation‑signal inference.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] SYSTEMS AND METHODS FOR INTERMEDIATE FEATURE COMPRESSION WITH ACTIVATION SIGNAL PRESERVATION

[0002] Statement of Related Applications

[0003] The present application claims the benefit of priority to U.S. provisional patent application Serial No. 63 / 776,725, filed on March 24, 2025, and titled Systems and Methods for Intermediate Feature Compression with Activation Signal Preservation, the entirety of which is hereby incorporated by reference.

[0004] Background of the Disclosure

[0005] As the number and scale of deployed video sensors / devices increases, an increasing amount of video is expected to be processed by machines and for machine use rather than human viewing. Indeed, an exemplary system or solution employing thousands of cameras would produce significant amounts of video that cannot be monitored by humans in a cost- effective manner. Because of the significant volume of the data, an efficient compression system is desirable for Machine-to-Machine (M2M) communication. To standardize the coding for machines to facilitate a M2M communication system more efficiently, Moving Picture Experts Group (MPEG) issued a Call for Proposal (CfP) [1] for Coding for Machines in 2022.

[0006] Currently, M2M communications are typically managed by two distinct systems, namely, edge computing or local computing systems and remote computing or Video Coding for Machines (VCM) systems. In the context of edge computing systems, the complete Convolutional Neural Network (CNN) is typically executed on edge devices. However, end devices often lack the computational capacity to run extensive networks, requiring billions of operations for inference result calculations. In contrast, in Video Coding for Machine systems, the video is typically compressed initially and transmitted to a base server or cloud server with more computing resources to execute the full CNN. The calculated results are then subsequently relayed back to the edge devices. However, Video Coding for machine systems may not be able to support compute offload and may fail to fully leverage the potential of end devices, as they execute the complete CNN on the base server. To overcome these limitations for M2M communication and to enable collaborative intelligence and split computing for M2M connections, MPEG has issued a Call for Proposal (CfP) [2] for Feature Coding for Machines (FCM) technology in 2023.An example FCM system is depicted in Fig. 1. In FCM systems, a Neural Network is first split into two parts, i.e., NN Part 1 (denoted as 120 in Fig. 1) and NN Part 2 (denoted as 170 in Fig. 1). In this split CNN configuration, instead of running the full CNN on the end devices, only the NN Part 1 is executed on the resource-limited device 101. The NN Part 2 is executed on the remote server 190 with more computing resources. Since the NN Part 1 is executed on the edge devices, instead of transmitting images / videos 190 to the remote server 190, the intermediate feature data 140, extracted from NN Part 1 is compressed by FCM Encoder 130 and sent to the remote server 190. On the remote server, the compressed intermediate feature data is decoded by FCM Decoder 160 and passes through NN Part 2 to generate the prediction results 180. However, the volume of intermediate feature data 140 is usually much larger than the input video / images 190 itself and needs efficient compression before sending it to the remote server 190.

[0007] Fig. 2 is a block diagram of an FCM Encoder 130 and FCM 160 Decoder system. A typical FCM Encoder consists of three main modules, e.g., Feature Reduction 220, Feature Conversion 230, and Inner Codec 240 for encoding sequences. Similarly, FCM Decoder consists of three modules correspondingly, e.g., Feature Restoration 270, Inverse Feature Conversion 260, and Inner Codec 250 for decoding the sequences. The input of the FCM Encoder is intermediate feature data (Xt) 201 extracted from NN Part 1 120, and the output of the FCM decoder is reconstructed intermediate feature data (Xt') 280, as shown in Error! Reference source not found.. After reconstruction of the intermediate feature data (Xt'), the NN Part 2 170 is executed to generate the inference results 180. Feature data usually consists of multiple layers. And one layer usually consists of multiple channels. For instance, the backbone of detectron2 [4] (NN Part 1) consists of 4 layers. And each layer consists of 256 channels. The details of all the FCM Encoder and Decoder modules are described in more detail in the subsequent sections.

[0008] Feature Reduction

[0009] The Feature Reduction 220 module of the FCM encoder takes original feature maps and / or layers as input and outputs a reduced number of feature maps and / or layers. The feature reduction module can be a neural network, auto-encoder, or any classical statistical approaches like PCA or clustering methods to reduce the number of feature maps and / or layers.Feature Conversion

[0010] The Feature Conversion 230 module of the FCM encoder takes reduced feature maps as input and converts the feature tensor data from floating point to unsigned integers of 8-bit or 10-bit. In some cases, depending on the data, it also packs all the feature maps spatially into one frame. Then, the packed frame or converted feature maps and / or layers are converted to a suitable video format, such as YUV 4:0:0, before sending it to the encoder.

[0011] Inner Codec (Encoder)

[0012] The Inner Codec (Encoder) 240 module takes the video, such as YUV 4:0:0, as input and outputs a bitstream to transmit to the receiver. The Inner Codec (Encoder) could be a Neural Network encoder or any traditional encoder like HEVC, VVC, VVenC, or a combination of both Neural Network encoder and traditional encoder.

[0013] Inner Codec (Decoder)

[0014] The Inner Codec (Decoder) 250 module takes the received bitstream as input and decompresses the bitstream file to a suitable format, such as YUV 4:0:0. Like the Inner Codec (Encoder) module, the Inner Codec (Decoder) could be a Neural Network based decoder or any traditional decoder like HEVC, VVC, VVdeC or the combination of both Neural Network based decoder and traditional decoder.

[0015] Inverse Feature Conversion

[0016] The Inverse Feature Conversion 260 module takes the decoded YUV 4:0:0 as input and unpacks (if necessary) and converts the unsigned integers to floating point values.

[0017] Feature Restoration

[0018] The Feature Restoration 270 module takes the feature tensors from the Inverse Feature Conversion module as input and restores the feature tensors from reduced size to the original size. Like the Feature Reduction module 220, the Feature Restoration module can also be a Neural Network, Auto-Decoder, or just classical statistical approaches like Inverse PCA or de-clustering to restore the original size of feature maps and layers.

[0019] Summary of the Disclosure

[0020] In one embodiment of the present disclosure, an FCM encoder with

[0021] activation-preserving compression of intermediate features in a split neural network employs a NN Part 1 configured to process an input image or video frame to generate an intermediate feature tensor. Optionally, a feature reduction module configured to reduce dimensionality of the intermediate feature tensor. A feature conversion module may be provided and configured to convert reduced features to an integer bit-depth and form a packed featureframe. An activation-signal inference module configured to identify active feature pixels in the packed feature frame. An activation-preserving preprocessing module configured to selectively reduce coding complexity of non-important pixels while preserving active feature pixels according to the activation-signal inference. An inner codec encoder is coupled to the activation preserving preprocessing module and is configured to encode the preprocessed packed feature frame into a bitstream.

[0022] The activation-signal inference module may be configured to identify active feature pixels using at least one of: distance from a per-channel mean, intra-region variability, gradient-based saliency, or attention scores. In some embodiments, the active feature pixels may be selected as the top 10% highest-value pixels and a bottom 10% lowest-value pixels.

[0023] In various embodiments, the activation-signal inference operates on at least one of an intermediate feature tensor, a reduced feature tensor, or the packed feature frame.

[0024] In some embodiments, the activation-preserving preprocessing may include applying a bilateral filter to the packed feature frame with spatial and range Gaussian parameters.

[0025] In some embodiments, the activation-preserving preprocessing may comprise a learned convolutional neural network trained to reduce complexity of non-important pixels while preserving active feature pixels.

[0026] Alternatively or additionally, the activation-preserving postprocessing comprises unsharp masking applied to a decoded packed feature frame.

[0027] In some embodiments, active feature pixel values are transmitted substantially losslessly in the bitstream to ensure preservation at the decoder.

[0028] In some embodiments, the activation-preserving Inner encoder is a learned end-to-end encoder jointly trained with a downstream task network and uses a learned soft mask for spatially adaptive quantization.

[0029] Alternatively or additionally, the activation-preserving inner encoder can be a statistically optimized video encoder that adjusts quantization parameters based on region importance or entropy.

[0030] In some embodiments, blocks with active pixels are encoded with lower quantization step sizes and blocks not including active pixels are encoded with higher quantization step sizes.

[0031] In some embodiments, the activation-preserving Inner decoder performs learned reconstruction of important regions using spatial attention to restore task-relevant information.In certain embodiments bitstream provided by the encoder includes an important-pixel list comprising, for each pixel, an x-coordinate, a y-coordinate, and a corresponding feature pixel value encoded losslessly or substantially losslessly. The encoder may also provide in the bitstream decoder augmentation instructions comprising one or more augmentation identifiers and associated parameters.

[0032] In some cases, the activation signal preprocessing module is integrated with the inner codec encoder.

[0033] The present disclosure also provides a method for encoding with activation-preserving compression of intermediate features in a split neural network, comprising generating, by NN Part 1, an intermediate feature tensor from an input image or video frame, reducing and converting the features to produce a packed feature frame, identifying, by an

[0034] activation-signal inference module, active feature pixels in the packed feature frame, selectively preserving the active feature pixels during compression. Selectively preserving active feature pixels may include at least one of applying activation-preserving preprocessing prior to Inner codec encoding, and encoding the packed feature frame with an activation-preserving Inner encoder.

[0035] A non-transitory computer-readable storage medium is disclosed for storing instructions that, when executed by one or more processors, cause the processors to perform the methods disclosed herein

[0036] An FCM decoder is provided for decoding the encoded bitstream generated by the various encoder embodiments and encoding methods described herein.

[0037] In one embodiment, and FCM decoder includes an inner codec decoder configured to decode an encoded bitstream, an activation-preserving postprocessing module configured to restore active feature pixels in a decoded packed feature frame, an inverse feature conversion module and a feature restoration module configured to reconstruct the feature tensor, and a NN Part 2 configured to consume the reconstructed features to produce a task output.

[0038] In some embodiments, the activation-preserving Inner decoder performs statistical reconstruction including dequantization and edge-enhancement or denoising filters guided by knowledge of active feature pixels.

[0039] In some cases, the activation-preserving postprocessing may include unsharp masking applied to a decoded packed feature frame.Alternatively or additionally, the activation-preserving inner decoder may perform statistical reconstruction including dequantization and edge-enhancement or denoising filters guided by knowledge of active feature pixels.

[0040] Brief Description of the Figures

[0041] To illustrate the invention, the drawings show aspects of one or more embodiments of the invention. However, it should be understood that the present disclosure is not limited to the precise arrangements and instrumentalities shown in the drawings, wherein:

[0042] Fig. l is a simplified block diagram of a system for Feature Coding for Machines (FCM);

[0043] Fig. 2 is a block diagram of a system for Feature Coding for Machines (FCM) of Fig.

[0044] 1 and further illustrating components of an exemplary FCM Encoder and FCM Decoder systems;

[0045] Fig. 3 is an image view of a monochrome, packed, feature channel map;

[0046] Fig. 4 is a representation of a section of a feature channel map showing the uncompressed signal, the state-of-the-art compressed signal, and the resulting distortion caused to active feature pixels;

[0047] Fig. 5 is a representation of a section of a feature channel map showing the uncompressed signal, the state-of-the-art compressed signal, and the compressed signal in accordance with the present disclosure and showing the preservation of the active feature pixel and the deterioration of unimportant feature pixels;

[0048] Fig. 6 is an embodiment of the present disclosure in which the exemplary system preprocesses a packed feature frame to reduce unimportant feature pixels while maintaining those pixels of high importance; and

[0049] Fig. 7 is an embodiment of the present disclosure in which the exemplary system uses a custom compression engine that compresses a packed feature frame to reduce unimportant feature pixels while maintaining those pixels of high importance;

[0050] Detailed Description of the Embodiments

[0051] Intermediate feature channels, such as those generated at the output of NN Part 1 in a split CNN FCM encoder typically contain a series of neural feature activations. These activations comprise differences in value from the mean. In the example of a lObit feature frame, the range of values is [0 - 1023], As illustrated in Figure 3, neural feature activationscan be seen as dark 305 (closer to 0) or bright pixels 310 (closer to 1023) in the packed frame 300 post feature conversion. Due to the nature of learned networks, pixels with a substantial difference from the mean will generally have a higher impact on the performance on the task network. Such pixels in the image data are referred to herein as active signal pixels. An active signal pixel in a packed feature frame following processing by NN part 1 is referred to as an active feature pixel.

[0052] An “active feature pixel” is a pixel of the packed feature frame identified by the activation-signal inference module as having high impact on downstream task performance, for example based on distance from a channel mean, intra-region variability, gradient-based saliency (e.g., Grad-CAM), attention scores, or learned / statistical criteria; such pixels are preferentially preserved with higher fidelity.

[0053] When a feature frame 300 is inserted into the inner codec 240 as described in Figure 2, existing FCM systems will typically use a standard codec such as VVC VTM that quantizes these regions without regard for the sensitive nature of highly active feature pixels. As illustrated in Figure 4, which graphically illustrates pixel intensity, this can result in a suboptimal output signal 420 as compared to input signal 410 because the inner codec introduces compression loss 430 to these high impact pixels which can ultimately reduce task performance.

[0054] Embodiments of the present disclosure preferably operate to reduce data complexity in lower impact feature pixels, i.e., those with values closer to the mean intensity, while preserving the high impact, active feature pixels in the context of split neural network inference when compared to the original feature frame signal. Such active feature pixels are deemed to be more important for a machine task. Figure 5 is a graphical illustration illustrating an example of an input signal 510, the input signal 510 as processed by conventional FCM systems 520, and that input signal 510 as processed by the present embodiments 530. Comparing signals 520 and 530, the present embodiments reduce the complexity of the lower impact pixels as shown by the increased compression 540, while preserving the active pixels as illustrated by the preservation of signal 550.

[0055] The present embodiments operate to selectively reduce non-important feature pixels while preserving edges between active feature pixels and non-important feature pixels, which can provide better compression to accuracy performance when compared to traditional compression.

[0056] The preservation of active feature pixels in intermediate feature maps 210 typically input to an FCM encoder generally requires three main stages. The first step is to understandwhat an active signal pixel looks like for the specific task network and split point from which the intermediate features are being extracted. The second step is to apply compression techniques such that active signals are preserved. This can be done via a preprocessing step or an inner encoder that applies quantization and other processes towards this goal. The final step is a decoder side optimization that further enhances these active feature signals.

[0057] In the present disclosure, these steps can be categorized as activation signal inference, active signal preserving compression and active signal restoration, each of which are discussed in further detail in connection with the embodiments disclosed in connection with Figures 6 and 7. For example, activation signal inference 655, 745, active signal preserving compression (preprocessing 660 and / or inner encoder 750), and active signal restoration (postprocessing 665 and / or inner decoder 755)

[0058] There are a multitude of ways a system with these qualities can be implemented, some of which are depicted in Figure 6 and Figure 7. Below a description of each system in accordance with the present disclosure is provided.

[0059] A first embodiment of an improved FCM encoder / decoder system in accordance with the present disclosure is illustrated in the simplified block diagram of Figure 6. A video frame input 605 is applied to a NN Part 1 610 which is a first part of a split neural network similar to those used in conventional FCM encoder systems. The output of NN Part 1 610 is a plurality of feature maps, such as illustrated in Fig. 3. A feature reduction module 615 and feature conversion module 620 similar to that described in connection with Fig. 2 are also provided. The output of feature conversion module 620 is coupled to an active signal preserving preprocessing module 660 in accordance with the present disclosure. The NN Part 1 610 and feature conversion module 620 are also coupled to an activation signal inference module 655, the output of which is coupled to the active signal preserving preprocessing module 660. The output of the active signal preserving preprocessing module 660 is coupled to an inner codec 625, which operates similarly to that discussed above in connection with Fig. 2.

[0060] The encoded bitstream from inner codec 625 is provided over a channel to the decoder. An inner codec 630, similar to that described in connection with Fig. 2, performs inverse operations to that of inner codec 625 to decode the encoded bitstream. An active signal preserving postprocessing module 665 is now interposed between the inner codec 630 and the inverse feature conversion module 635. The output of active signal postprocessing module 665 is coupled to inverse feature conversion module 635, feature restoration module640 and Neural Network Part 2645 which operate in a manner substantially as described above in connection with Figures 1 and 2.

[0061] The processing at the encoder will apply augmentations to the intermediate feature frame such that it reduces its complexity for the encoder while maintaining active feature pixels (e.g., Fig. 5, 550). In connection with Fig. 6, the embodiment does so in the following stages.

[0062] Activation Signal Inference 655: Some level of inference is required to determine what constitutes an active feature for the specific task network. This can involve applying methods such as measuring the distance from the mean pixel value, analyzing variability within feature map regions, monitoring activations using techniques like Grad-CAM, or applying other statistical or learned approaches to identify which outputs from the first part of the neural network are important for the final computation performed by the second part. An example of this would be analyzing 10% highest and 10% lowest feature pixels by pixel value, these will provide a baseline of highly important pixel values that will preferably be transmitted with substantially no signal loss.

[0063] Active Signal Preserving Preprocessing 660: This stage operates to reduce complexity of the intermediate feature frame for better compression performance by the Inner Codec while maintaining the active feature pixels that will heavily influence the neural network part 2 computation. This objective can be completed in the following ways:

[0064] In one exemplary embodiment, a learned approach can be employed where a network is trained to reduce the complexity of the feature pixels while preserving important task related features. For example, a convolutional neural network can be trained such that a packed feature frame 300 is presented as an input and the output would be a denoised version that maintains task network performance by selectively reducing complexity from unimportant pixels the network learned over training iterations.

[0065] In addition or alternatively, statistical approaches such as edge preserving denoising filters can be performed on the intermediate feature frame. For example, one such approach can utilize a bilateral filter on the packed frame of reduced features prior to the inner codec. With x G IRXHXM / denoting a packed frame of reduced features, a frame with the bilateral filter applied x_f is given by

[0066]

[0067] where S is a window of pixels surrounding x[i], gpis the Gaussian function parameterized by p for coordinate range between surrounding pixels, and gqis the Gaussian function parameterized by q for value range between surrounding pixels. The resulting frame Xf is passed to the inner codec instead of x. The application of the bilateral filter removes unnecessary patterns that increase coding complexity and softens sharp edges between channels. These operations could be completed at the tensor level or at the image format level post feature conversion.

[0068] The decoder of Figure 6 includes a new active signal preserving postprocessing module 665. Postprocessing the decoded feature pixels is a means of increasing the task network performance by restoring high importance feature pixels. This can be completed in the following ways:

[0069] In one embodiment, a learned approach where a network is trained to reconstruct feature pixels that may have undergone loss in the encoding stage can be used. For example, a convolutional neural network can be trained such that a decoded packed feature frame is presented as an input and the output would be an altered version that has certain feature pixels changed such that performance of the task network increases. The pixel changes would be learned over training iterations.

[0070] In addition or alternatively, statistical approaches such as unsharp masking and other enhancement filters can be employed to improve the visual quality of regions with active or important pixels. For instance, unsharp masking can be applied on the decoder side using a Gaussian blur with a radius of 1.0 and an amount of 1.5 to enhance edges and fine details. In addition to such enhancement techniques, another approach involves directly transmitting pixel values or critical feature information to the decoder, preserving a lossless representation of selected high-importance feature pixels. This ensures that regions critical to object detection or interpretation remain visually and semantically intact.

[0071] It will be appreciated that feature pixels can be at the tensor level directly after extraction from NN Part 1 all the way up to the image level post feature conversion and active signal inference can be applied at different points in the exemplary systems. For example, activation signal inference 655 can be performed directly on the intermediate feature maps, giving direct insight into output of NN Part 1 610. It could also be performed post feature reduction 615, which will give a direct inference to the reduced feature map. Active signal inference may also be performed on the post conversion 620 packed feature frame, which can give insight of what pixels to preserve on the final pixels right before theencoder. The same logic can also be applied to other stages of this system such as preprocessing 660 and postprocessing 665.

[0072] An alternate embodiment of an improved FCM encoder / decoder system in accordance with the present disclosure is illustrated in the simplified block diagram of Figure 7. A video frame input 705 is applied to a NN Part 1 710 which is a first part of a split neural network similar to those used in conventional FCM encoder systems. The output of NN Part 1 710 is a plurality of feature maps, such as illustrated in Fig. 3. A feature reduction module 715 and feature conversion module 720 similar to that described in connection with Fig. 2 are also provided. The output of feature conversion module 720 is coupled to an active signal preserving inner encoder 750 in accordance with the present disclosure. The NN Part 1 710 and feature conversion module 720 are also coupled to an activation signal inference module 745, the output of which is coupled to the active signal preserving inner encoder 750.

[0073] The encoded bitstream from inner codec 750 is provided over a channel to the decoder. An active signal preserving inner decoder 755 performs inverse operations to that of inner codec 750 to decode the encoded bitstream. The decoded bitstream from the inner codec 755 is coupled to an inverse feature conversion module 725, feature restoration module 730 and Neural Network Part 2735 which operate in a manner described above in connection with Figures 1 and 2.

[0074] The embodiment of Figure 7 employs an optimized inner codec mode that compresses the intermediate feature frame directly while maintaining active feature pixels. It can do so in the following stages.

[0075] The activation signal inference module 745 performs some inference to determine what constitutes an active feature for the task network. This can involve methods such as measuring the distance from the mean pixel value, analyzing variability within feature map regions, monitoring activations using techniques like Grad-CAM, or applying other statistical or learned approaches to identify which outputs from the first part of the neural network are important for the final computation performed by the second part.

[0076] The embodiment of Figure 7 employs an active signal preserving inner encoder 750. This module operates to compress the intermediate feature frame while preserving the feature pixels that significantly impact the downstream computation in NN part 2. The inner encoder 750 can be implemented as an optimized video encoder, in one or both of the following ways:In one embodiment, a learned end-to-end encoder can be employed where the inner encoder is trained end-to-end alongside the task network (e.g., an object detector or classifier), enabling it to learn a feature importance from supervision. The encoder 750 preferably selectively compresses the feature maps by applying spatially adaptive quantization or attention-based gating mechanisms. For example, using a learned soft mask, feature values corresponding to high-importance pixels (e.g., those with high gradient magnitudes or attention scores) are preserved with higher fidelity, while less relevant regions are downscaled or quantized more aggressively. The feature importance can be modeled via an auxiliary branch trained with the main task loss to emphasize semantic consistency in critical areas.

[0077] In addition or alternatively, the inner encoder 750 can be implemented with a statistically optimized video encoder, for which the rate-distortion optimization is modified to calculate quantization (e.g. QP - a quantization parameter in the VVC). This method relies on algorithms to identify structurally important regions during compression. For instance, quantization parameters can be tuned, for example, assigning lower quantization step sizes (e.g., QP = 18) to blocks with active feature pixels 550 or higher entropy, while using higher step sizes (e.g., QP = 32) for regions with low impact activations, or lower entropy. The relationship between the statistical measure (e.g. entropy) and QP can be linear (inverse) or represented with an optimization formula that is derived from the extensive testing on the representative video datasets.

[0078] These operations can be performed either directly on the feature tensors in the end-to-end training case or after converting the features into an image-compatible format.

[0079] The decoder of Figure 7 includes an active signal preserving inner decoder 755. The inner decoder 755 reconstructs feature pixels encoded at the encoder during the active signal preserving inner encoder 750 stage. Depending on the implementation of the inner encoder 750, the bitstream could be decoded in the following ways:

[0080] In one embodiment, a learned approach can be applied in which the inner decoder 755 is jointly trained with the inner encoder 750 in an end-to-end fashion to reconstruct the active feature pixel 550-optimized intermediate feature frame 330. This reconstruction process typically involves a lightweight convolutional decoder network that leverages spatial attention or learned masks to focus on restoring task-relevant information. The decoder may use task loss (e.g., detection / classification loss), to ensure both pixel-level accuracy and downstream task performance are preserved.Additionally or alternatively, a statistical approach can be employed where traditional bitstream decoding is performed on the active feature pixel optimized intermediate feature frame such as dequantizing, while maintaining knowledge of the active feature pixels for optimal decoding and partial enhancement, such as applying certain filters that can enhance task network performance, such as edge enhancement, denoising, etc.

[0081] Exemplary Bitstream Syntax:

[0082]

[0083] This bitstream sends active feature pixel values to the decoder for lossless reconstruction.

[0084] important_feature_pixels_len: total active feature pixels that need lossless transmission.

[0085] Pixel coordinate x: active feature pixel coordinate on image x plane.

[0086] Pixel_coordinate_y: active feature pixel coordinate on image y plane.

[0087] Feature Pixel value: lossless value for pixel.

[0088]

[0089]

[0090] This bitstream can be conveyed to the decoder a series of augmentation instructions for filters or further processing based on processing done on the encoder.

[0091] decoder augmentations len: total operations needed to be performed by decoder for feature pixel enhancement.

[0092] Decoder augmentation id: specific decoder process to be performed.

[0093] Parameter application value: strength of filter or other parameter for augmentation application.

[0094] Some embodiments of the present disclosure may include / and or be embodied by non- transitory computer program products (i.e., physically embodied computer program products) that store instructions, which when executed by one or more data processors of one or more computing systems, cause at least one data processor to perform operations herein. Such non-transitory computer program products (i.e., physically embodied computer program products) may store instructions, which when executed by one or more data processors of one or more computing systems, causes at least one data processor to perform operations, and / or steps thereof described in this disclosure, including without limitation any operations described above and / or any operations of the FCM decoder and / or FCM encoder may be configured to perform. Similarly, computer systems are also described that may include one or more data processors and memory coupled to the one or more data processors. The memory may temporarily or permanently store instructions that cause at least one processor to perform one or more of the operations described herein. In addition, methods can be implemented by one or more data processors either within a single computing system or distributed among two or more computing systems. Such computing systems can be connected and can exchange data and / or commands or other instructions orthe like via one or more connections, including a connection over a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, or the like), via a direct connection between one or more of the multiple computing systems, or the like.

[0095] Any one or more of the aspects and embodiments described herein may be conveniently implemented using one or more machines (e.g., one or more computing devices that are utilized as a user computing device for an electronic document, one or more server devices, such as a document server, etc.) programmed according to the teachings of the present specification, as will be apparent to those of ordinary skill in the computer art. Appropriate software coding can readily be prepared by skilled programmers based on the teachings of the present disclosure, as will be apparent to those of ordinary skill in the software art. Aspects and implementations discussed above employing software and / or software modules may also include appropriate hardware for assisting in the implementation of the machine executable instructions of the software and / or software module.

[0096] Such software may be a computer program product that employs a machine-readable storage medium. A machine-readable storage medium may be any medium that is capable of storing and / or encoding a sequence of instructions for execution by a machine (e.g., a computing device) and that causes the machine to perform any one of the methodologies and / or embodiments described herein. Examples of a machine-readable storage medium include, but are not limited to, a magnetic disk, an optical disc (e.g., CD, CD-R, DVD, DVD-R, etc.), a magneto-optical disk, a read-only memory “ROM” device, a random-access memory “RAM” device, a magnetic card, an optical card, a solid-state memory device, an EPROM, an EEPROM, and any combinations thereof. A machine-readable medium, as used herein, is intended to include a single medium as well as a collection of physically separate media, such as, for example, a collection of compact discs or one or more hard disk drives in combination with a computer memory. As used herein, a machine-readable storage medium does not include transitory forms of signal transmission.

[0097] Such software may also include information (e.g., data) carried as a data signal on a data carrier, such as a carrier wave. For example, machine-executable information may be included as a data-carrying signal embodied in a data carrier in which the signal encodes a sequence of instruction, or portion thereof, for execution by a machine (e.g., a computing device) and any related information (e.g., data structures and data) that causes the machine to perform any one of the methodologies and / or embodiments described herein.Examples of a computing device include, but are not limited to, an electronic book reading device, a computer workstation, a terminal computer, a server computer, a handheld device (e.g., a tablet computer, a smartphone, etc.), a web appliance, a network router, a network switch, a network bridge, any machine capable of executing a sequence of instructions that specify an action to be taken by that machine, and any combinations thereof. In one example, a computing device may include and / or be included in a kiosk.

[0098] The foregoing has been a detailed description of illustrative embodiments of the invention. Various modifications and additions can be made without departing from the spirit and scope of this invention. Features of each of the various embodiments described above may be combined with features of other described embodiments as appropriate in order to provide a multiplicity of feature combinations in associated new embodiments. Furthermore, while the foregoing describes a number of separate embodiments, what has been described herein is merely illustrative of the application of the principles of the present invention. Additionally, although particular methods herein may be illustrated and / or described as being performed in a specific order, the ordering is highly variable within ordinary skill to achieve methods, systems, and software according to the present disclosure. Accordingly, this description is meant to be taken only by way of example, and not to otherwise limit the scope of this invention.

[0099] The following references cited above may be relevant as background or contextual material and are hereby incorporated by reference in their entireties:

[0100] [1] Yuan Zhang, Manouchehr Rafie, Shan Liu, Christopher Hohmann, “[N00164] Call for Proposals on Video Coding for Machines.” ISO / IEC JTC 1 / SC 29 / WG 2, Jan. 2022.

[0101] [2] C. Rosewame (Canon) and Y. Zhang (China Telecom), “[N00282] Call for Proposals on Feature Compression for Video Coding for Machines.” ISO / IEC JTC 1 / SC 29 / WG 2, Jan. 2023.

[0102] [3] WG 04 MPEG Video coding, “[N00460] Algorithm description of FCTM.” Jan. 2024.

[0103] [4] Y. Wu, A. Kirillov, F. Massa, W.-Y. Lo, and R. Girshick, “Detectron2,” 2019, Feb. 2024. [Online],

[0104] Exemplary embodiments have been disclosed above and illustrated in the accompanying drawings. It will be understood by those skilled in the art that variouschanges, omissions and additions may be made to that which is specifically disclosed herein without departing from the spirit and scope of the present invention.

Claims

What is Claimed is:

1. An encoder with activation-preserving compression of intermediate features in a split neural network, comprising:a NN Part 1 configured to process an input image or video frame to generate an intermediate feature tensor;a feature reduction module configured to reduce dimensionality of the intermediate feature tensor;a feature conversion module configured to convert reduced features to an integer bit-depth and form a packed feature frame;an activation-signal inference module configured to identify active feature pixels in the packed feature frame;an activation-preserving preprocessing module configured to selectively reduce coding complexity of non-important pixels while preserving active feature pixels according to the activation-signal inference; andan inner codec encoder configured to encode the preprocessed packed feature frame into a bitstream.

2. The encoder of claim 1, wherein the activation-signal inference module identifies active feature pixels using at least one of: distance from a per-channel mean, intra-region variability, gradient-based saliency, or attention scores.

3. The encoder of claim 1, wherein the active feature pixels comprise at least one of a top 10% highest-value pixels and a bottom 10% lowest-value pixels.

4. The encoder of claim 1, wherein the activation-signal inference operates on at least one of: the intermediate feature tensor, the reduced feature tensor, or the packed feature frame.

5. The encoder of claim 1, wherein the activation-preserving preprocessing comprises applying a bilateral filter to the packed feature frame with spatial and range Gaussian parameters.

6. The encoder of claim 1, wherein the activation-preserving preprocessing comprises a learned convolutional neural network trained to reduce complexity of non-important pixels while preserving active feature pixels.

7. The encoder of claim 1, wherein the activation-preserving postprocessing comprises unsharp masking applied to a decoded packed feature frame.

8. The encoder of claim 1, wherein active feature pixel values are transmitted substantially losslessly in the bitstream to ensure preservation at the decoder.

9. The encoder of claim 1, wherein the activation-preserving Inner encoder is a learned end-to-end encoder jointly trained with a downstream task network and uses a learned soft mask for spatially adaptive quantization.

10. The encoder of claim 1, wherein the activation-preserving Inner encoder is a statistically optimized video encoder that adjusts quantization parameters based on region importance or entropy.

11. The encoder of claim 1, wherein blocks with active pixels are encoded with lower quantization step sizes and blocks not including active pixels are encoded with higher quantization step sizes.

12. The encoder of claim 1, wherein the activation-preserving Inner decoder performs learned reconstruction of important regions using spatial attention to restore task-relevant information.

13. The encoder of claim 1, wherein the bitstream includes an important-pixel list comprising, for each pixel, an x-coordinate, a y-coordinate, and a corresponding feature pixel value encoded losslessly.

14. The encoder of claim 1, wherein the bitstream further includes decoder augmentation instructions comprising one or more augmentation identifiers and associated parameters.

15. The encoder of claim 1, wherein the activation signal preprocessing module is integrated with the inner codec encoder.

16. A method for encoding with activation-preserving compression of intermediate features in a split neural network, comprising:generating, by NN Part 1, an intermediate feature tensor from an input image or video frame;reducing and converting the features to produce a packed feature frame; identifying, by an activation-signal inference module, active feature pixels in the packed feature frame;selectively preserving the active feature pixels during compression by at least one of: applying activation-preserving preprocessing prior to Inner codec encoding; and encoding the packed feature frame with an activation-preserving Inner encoder;17. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the processors to perform the method of claim 16.

18. An F CM decoder compri sing :an inner codec decoder configured to decode an encoded bitstream;an activation-preserving postprocessing module configured to restore active feature pixels in a decoded packed feature frame;an inverse feature conversion module and a feature restoration module configured to reconstruct the feature tensor; andNN Part 2 configured to consume the reconstructed features to produce a task output.

19. The FCM decoder of claim 18, wherein the activation-preserving Inner decoder performs statistical reconstruction including dequantization and edge-enhancement or denoising filters guided by knowledge of active feature pixels.

20. The decoder of claim 18, wherein the activation-preserving postprocessing comprises unsharp masking applied to a decoded packed feature frame.

21. The decoder of claim 18, wherein the activation-preserving Inner decoder performs statistical reconstruction including dequantization and edge-enhancement or denoising filters guided by knowledge of active feature pixels.