Neural Video Compression With Controlled Auxiliary Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image and video compression technologies struggle to maintain quality while efficiently compressing data, especially when machines analyze video for events and objects, as they often prioritize bitrate over human perceptual ability.

Innovation Solution

Utilizing a neural network-based processor to generate and combine features and auxiliary features, with controlling values to optimize data units and enhance compression efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If traditional compression technologies are used to compress video data, then the bitrate is reduced, but the visual quality deteriorates

Engineering Contradiction:
ImprovebitrateVSAvoidvisual quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent transforms the compression approach by changing from traditional bitrate-based parameters to perceptual quality parameters. The neural network processes video data based on human visual system characteristics, transforming the optimization target from file size to perceived quality metrics, thereby achieving quality preservation at lower bitrates

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces traditional mechanical compression algorithms with a neural network-based system. The neural network learns optimal compression strategies by training on large datasets, substituting rule-based mechanical compression with adaptive intelligent compression that maintains visual quality while reducing bitrate

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If compression is optimized for human perceptual ability, then visual quality is maintained, but machine analysis capability deteriorates

Engineering Contradiction:
Improvevisual qualityVSAvoidmachine analysis capability
Core Design Contradiction:
Manufacturing precisionVSMeasurement precision

Solution Approach 1:

The patent segments the video processing into distinct neural network components: one branch optimized for perceptual quality reconstruction and another branch preserving machine-analysis-relevant features. This segmentation allows simultaneous optimization for both human viewing and machine analysis without compromise

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different processing strategies to different parts of the video data based on their importance. Regions critical for machine analysis (such as object boundaries, text, key features) are processed with higher fidelity to preserve analysis capability, while less critical regions use more aggressive compression

Inventive Principle:
Principle #3Local quality

3Productivity

If neural network processing is applied to generate features, then compression efficiency is improved, but computational complexity increases

Engineering Contradiction:
Improvecompression efficiencyVSAvoidcomputational complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent performs feature extraction and compression optimization during an offline training phase using large datasets. The neural network learns optimal compression parameters and feature representations in advance, so that during actual video compression, the pre-trained model can efficiently process videos without requiring complex real-time computations

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses a trained neural network model that can be copied and deployed across multiple devices. Once trained on powerful hardware, the model weights and architecture can be copied to various deployment environments, allowing complex compression tasks to be performed efficiently without requiring the original training computational resources

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20260095581A1A method, an apparatus and a computer program product for image and video processing using a neural network
Publication Date: 2026.04.02 NOKIA TECHNOLOGIES OY
  • US20260095581A1 patent drawing
  • US20260095581A1 patent drawing
  • US20260095581A1 patent drawing

AI summary

The embodiments relate to a method comprising receiving one or more data units; receiving one or more auxiliary data units; processing the data units by a first portion of a neural network based processor to generate a set of features; processing the auxiliary data units by second respective portions of the neural network based processor to generate sets of auxiliary features; determining controlling values associated to auxiliary data units or to the one or more sets of auxiliary features; controlling the auxiliary data units or the sets of auxiliary features according to the controlling values to generate sets of controlled auxiliary features; combining an output corresponding to the set of features and output corresponding to the sets of controlled auxiliary features into sets of combined features to be processed by further portions of the neural network based processor, and generating a signal comprising the one or more controlling values.