Visual Data Coding With Content-Specific Neural Weight Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current neural network-based image and video compression methods face challenges in optimizing synthesis networks for varying content types, leading to degraded performance and increased computational and storage burdens.

Innovation Solution

Implement an independent subnetwork weight selection scheme, where synthesis networks are chosen from a set of pretrained weights based on content type, and enhance entropy coding to support variable rate coding with fewer networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple synthesis networks are used to handle different content types, then coding performance is improved, but device complexity and storage requirements increase

Engineering Contradiction:
Improvecoding performanceVSAvoidnumber of synthesis networks
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the synthesis network into multiple specialized subnetworks, each trained on specific content types (e.g., natural images, screen content, medical images). Instead of using one general-purpose synthesis network, the system divides the synthesis task into specialized segments that can be selectively applied based on content type, improving performance without requiring all networks to be active simultaneously.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic selection of synthesis networks based on content type classification. The system dynamically chooses which subnetwork to apply by analyzing the input content characteristics and selecting the most appropriate pre-trained network, allowing the system to adapt to different content types without maintaining all networks in an active state, thus reducing computational complexity.

Inventive Principle:
Principle #15Dynamics

2Reliability

If multiple synthesis networks are used to handle different content types, then coding performance is improved, but storage requirements increase

Engineering Contradiction:
Improvecoding performanceVSAvoidmodel size
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent segments the synthesis network into multiple specialized subnetworks, each trained on specific content types (e.g., natural images, screen content, medical images). Instead of using one general-purpose synthesis network, the system divides the synthesis task into specialized segments that can be selectively applied based on content type, improving performance without requiring all networks to be active simultaneously.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic selection of synthesis networks based on content type classification. The system dynamically chooses which subnetwork to apply by analyzing the input content characteristics and selecting the most appropriate pre-trained network, allowing the system to adapt to different content types without maintaining all networks in an active state, thus reducing computational complexity.

Inventive Principle:
Principle #15Dynamics

3Device complexity

If a single synthesis network is used for all content types, then device complexity is reduced, but coding performance degrades

Engineering Contradiction:
Improvenumber of synthesis networksVSAvoidcoding performance
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent applies local quality by creating specialized synthesis subnetworks optimized for specific content types (natural images, screen content, medical images, etc.). Each subnetwork has locally optimized parameters and architectures tailored to its designated content type, ensuring high performance for that specific domain rather than using a generic one-size-fits-all approach.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements dynamic selection of synthesis networks based on content type classification. The system dynamically chooses which subnetwork to apply by analyzing the input content characteristics and selecting the most appropriate pre-trained network, allowing the system to adapt to different content types without maintaining all networks in an active state, thus reducing computational complexity.

Inventive Principle:
Principle #15Dynamics

4Productivity

If content-type-specific synthesis networks are implemented, then coding efficiency is improved, but ease of operation decreases

Engineering Contradiction:
Improvecoding efficiencyVSAvoidnetwork selection complexity
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent implements self-service by enabling the synthesis network selection process to be automatic and autonomous. The system automatically classifies the input content type and selects the appropriate subnetwork without requiring manual intervention or complex configuration by the user, making the system easy to operate despite having multiple specialized networks.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent implements dynamic selection of synthesis networks based on content type classification. The system dynamically chooses which subnetwork to apply by analyzing the input content characteristics and selecting the most appropriate pre-trained network, allowing the system to adapt to different content types without maintaining all networks in an active state, thus reducing computational complexity.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250247542A1Method, apparatus, and medium for visual data processing
Publication Date: 2025.07.31 DOUYIN VISION CO LTD
  • US20250247542A1 patent drawing
  • US20250247542A1 patent drawing
  • US20250247542A1 patent drawing

AI summary

Embodiments of the present disclosure provide a solution for visual data processing. A method for visual data processing is proposed. The method comprises: determining, for a conversion between visual data and a bitstream of the visual data, a target weight for use by a target module in a coding system based on information associated with the visual data, the coding system being implemented with at least one neural network; and performing the conversion by using the coding system based on the target weight.