Music Compression Training with Decoupled Discrete Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The quality of discrete features in music processing tasks significantly affects the results, and existing methods struggle to effectively encode and decode music content to ensure high fidelity and rich expressiveness.

Innovation Solution

A music compression system is trained by obtaining encoded representations, processing them with discrete encoders to generate features, decoding these features with discrete decoders, and adjusting parameters based on training loss to improve fidelity and expressiveness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If music content is encoded into discrete features using existing methods, then the encoding process can be completed, but the quality of discrete features is insufficient to ensure high fidelity and rich expressiveness

Engineering Contradiction:
Improvequality of discrete featuresVSAvoidfidelity and expressiveness of music content
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent segments the music content into multiple independent tracks (e.g., vocals, instruments, drums) and processes each track separately through dedicated encoders and decoders. This segmentation allows each component to be optimized independently, improving the overall quality of discrete features while maintaining high fidelity and expressiveness of the complete music content.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different encoding and decoding parameters to different music tracks based on their specific characteristics. Each track receives localized quality optimization through dedicated discrete encoders and decoders that are tuned to the specific audio characteristics of that track, thereby improving the overall quality of discrete features while preserving the unique qualities of each music component.

Inventive Principle:
Principle #3Local quality

2Manufacturing precision

If a music compression system is trained with multiple discrete encoders and decoders for different music tracks, then the quality of audio compression is enhanced, but the device complexity increases

Engineering Contradiction:
Improvequality of audio compressionVSAvoidnumber of discrete encoders and decoders
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent designs a universal training framework that can accommodate multiple discrete encoders and decoders for different music tracks while using a unified loss function and optimization process. The system learns to handle multiple track types through a common architectural paradigm, reducing the practical complexity burden despite having multiple specialized components.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements a comprehensive feedback mechanism through the training process where the system continuously adjusts encoder and decoder parameters based on reconstruction quality metrics. This feedback loop allows the system to optimize performance automatically, reducing the need for manual tuning and simplifying the operational complexity of managing multiple encoders and decoders.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20260073927A1Method, apparatus, device and storage medium of training a music compression system
Publication Date: 2026.03.12 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20260073927A1 patent drawing
  • US20260073927A1 patent drawing
  • US20260073927A1 patent drawing

AI summary

Embodiments of the disclosue relate to a method, apparatus, device, and storage medium of training a music compression system. The method includes: obtaining a first encoded representation associated with training music content; processing the first encoded representation with the discrete encoder to generate a first set of discrete features corresponding to first music data and a second set of discrete features corresponding to second music data; decoding the first set of discrete features with the first discrete decoder to obtain a first audio feature corresponding to the first music data, and decoding the second set of discrete features with the second discrete decoder to obtain a second audio feature corresponding to the second music data; and determining a training loss based on the first audio feature, the second audio feature and the training music content, and adjusting parameters of the discrete encoder and the discrete decoders based on training loss.