Neural Network Audio Labeling with Dual Loss Functions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in automatically identifying music theory labels and segment boundaries in audio tracks due to subjective timing and selection variations among users, making it difficult for computing devices to accurately analyze and label new songs or audio tracks.

Innovation Solution

A neural network model is trained using a deep learning approach with a first loss function for music theory label identifications and a second loss function for segment boundary identifications, allowing for the generation of music theory labels and segment boundaries in audio tracks, even for shorter tracks or those with non-standard durations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual labeling of music structural elements is used, then labeling accuracy reflects user expertise, but subjectivity varies timing and selection among users

Engineering Contradiction:
Improvelabeling accuracyVSAvoidconsistency among users
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The system enables automatic self-labeling of music structural elements through trained neural network models. The model autonomously identifies segment boundaries and assigns music theory labels (verse, chorus, bridge, etc.) without requiring manual user input, thereby eliminating inter-user subjectivity while maintaining high labeling accuracy through learned patterns from training data.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual labeling process with an automated neural network-based system. Instead of relying on human users to manually identify and label music structural elements, the system uses trained deep learning models that process audio signals to automatically detect segments and assign labels, substituting human judgment with algorithmic analysis for consistent results.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated identification is implemented, then productivity increases, but measurement precision decreases due to subjective variations

Engineering Contradiction:
Improveautomation speedVSAvoididentification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system performs preliminary training of neural network models using extensively labeled training datasets before deployment. During this offline preparation phase, the models learn accurate patterns of music structural elements from diverse examples. When deployed for automated identification, these pre-trained models maintain high measurement precision while enabling rapid processing of new audio tracks without manual intervention.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual music structural element identification with automated neural network analysis. The system processes audio signals through trained models that detect segment boundaries and classify musical sections, achieving both high productivity through automation and maintained precision through learned patterns from training data, eliminating the trade-off between speed and accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Ease of operation

If traditional methods are used for audio thumbnail extraction, then random excerpts are selected, but familiar elements like choruses cannot be prioritized

Engineering Contradiction:
Improvethumbnail generation simplicityVSAvoidthumbnail quality
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent replaces random excerpt selection with automated neural network-based identification of music structural elements. The system analyzes the audio track to detect segment boundaries and identify musical sections (verse, chorus, bridge), then uses this structured understanding to select representative thumbnails from prominent sections like choruses, ensuring high-quality previews that reflect the song's key elements rather than random portions.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Measurement precision

If multiple loss functions are used for training, then identification accuracy improves, but device complexity increases

Engineering Contradiction:
Improveidentification accuracyVSAvoidtraining model complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The training process is segmented into distinct objective functions: one loss function specifically for segment boundary identification and another for music theory label identification. This segmentation allows each component to be optimized independently while working together through the unified neural network architecture, achieving high overall identification accuracy without overwhelming complexity in the training procedure.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20230386437A1Neural network model for audio track label generation
Publication Date: 2023.11.30 LEMON INC(GB)
  • US20230386437A1 patent drawing
  • US20230386437A1 patent drawing
  • US20230386437A1 patent drawing

AI summary

System and methods directed to identifying music theory labels for audio tracks are described. More specifically, a first training set of audio portions may be generated from a plurality of audio tracks, segments within the plurality of audio tracks being labeled according to a plurality of music theory labels. A deep neural network model may then be trained using the first training set as an input, a first loss function for music theory label identifications of audio portions of the first training set, and a second loss function for segment boundary identifications within the audio portions of the first training set. In examples, the music theory label identifications and the segment boundary identifications are generated by the deep neural network model. A first audio track is received and segment boundary identifications and music theory labels for segments within the first audio track are generated using the deep neural network model.