Reference-Guided Vocal Loudness Matching for Audio Mixing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio synthesis methods require professional expertise and are inefficient due to manual processing, leading to low synthesis efficiency and suboptimal sound quality, especially when mixing human voice with accompaniment.

Innovation Solution

An audio synthesis method that adjusts the loudness and spectral characteristics of human voice audio based on a reference audio's loudness range, using signal processing to align and mix the voice with accompaniment, enhancing sound quality and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual synthesis by professional audio engineers is used, then sound quality can be improved, but synthesis efficiency deteriorates due to long synthesis duration

Engineering Contradiction:
Improvesound qualityVSAvoidsynthesis efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system performs automatic audio synthesis by extracting human voice and accompaniment from reference audio, automatically determining loudness ranges, comparing and adjusting audio parameters, and mixing the final output without requiring professional audio engineers to manually intervene in each synthesis process

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system automatically determines loudness ranges for both human voice and accompaniment, compares these ranges, and adjusts the audio parameters (particularly loudness) to ensure proper mixing ratios and sound quality without manual intervention

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If manual audio synthesis process is used, then audio quality can be controlled, but time consumption increases leading to low efficiency

Engineering Contradiction:
Improveaudio qualityVSAvoidsynthesis duration
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical process of audio engineering with an automated computer-based system that uses signal processing algorithms to extract audio components, analyze loudness characteristics, and perform mixing operations automatically

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system determines loudness ranges for both human voice and accompaniment, compares these ranges to identify discrepancies, and uses this comparison feedback to automatically adjust the mixing parameters and achieve balanced audio output

Inventive Principle:
Principle #23Feedback

3Manufacturing precision

If professional audio engineering systems are used, then synthesis quality can be maintained, but operation complexity increases due to tight professional restrictions

Engineering Contradiction:
Improvesynthesis qualityVSAvoidoperation complexity
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The system automatically performs all audio synthesis operations including voice extraction, accompaniment separation, loudness analysis, parameter comparison, and mixing without requiring professional audio engineers to manually control each parameter or make artistic decisions

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system provides a universal audio synthesis solution that can process various types of reference audio (songs, speeches, etc.) and automatically adapt to different audio characteristics through automated loudness range determination and parameter adjustment, making it accessible to non-professionals

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20260045270A1Audio synthesis method and apparatus, electronic device, medium, and program product
Publication Date: 2026.02.12 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20260045270A1 patent drawing
  • US20260045270A1 patent drawing
  • US20260045270A1 patent drawing

AI summary

The present disclosure relates to the technical field of music engineering, and provides an audio synthesis method and apparatus, an electronic device, a medium, and a program product. The present disclosure provides an audio synthesis method, including: acquiring a first human voice audio and an accompaniment audio from reference audio; determining a first loudness range based on loudness of the first human voice audio; acquiring second human voice audio corresponding to the reference audio; determining a second loudness range based on loudness of the second human voice audio; adjusting the loudness of the second human voice audio based on a first comparison result between the second loudness range and the first loudness range to obtain first target human voice audio; and mixing the first target human voice audio and the accompaniment audio to obtain target audio.