Reference-Guided Vocal Loudness Matching for Audio Mixing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio synthesis methods require professional expertise and are inefficient due to manual processing, leading to low synthesis efficiency and suboptimal sound quality, especially when mixing human voice with accompaniment.
Innovation Solution
An audio synthesis method that adjusts the loudness and spectral characteristics of human voice audio based on a reference audio's loudness range, using signal processing to align and mix the voice with accompaniment, enhancing sound quality and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If manual synthesis by professional audio engineers is used, then sound quality can be improved, but synthesis efficiency deteriorates due to long synthesis duration
Solution Approach 1:
The system performs automatic audio synthesis by extracting human voice and accompaniment from reference audio, automatically determining loudness ranges, comparing and adjusting audio parameters, and mixing the final output without requiring professional audio engineers to manually intervene in each synthesis process
Solution Approach 2:
The system automatically determines loudness ranges for both human voice and accompaniment, compares these ranges, and adjusts the audio parameters (particularly loudness) to ensure proper mixing ratios and sound quality without manual intervention
2Manufacturing precision
If manual audio synthesis process is used, then audio quality can be controlled, but time consumption increases leading to low efficiency
Solution Approach 1:
The patent replaces the manual mechanical process of audio engineering with an automated computer-based system that uses signal processing algorithms to extract audio components, analyze loudness characteristics, and perform mixing operations automatically
Solution Approach 2:
The system determines loudness ranges for both human voice and accompaniment, compares these ranges to identify discrepancies, and uses this comparison feedback to automatically adjust the mixing parameters and achieve balanced audio output
3Manufacturing precision
If professional audio engineering systems are used, then synthesis quality can be maintained, but operation complexity increases due to tight professional restrictions
Solution Approach 1:
The system automatically performs all audio synthesis operations including voice extraction, accompaniment separation, loudness analysis, parameter comparison, and mixing without requiring professional audio engineers to manually control each parameter or make artistic decisions
Solution Approach 2:
The system provides a universal audio synthesis solution that can process various types of reference audio (songs, speeches, etc.) and automatically adapt to different audio characteristics through automated loudness range determination and parameter adjustment, making it accessible to non-professionals
Data Source
AI summary
The present disclosure relates to the technical field of music engineering, and provides an audio synthesis method and apparatus, an electronic device, a medium, and a program product. The present disclosure provides an audio synthesis method, including: acquiring a first human voice audio and an accompaniment audio from reference audio; determining a first loudness range based on loudness of the first human voice audio; acquiring second human voice audio corresponding to the reference audio; determining a second loudness range based on loudness of the second human voice audio; adjusting the loudness of the second human voice audio based on a first comparison result between the second loudness range and the first loudness range to obtain first target human voice audio; and mixing the first target human voice audio and the accompaniment audio to obtain target audio.


