Voice Over Music Recording Audio Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic devices lack the capability to efficiently record a user's voice over music without requiring specialized recording equipment, and they struggle to separate the user's voice from the music in real-time.
Innovation Solution
The system architecture enables electronic devices to record a user's voice over music by using one or more microphones to capture both the music output from a speaker and the user's singing, then separates the user's voice from the music using digital signal processing or machine learning models, allowing for recombination with digital audio content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If electronic devices use standard microphones and speakers to record voice over music, then the ease of operation and accessibility are improved, but the ability to separate voice from music and achieve high-quality recordings deteriorates
Solution Approach 1:
The patent replaces traditional mechanical audio recording systems with digital signal processing and machine learning models. The system uses software-based voice separation algorithms (such as neural networks) to distinguish and isolate the user's voice from the music, substituting complex mechanical separation equipment with intelligent digital processing that can accurately separate audio sources using standard device microphones.
2Manufacturing precision
If electronic devices use specialized recording equipment to capture high-quality voice over music, then the recording quality is improved, but the device complexity and cost increase
Solution Approach 1:
The patent creates a virtual copy of the professional recording experience by using machine learning models trained on high-quality recording data. These models learn from extensive audio datasets to replicate the separation and processing capabilities of specialized equipment, allowing standard devices to achieve professional-quality results through software intelligence rather than hardware complexity.
Solution Approach 2:
The system dynamically adjusts audio processing parameters in real-time based on the detected audio environment. The machine learning models analyze audio characteristics and automatically optimize processing parameters such as voice separation thresholds, noise suppression levels, and mixing ratios, eliminating the need for manual configuration and specialized equipment setup.
3Productivity
If electronic devices process and separate voice from music in real-time, then the productivity and user experience are improved, but the energy consumption and computational load increase
Solution Approach 1:
The patent implements adaptive processing that applies full computational power only when voice separation is actively needed, rather than continuously processing all audio data. The system can switch between processing modes based on user interaction, applying intensive machine learning inference only during recording segments and using lighter processing for playback or idle states, thus reducing overall energy consumption while maintaining high productivity when required.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This solution allows users to create high-quality recordings of their singing voice with prerecorded music without needing specialized equipment, enabling real-time playback and processing such as noise suppression and style matching.
Implementation Method 1
obtain, using one or more microphones of an electronic device, an audio input stream comprising: a voice of a person from a first source, and music from a second source that differs from the first source
Data Source
AI summary
Systems, devices, and methods for voice over music recording are provided. A person may sing along with music being output into the physical environment of the person by a loudspeaker. One or more microphones may be used to record both the voice of the person singing, and the music output by the speaker. An electronic device may cancel the portion of the microphone signals containing the music output by the speaker, and remix the voice of the person with digital audio content corresponding to the music.


