Karaoke Audio Processing for Multi-User Singing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Karaoke applications do not allow multiple users to sing together simultaneously, limiting the social Karaoke experience.
Innovation Solution
An audio processing method and system that enables users to sing together by playing and recording audio data during specific lyrics parts, mixing user audio with accompaniment or original audio files, and switching between accompaniment and original audio to create a collaborative singing experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If only one user records audio during the entire song, then the recording process is simple, but the social Karaoke experience of singing together with others is lost
Solution Approach 1:
The lyrics file is divided into multiple parts (first lyrics part, second lyrics part, third lyrics part) with different singers assigned to each part. The audio recording process is segmented to record different users' audio data during different time periods corresponding to different lyrics parts, enabling multiple users to sing together while maintaining manageable processing complexity
Solution Approach 2:
The lyrics file is pre-divided into multiple parts with designated singers for each part before the recording process begins. This preliminary segmentation allows the system to automatically switch between different audio files (accompaniment, original, or user-recorded audio) during playback, creating a collaborative singing experience without requiring complex real-time decision-making
2Adaptability or versatility
If the audio file is played continuously without switching, then the playback is simple, but the collaborative singing experience with different singers cannot be achieved
Solution Approach 1:
The audio playback is segmented into different segments corresponding to different lyrics parts. During the first lyrics part, the system plays accompaniment audio and records user audio. During the second lyrics part, the system plays original audio files. During the third lyrics part, the system plays recorded audio from other users. This segmentation enables collaborative singing while keeping the switching logic manageable and predictable
Solution Approach 2:
The audio playback system dynamically switches between different audio sources (accompaniment audio, original audio, recorded user audio) based on the current lyrics part. This dynamic switching capability allows the system to adapt to different singing scenarios and create a collaborative experience without requiring permanent complex configurations
3Adaptability or versatility
If all lyrics parts are recorded by the same user, then the recording process is straightforward, but the experience of singing with others or a star is lost
Solution Approach 1:
The recording process is segmented by lyrics parts, with each part assigned to a different singer. The system automatically manages which user records audio during which lyrics part, eliminating the need for manual configuration while enabling multiple users to participate. This segmentation maintains operational simplicity while achieving multi-user collaboration
Solution Approach 2:
The system acts as an intermediary that coordinates between multiple users and the audio playback/recording process. It automatically switches between different audio sources and manages the recording of different users' audio data during different lyrics parts, creating a seamless collaborative experience without requiring users to manually manage the complexity
Data Source
AI summary
An audio processing method, apparatus and system, capable of realizing the experience of singing Karaoke with other people. The method comprises: acquiring an audio file of a song and a lyric file of the song; playing the audio file at display time corresponding to a first lyric part of the lyric file and recording audio data of a user; playing the audio file at display time corresponding to a second lyric part of the lyric file; and performing audio mixing on the audio data of the user and audio data of the audio file at the display time corresponding to the first lyric part.


