Multi-modal Music Generation via Embedding Vector Stem Search
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing online music services are unable to create new music and can only recommend existing music based on user preferences, limiting their ability to generate music that reflects diverse user inputs such as text, images, videos, and reference music information.
Innovation Solution
A multi-modal based automatic music generation method and device that generates input embedding vectors from user-provided information, combines audio stems from a database to create song-level music, and selects reference and similar stems to produce natural and complete music.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing online music services recommend existing music based on user preferences, then music recommendation accuracy is improved, but music creativity and diversity are limited
Solution Approach 1:
The system uses embedding vectors to capture the essence of user preferences and reference music, then generates new music that copies these patterns while creating something original. The embedding extraction module creates numerical representations of music features, allowing the system to replicate successful patterns without simply repeating existing songs.
Solution Approach 2:
The system transforms music parameters into embedding vectors through the embedding extraction module, then uses these vectors to guide the music generation process. By changing from traditional music parameters to vector space representations, the system can explore new combinations and create diverse music while maintaining user preference alignment.
2Measurement precision
If text information is input from the user and music is recommended based on it, then user preference matching is improved, but the ability to receive and process various types of user commands is limited
Solution Approach 1:
The embedding extraction module serves multiple functions: it processes text information, image information, video information, and music information uniformly by converting them all into embedding vectors. This universal approach allows the system to handle diverse user commands and input modalities through a single integrated mechanism.
Solution Approach 2:
The embedding vector acts as an intermediary representation that bridges different input modalities and the music generation process. Whether the user inputs text, images, videos, or existing music, the embedding vector translates all these diverse inputs into a common representation that can guide music generation.
3Adaptability or versatility
If multi-modal information is processed to generate music, then music diversity and user input reflection are improved, but system complexity increases
Solution Approach 1:
The embedding extraction module extracts and isolates the essential features from multi-modal inputs (text, image, video, music) by converting them into embedding vectors. This extraction process separates the complex input processing from the music generation process, allowing each component to focus on its specific function while maintaining overall system manageability.
Solution Approach 2:
The system segments the music generation process into distinct modules: embedding extraction for processing multi-modal inputs, stem search for selecting audio elements, and music generation for assembling the final output. This segmentation allows each module to be optimized independently while working together to achieve music diversity and user input reflection.
Data Source
AI summary
A multi-modal based automatic music generation device according to an embodiment of present invention comprises an embedding extraction module configured to generate an input embedding vector from input information received from a user terminal, a stem search module configured to select a reference stem based on the input embedding vector and select a plurality of similar stems based on the input embedding vector and an embedding vector included in the reference stem, a reference section generation module configured to create a reference section by editing and mixing the plurality of similar stems and a music generation module configured to generate music in song units by generating a plurality of audio sections based on the reference section.


