Multi-modal Music Generation via Embedding Vector Stem Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing online music services are unable to create new music and can only recommend existing music based on user preferences, limiting their ability to generate music that reflects diverse user inputs such as text, images, videos, and reference music information.

Innovation Solution

A multi-modal based automatic music generation method and device that generates input embedding vectors from user-provided information, combines audio stems from a database to create song-level music, and selects reference and similar stems to produce natural and complete music.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing online music services recommend existing music based on user preferences, then music recommendation accuracy is improved, but music creativity and diversity are limited

Engineering Contradiction:
Improvemusic recommendation accuracyVSAvoidmusic creativity and diversity
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system uses embedding vectors to capture the essence of user preferences and reference music, then generates new music that copies these patterns while creating something original. The embedding extraction module creates numerical representations of music features, allowing the system to replicate successful patterns without simply repeating existing songs.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system transforms music parameters into embedding vectors through the embedding extraction module, then uses these vectors to guide the music generation process. By changing from traditional music parameters to vector space representations, the system can explore new combinations and create diverse music while maintaining user preference alignment.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If text information is input from the user and music is recommended based on it, then user preference matching is improved, but the ability to receive and process various types of user commands is limited

Engineering Contradiction:
Improveuser preference matchingVSAvoiduser command processing capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The embedding extraction module serves multiple functions: it processes text information, image information, video information, and music information uniformly by converting them all into embedding vectors. This universal approach allows the system to handle diverse user commands and input modalities through a single integrated mechanism.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The embedding vector acts as an intermediary representation that bridges different input modalities and the music generation process. Whether the user inputs text, images, videos, or existing music, the embedding vector translates all these diverse inputs into a common representation that can guide music generation.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If multi-modal information is processed to generate music, then music diversity and user input reflection are improved, but system complexity increases

Engineering Contradiction:
Improvemusic diversity and user input reflectionVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The embedding extraction module extracts and isolates the essential features from multi-modal inputs (text, image, video, music) by converting them into embedding vectors. This extraction process separates the complex input processing from the music generation process, allowing each component to focus on its specific function while maintaining overall system manageability.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system segments the music generation process into distinct modules: embedding extraction for processing multi-modal inputs, stem search for selecting audio elements, and music generation for assembling the final output. This segmentation allows each module to be optimized independently while working together to achieve music diversity and user input reflection.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250201217A1Multi-modal based automatic music generation method and device
Publication Date: 2025.06.19 NEUTUNE CO LTD
  • US20250201217A1 patent drawing
  • US20250201217A1 patent drawing
  • US20250201217A1 patent drawing

AI summary

A multi-modal based automatic music generation device according to an embodiment of present invention comprises an embedding extraction module configured to generate an input embedding vector from input information received from a user terminal, a stem search module configured to select a reference stem based on the input embedding vector and select a plurality of similar stems based on the input embedding vector and an embedding vector included in the reference stem, a reference section generation module configured to create a reference section by editing and mixing the plurality of similar stems and a music generation module configured to generate music in song units by generating a plurality of audio sections based on the reference section.