Method, device, and system for remixing, editing, and generating music on basis of artificial intelligence
The deep learning-based method addresses the limitations of conventional music editing and remixing by segmenting music structure and instrument elements, enabling natural and precise editing and remixing that reflects user intention and feedback.
Patent Information
- Application Number
- PCT/KR2024/020314
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-12-12
- Filing Date
- 2024-12-13
- Publication Date
- 2025-06-19
AI Technical Summary
Conventional music editing and remixing technologies struggle to accurately reflect the complex structure and instrumental elements of music, leading to limitations in feature extraction and natural result generation.
A deep learning-based method that extracts various music features such as genre, mood, tempo, and instrument composition, allowing for segmentation of music structure and instrument elements, enabling natural editing, remixing, and generation based on user intention.
Enables advanced music analysis and generation by automating manual music element analysis, allowing for flexible and precise remixing and editing, while reflecting user feedback and preferences in real-time.
Smart Images

Figure KR2024020314_19062025_PF_FP_ABST
Abstract
Description
Method, device and system for remixing, editing and generating music based on artificial intelligence
[0001] The present invention relates to music and audio processing technology using artificial intelligence, and more particularly, to a method, device, system and method for remixing, editing and generating music using the structure and instrumental elements of music based on artificial intelligence technology.
[0002] The present invention relates to music and audio processing technology using deep learning, and more particularly, to a system and method for remixing, editing, and generating music according to a user's intention using deep learning technology to determine the structure and instrumental elements of the music.
[0003] The following documents form the background or part of the present invention, and the contents disclosed in the documents are referenced in describing the present invention.
[0004] 1. Kim et al., All-In-One Metrical And Functional Structure Analysis With Neighborhood Attentions on Demixed Audio. IEEE, WASPAA, 2023
[0005] 2. Lee et al., "Metric learning vs classification for disentangled music representation learning", ISMIR, 2020
[0006] 3. Doh, Seung-Heon, et al. "LP-MusicCaps: LLM-Based Pseudo Music Captioning. ISMIR, 2023
[0007] 4. Doh, Seung-Heon, et al. “ENRICHING MUSIC DESCRIPTIONS WITH A FINETUNED-LLM AND METADATA FOR TEXT-TO-MUSIC RETRIEVAL”. ICASSP, 2024
[0008] Conventional music editing and remixing techniques primarily rely on simple track-level processing, limiting their ability to accurately capture the complex structure or detailed instrumental elements of music. Furthermore, existing technologies struggle to precisely extract musical characteristics or deliver natural-looking results during remixing.
[0009] Furthermore, conventional music information retrieval (MIR) techniques have primarily focused on individual tasks, such as beat / downbeat tracking, segmentation, and functional structure labeling. This often misses the potential benefits of inter-structural interactions. Furthermore, integrating all hierarchical information into a single dataset requires processing the length of a song and high-dimensional data, limiting conventional deep learning models' ability to handle such data.
[0010] The automatic multi-channel music mix from audio stems (or component music files) disclosed in [Patent Document 0001] automatically mixes audio stems (or component music files) based on preset rules, but has limited functionality for reflecting the user's intention or preference in detail.
[0011] [Patent Document 0002] focuses on generating music by having a machine learning model learn the music generation method of a rule-based expert system. However, while incorporating user feedback and preferences into model training, the ability to directly reflect user intent during the music generation process itself is relatively limited.
[0012] [Prior Art Literature]
[0013] [Patent Document 0001] Korean Patent No. 10-2268933 (Published on June 25, 2021) - 'Automatic Multi-Channel Music Mix from Multiple Audio Stems'
[0014] [Patent Document 0002] International Publication No. WO 2024 / 178038 A1 (Published on February 21, 2024) - 'Machine learning model learned based on music and decisions generated by an expert system'
[0015] The present invention is designed to solve the above-described problems, and more specifically, to provide a method, device and system for remixing, editing and generating music using the structure and instrumental elements of music based on artificial intelligence technology.
[0016] According to the present invention, a method, device and system for generating an embedding based on at least one of text information, image information, music information and video information, and remixing, editing and generating music based thereon can be provided.
[0017] The present invention provides advanced music analysis and generation technology based on deep learning. The present invention utilizes deep learning to extract various characteristics of music, such as genre, mood, tempo, and instrumental composition, thereby segmenting the structure and instrumental elements of the music, enabling the user to naturally edit, remix, and generate music as intended. By utilizing deep learning, the present invention can automate music elements previously analyzed manually and provide advanced analysis.
[0018] The present invention provides a technique for segmenting and componentizing various musical elements. The present invention provides a technique for segmenting music along time axis (e.g., intro section, verse section, chorus section, etc.) and instrument axis (e.g., vocals, bass, drums, etc.) and then dividing the music into component blocks. Through this structural analysis, users can easily edit and reconstruct various parts of the music. This allows for natural changes in the music and enables more flexible and precise remixing.
[0019] The present invention provides remixing and generation technology based on user intent. The present invention provides deep learning-based music captioning and embedding technology that enables music remixing and generation based on user input prompts or intent. According to the present invention, more flexible and customized music editing and generation is possible by analyzing and reflecting musical elements tailored to the user's needs.
[0020] The present invention provides technology capable of incorporating user feedback in real time during music creation. The present invention utilizes a deep learning-based model to generate music by incorporating user feedback in real time. User preferences and feedback are analyzed through embedding, and this information is then reflected in the music creation and editing process, providing users with optimized music.
[0021] The present invention provides a music feature analysis technique using multidimensional embedding and contrastive learning. This technique utilizes deep learning to generate multidimensional embeddings of music content and analyzes feature differences between music tracks through contrastive learning. This allows for more precise analysis of various aspects of music, thereby enhancing the quality of music creation and editing.
[0022] The present invention provides a technology that utilizes deep learning to generate embeddings between music and text. This allows users to describe the characteristics of music in natural language, search for appropriate music, and remix it. The music text embedding model effectively reflects the user's desired musical style or characteristics, providing a more intuitive remixing and generation process.
[0023] The present invention can achieve musical harmony through frequency distribution and chord progression embedding. The present invention analyzes and reflects the detailed range and progression of music through frequency distribution and chord progression embedding. This allows the frequency distributions of various component music files to be adjusted to avoid overlapping during remixing, thereby maintaining musical harmony and effectively enabling the search for music with similar chord progressions. The present invention analyzes and reflects the detailed range and progression of music through frequency band embedding and chord progression embedding. This allows the frequency distributions of various component music files to be adjusted to avoid overlapping during remixing, and music with similar chord progressions can be found. This enhances the completeness of the remix.
[0024] A method for generating music based on artificial intelligence according to one embodiment of the present invention may include the steps of: receiving at least one of text information, image information, music information, and video information; extracting an embedding based on the received information; calculating a similarity between the extracted input embedding and an embedding of a component block included in a database; selecting a reference component block from the database based on the similarity; filtering the database based on the selected reference component block; selecting one or more component blocks from the filtered database; generating a reference section based on the reference component and the selected one or more component blocks; and generating one or more audio sections based on the generated reference section, and listing the reference section and the one or more audio sections in time series.
[0025] In a method for generating music based on artificial intelligence according to one embodiment of the present invention, the step of extracting an embedding based on the received information may include at least one of a step of generating a text embedding based on at least one of caption information and text information generated for at least one of the image information and the music information, a step of generating a music feature embedding from the music information, a step of generating a chord progression embedding from the music information, and a step of generating a frequency band embedding from the music information.
[0026] A device for generating music based on artificial intelligence according to one embodiment of the present invention may include means for receiving at least one of text information, image information, music information, and video information, means for extracting an embedding based on the received information, means for calculating a similarity between the extracted input embedding and an embedding of a component block included in a database, means for selecting a reference component block from the database based on the similarity, means for filtering the database based on the selected reference component block, means for selecting one or more component blocks from the filtered database, means for generating a reference section based on the reference component and the one or more selected component blocks, and means for generating one or more audio sections based on the generated reference section, and listing the reference section and the one or more audio sections in time series.
[0027] In a device for generating music based on artificial intelligence according to one embodiment of the present invention, the means for extracting an embedding based on the received information may perform at least one of an operation of generating a text embedding based on at least one of caption information and text information generated for at least one of the image information and the music information, an operation of generating a music feature embedding from the music information, an operation of generating a chord progression embedding from the music information, and an operation of generating a frequency band embedding from the music information.
[0028] The present invention relates to a method for providing music generated based on artificial intelligence, comprising: receiving at least one of text information, tag information, image information, and music information input from a user terminal; extracting one or more embeddings based on the received information through one or more embedding extraction models; calculating a similarity between one or more embeddings based on the received information and one or more embeddings extracted from component blocks included in a database, the component blocks being extracted from original music files, through a similarity calculation model; selecting a first reference component block from the database based on the calculated similarity; filtering the database based on the selected first reference component block; selecting one or more component blocks from the filtered database; and generating a reference section based on the first reference component and the selected one or more component blocks.
[0029] In addition, the method includes: selecting a second reference component block from the raw music file, the second reference component block having the same instrumental elements as the first reference component block; selecting one or more additional component blocks from the filtered database based on the second reference component block; generating an audio section based on the second reference component block and the one or more additional component blocks; and generating a music file by arranging the reference section and the audio section in time series.
[0030] Additionally, one or more embeddings extracted based on the received information and one or more embeddings extracted from component blocks included in the database each include at least one of a text embedding, a music text embedding, a music feature embedding, a chord progression embedding, and a frequency band embedding.
[0031] In addition, the step of calculating the similarity between one or more embeddings extracted based on the received information and one or more embeddings extracted from component blocks included in the database through a similarity calculation model further includes the step of calculating one or more similarity scores based on the text embedding, the music text embedding, the music feature embedding, the chord progression embedding, and the frequency band embedding; and the step of reflecting a weight according to the type of the embedding for the one or more similarity scores and calculating the similarity by adding up the one or more similarity scores to which the weight is reflected.
[0032] In addition, the step of filtering the database based on the selected first reference component block further includes the step of filtering the database based on music basic information related to the first reference component block.
[0033] The present invention relates to a server for providing music generated based on artificial intelligence, comprising: a communication module; a memory for storing a database; and a processor; wherein the processor receives at least one of text information, tag information, image information, and music information input from a user terminal through the communication module, extracts one or more embeddings based on the received information through one or more embedding extraction models stored in the memory, calculates a similarity between one or more embeddings based on the received information and one or more embeddings extracted from component blocks included in a database, the component blocks being extracted from original music files, through a similarity calculation model stored in the memory, selects a first reference component block from the database based on the calculated similarity, filters the database based on the selected first reference component block, selects one or more component blocks from the filtered database, and generates a reference section based on the first reference component and the selected one or more component blocks.
[0034] Additionally, the processor selects, from the raw music file, a second reference component block having the same instrumental elements as the first reference component block, selects one or more additional component blocks from the filtered database based on the second reference component block, generates an audio section based on the second reference component block and the one or more additional component blocks, and arranges the reference section and the audio section in time series to generate a music file.
[0035] Additionally, one or more embeddings extracted based on the received information and one or more embeddings extracted from component blocks included in the database each include at least one of a text embedding, a music text embedding, a music feature embedding, a chord progression embedding, and a frequency band embedding.
[0036] In addition, the processor calculates one or more similarity scores based on the text embedding, the music text embedding, the music feature embedding, the chord progression embedding, and the frequency band embedding through a similarity calculation model stored in the memory, reflects a weight according to the type of the embedding for the one or more similarity scores, and calculates the similarity by adding up the one or more similarity scores to which the weight is reflected.
[0037] Additionally, the processor filters the database based on music basic information related to the first reference component block.
[0038] The present invention relates to a method for providing music generated based on artificial intelligence, comprising: receiving text information, tag information, image information, and music information input from a user terminal; generating image caption information for the image information through an image captioning model; generating music caption information for the music information through a music captioning model; performing preprocessing on the text information, the tag information, the image caption information, and the music caption information through a text preprocessing model; and generating at least one of a text embedding and a music text embedding based on the preprocessed information through one or more embedding extraction models.
[0039] In addition, the music captioning model is an artificial intelligence model that is pre-trained on a mock caption data set generated based on an artificial intelligence model utilizing a system prompt including one or more tags and one or more instructions, and is an artificial intelligence model that is transferred based on a label data set.
[0040] Additionally, the one or more directives include at least one of a directive for specifying a purpose or role of the music captioning model, and a directive for specifying an output format of a simulated caption.
[0041] In addition, the method further includes a step of calculating a similarity between at least one of a text embedding and a music text embedding based on the preprocessed information and at least one of a text embedding and a music text embedding extracted from a component block included in a database, using a similarity calculation model.
[0042] In addition, the step of calculating the similarity further includes the step of calculating one or more similarity scores between at least one of a text embedding and a music text embedding based on the preprocessed information and at least one of a text embedding and a music text embedding extracted from a component block included in a database, reflecting a weight according to the type of the embedding for the one or more similarity scores, and calculating the similarity by adding the one or more similarity scores to which the weight is reflected.
[0043] The present invention relates to a server for providing music generated based on artificial intelligence, comprising: a communication module; a memory for storing a database; and a processor; wherein the processor receives text information, tag information, image information, and music information input from a user terminal through the communication module, generates image caption information for the image information through an image captioning model stored in the memory, generates music captioning information for the music information through a music captioning model stored in the memory, performs preprocessing on the text information, the tag information, the image caption information, and the music caption information through a text preprocessing model stored in the memory, and generates at least one of a text embedding and a music text embedding based on the preprocessed information through one or more embedding extraction models stored in the memory.
[0044] In addition, the music captioning model is an artificial intelligence model that is pre-trained on a mock caption data set generated based on an artificial intelligence model utilizing a system prompt including one or more tags and one or more instructions, and is an artificial intelligence model that is transferred based on a label data set.
[0045] In addition, the music captioning model is a natural language processing-based artificial intelligence model that is pre-trained on a mock caption data set generated based on a natural language processing-based artificial intelligence model utilizing a system prompt including one or more tags and one or more instructions, and is a natural language processing-based artificial intelligence model that is transferred based on a label data set.
[0046] Additionally, the one or more directives include at least one of a directive for specifying a purpose or role of the music captioning model, and a directive for specifying an output format of a simulated caption.
[0047] Additionally, the processor calculates, through a similarity calculation model stored in the memory, a similarity between at least one of a text embedding and a music text embedding based on the preprocessed information and at least one of a text embedding and a music text embedding extracted from a component block included in a database.
[0048] The present invention was conceived to address the aforementioned problems of the prior art. It utilizes deep learning to extract various musical characteristics, such as genre, mood, tempo, and instrumental composition, thereby segmenting the musical structure and instrumental elements, enabling the user to naturally edit, remix, and generate music as intended. By leveraging deep learning, the present invention automates musical elements previously analyzed manually, providing advanced analysis.
[0049] According to the present invention, a technique can be provided for dividing music into component blocks by dividing it into time axis (e.g., intro section, verse section, chorus section, etc.) and instrument axis (e.g., vocals, bass, drums, etc.). This structural analysis allows users to easily edit and reconstruct various parts of the music. This allows for natural changes in the music and enables more flexible and precise remixing.
[0050] The present invention provides a deep learning-based music captioning and embedding technology that enables music remixing and creation based on user input prompts or intentions. The present invention enables more flexible and customized music editing and creation by analyzing and reflecting musical elements tailored to the user's needs.
[0051] According to the present invention, a deep learning-based model can generate music by reflecting user feedback in real time. User preferences and feedback are analyzed through embedding, and this information is then reflected in the music creation and editing process, providing users with optimized music.
[0052] According to the present invention, deep learning can be used to generate multidimensional embeddings of music content and contrastive learning can be used to analyze feature differences between music tracks. This allows for more precise analysis of various aspects of music, which can then be used to improve the quality of music creation and editing.
[0053] According to the present invention, by providing a technology that generates embeddings between music and text using deep learning, users can describe the characteristics of music in natural language and search for and remix music that fits their needs. The music text embedding model effectively reflects the user's desired musical style or characteristics, providing a more intuitive remixing and generation process.
[0054] According to the present invention, the detailed range and progression of music can be analyzed and reflected through frequency distribution and chord progression embedding. This allows for maintaining musical harmony by adjusting the frequency distributions of various component music files to avoid overlapping during remixing, and effectively searching for music with similar chord progressions. The present invention analyzes and reflects the detailed range and progression of music through frequency band embedding and chord progression embedding. This allows for adjusting the frequency distributions of various component music files to avoid overlapping during remixing, and for searching for music with similar chord progressions. This enhances the completeness of the remix.
[0055] According to the present invention, a method, device and system for remixing, editing and generating music on a multi-modal basis based on at least one of text, image, video and music information can be provided.
[0056] The effects according to the present invention are not limited thereto, and a person having ordinary skill in the art of the present invention will be able to recognize the effects of the present invention devised to solve the problems of the prior art described above based on the configuration and purpose described in the specification of the present invention.
[0057] Figures 1a and 1b are conceptual diagrams illustrating a system environment for providing a service according to one embodiment of the present invention.
[0058] FIG. 2a is a drawing for explaining the configuration of a user terminal (100a) and a local terminal (100b) according to one embodiment of the present invention.
[0059] FIG. 2b is a drawing for explaining the configuration of a cloud server (200a) and a local server (200b) according to one embodiment of the present invention.
[0060] FIG. 3a is a flowchart illustrating an artificial intelligence model for providing a service according to one embodiment of the present invention and an overall method for providing a service according to one embodiment of the present invention.
[0061] FIG. 3b is a diagram illustrating an artificial intelligence model for obtaining information and / or files according to one embodiment of the present invention.
[0062] FIG. 4a is a table that organizes one or more files for providing a service according to one embodiment of the present invention.
[0063] FIG. 4b is a table that organizes one or more pieces of information for providing a service according to one embodiment of the present invention.
[0064] FIG. 4c is a table that organizes one or more embeddings for providing a service according to one embodiment of the present invention.
[0065] FIG. 5a is a conceptual diagram illustrating a process of obtaining a component music file from a raw music file according to one embodiment of the present invention.
[0066] FIG. 5b is a conceptual diagram illustrating a database of component music files according to one embodiment of the present invention.
[0067] Figure 6 is a conceptual diagram for explaining the configuration of a hierarchical structure analysis model according to one embodiment of the present invention.
[0068] FIG. 7 is a conceptual diagram illustrating the configuration of a transformer-based artificial intelligence model included in a hierarchical structure analysis model according to one embodiment of the present invention.
[0069] FIG. 8 is a conceptual diagram for explaining a neighboring attention model included in a hierarchical structure analysis model according to one embodiment of the present invention.
[0070] FIG. 9 is a diagram for explaining a process of generating a mock caption according to one embodiment of the present invention.
[0071] FIG. 10 is a drawing for explaining an example of a directive input to generate a caption according to one embodiment of the present invention.
[0072] FIG. 11 is a drawing for explaining an example of a directive input to generate a caption according to one embodiment of the present invention.
[0073] FIG. 12 is a conceptual diagram illustrating the configuration of a cross-modal encoder-decoder transformer according to one embodiment of the present invention.
[0074] FIG. 13 is a diagram for explaining a process of obtaining a music captioning model (302) according to one embodiment of the present invention.
[0075] FIG. 14 is a diagram for explaining a music text embedding extraction model according to one embodiment of the present invention.
[0076] FIG. 15a is a diagram for explaining a triplet-based music feature embedding extraction model according to one embodiment of the present invention.
[0077] FIG. 15b is a diagram for explaining a proxy-based music feature embedding extraction model according to one embodiment of the present invention.
[0078] FIG. 15c is a diagram for explaining a classification-based music feature embedding extraction model according to one embodiment of the present invention.
[0079] FIG. 16a and FIG. 16b are diagrams for explaining a code progression embedding extraction model according to one embodiment of the present invention.
[0080] FIG. 17a and FIG. 17b are diagrams for explaining a frequency band embedding extraction model according to one embodiment of the present invention.
[0081] FIG. 18 is a diagram for explaining a similarity determination method based on a music feature embedding space according to one embodiment of the present invention.
[0082] FIG. 19a is a flowchart illustrating a component block filtering process according to one embodiment of the present invention.
[0083] Figure 19b is a flowchart illustrating a component block search process according to one embodiment of the present invention.
[0084] Figure 20 is a flowchart illustrating a section expansion process according to one embodiment of the present invention.
[0085] FIG. 21 is a flowchart illustrating a component block editing and mixing process according to one embodiment of the present invention.
[0086] FIG. 22a and FIG. 22b are flowcharts illustrating a method for generating music on a multi-modal basis according to one embodiment of the present invention.
[0087] Hereinafter, embodiments of the present invention will be described with reference to the attached drawings. When adding reference numerals to components in each drawing, it should be noted that identical components are given the same numerals as much as possible even if they are shown in different drawings. In addition, when describing embodiments of the present invention, if a detailed description of a related known configuration or function is judged to hinder understanding of the embodiments of the present invention, a detailed description thereof will be omitted. In addition, although embodiments of the present invention will be described below, the technical idea of the present invention is not limited or restricted thereto, and may be modified and implemented in various ways by those skilled in the art.
[0088] The advantages and features of the present invention, and the methods for achieving them, will become clearer with reference to the embodiments described in detail below together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided solely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the present invention, and the present invention is defined solely by the scope of the claims.
[0089] The terminology used herein is for the purpose of describing embodiments only and is not intended to limit the present invention. In this specification, the singular also includes the plural unless specifically stated otherwise. As used herein, the terms "comprises" and / or "comprising" do not exclude the presence or addition of one or more other components in addition to the mentioned components. Like reference numerals refer to like components throughout the specification, and "and / or" includes each and any combination of one or more of the mentioned components. Although "first", "second", etc. are used to describe various components, these components are not limited by these terms. These terms are only used to distinguish one component from another. Therefore, it should be understood that a first component mentioned below may also be a second component within the technical spirit of the present invention.
[0090] Unless otherwise defined, all terms (including technical and scientific terms) used herein may be used in their common sense to those skilled in the art to which the present invention pertains. Furthermore, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise.
[0091] The term "part" or "module" as used herein refers to a software or hardware component such as an FPGA or ASIC, and the "part" or "module" performs certain functions. However, the "part" or "module" is not limited to software or hardware. The "part" or "module" may be configured to reside on an addressable storage medium and may be configured to execute one or more processors. Thus, by way of example, the "part" or "module" includes components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables. The functionality provided within the components and "parts" or "modules" may be combined into fewer components and "parts" or "modules," or further separated into additional components and "parts" or "modules."
[0092] In this specification, the term "computer" refers to any type of hardware device including at least one processor, and may also be understood to encompass software components operating on the hardware device, depending on the embodiment. For example, the term "computer" may be understood to encompass, but is not limited to, smartphones, tablet PCs, desktops, laptops, and all user clients and applications running on each device.
[0093] Those skilled in the art should further appreciate that the various illustrative logical blocks, configurations, modules, circuits, means, logics, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, configurations, means, logics, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application. However, such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0094] Overall system environment
[0095] Figures 1a and 1b are conceptual diagrams illustrating a system environment for providing a service according to one embodiment of the present invention.
[0096] Referring to FIG. 1a, a system environment for providing a service according to one embodiment of the present invention may include a user terminal (100a), a cloud server (200a), and a local server (200b). Referring to FIG. 1b, a system environment for providing a service according to one embodiment of the present invention may further include a local terminal (100b) connected to a local server (200b) via one or more networks according to one embodiment of the present invention.
[0097] An artificial intelligence model for providing music remixing, editing, and / or creation services according to one embodiment of the present invention may be implemented in a cloud server (200a) and / or a local server (200b). According to one embodiment of the present invention, the artificial intelligence model may be implemented in a distributed manner across the cloud server (200a) or the local server (200b). The artificial intelligence model according to one embodiment of the present invention may include one or more artificial neural networks.
[0098] According to one embodiment of the present invention, an application program may be stored in a user terminal (100a) and / or a local terminal (100b). Based on user input (e.g., prompt input, etc.) entered into the user terminal (100b) and / or the local terminal (100b), a music remixing, editing, and / or creation service according to one embodiment of the present invention may be provided by utilizing one or more artificial intelligence models on a cloud server (200a) and / or a local server (200b).
[0099] User terminal (100a) and local terminal (100b)
[0100] FIG. 2A is a diagram for explaining the configuration of a user terminal (100a) and a local terminal (100b) according to one embodiment of the present invention. Referring to FIG. 2A, the user terminal (100a) may include at least one of a processor (110a), a memory (120a), an output unit (130a), an input unit (140a), an interface (150a), a service module (160a), and a communication module (170a). Referring to FIG. 2A, the local terminal (100b) may include at least one of a processor (110b), a memory (120b), an output unit (130b), an input unit (140b), an interface (150b), a service module (160b), and a communication module (170b). The configuration of the user terminal (100a) and the local terminal (100b) of the present invention is not limited to the components illustrated in FIG. 2A. That is, depending on the implementation aspect of the user terminal (100a) and / or local terminal (100b) according to the present invention, additional components may be included or some of the components illustrated in FIG. 2a may be omitted.
[0101] According to one embodiment of the present invention, the user terminal (100a) and / or the local terminal (100b) may refer to a terminal possessed by a user, which may provide a music remixing, editing and / or creation service according to the present invention through information exchange with a cloud server (200a) and / or a local server (200b). The user terminal (100a) and / or the local terminal (100b) according to one embodiment of the present invention may refer to any type of entity(ies) in a system having a mechanism for communication with other electronic devices disclosed in the present invention. For example, such user terminal (100a) and / or the local terminal (100b) may include a personal computer (PC), a notebook, a mobile terminal, a smart phone, a tablet PC, an artificial intelligence (AI) speaker, an artificial intelligence TV, a wearable device, etc., and may include all types of terminals capable of connecting to a wired / wireless network. Additionally, the user terminal (100a) and / or local terminal (100b) may include any server implemented by at least one of an agent, an Application Programming Interface (API), and a plug-in. Additionally, the user terminal (100a) and / or local terminal (100b) may include an application source and / or a client application.
[0102] According to one embodiment of the present invention, an application program may be installed on a user terminal (100a) and / or a local terminal (100b). Various prompts may be input through the application program installed on the user terminal (100a) and / or the local terminal (100b), and various related information, including music, may be generated from the input various prompts using an artificial intelligence model on a cloud server (200a) and / or a local server (200b).
[0103] Processor (110a, 110b)
[0104] A processor (110a, 110b) according to one embodiment of the present invention may include one or more application processors (APs), one or more communication processors (CPs), or at least one artificial intelligence processor (AI processor). The application processors, communication processors, or AI processors may be contained within different integrated circuit (IC) packages, or may be contained within a single IC package.
[0105] An application processor can control multiple hardware or software components connected to the application processor by running an operating system or application program, and perform various data processing / operations, including multimedia data. For example, the application processor may be implemented as a system on chip (SoC). The processors (110a, 110b) may further include a graphics processing unit (GPU, not shown).
[0106] The communication processor may perform functions such as managing data links and converting communication protocols in communication between the user terminal (100a) and / or the local terminal (100b) and other computing devices connected to the network (e.g., a cloud server (200a), a local server (200b), etc.). For example, the communication processor may be implemented as an SoC. The communication processor may perform at least a portion of the multimedia control functions. In addition, the communication processor may control data transmission and reception of the communication modules (170a, 170b). The communication processor may also be implemented to be included as at least a part of the application processor.
[0107] An application processor or a communication processor may load commands or data received from non-volatile memory or at least one of the other components connected thereto into volatile memory for processing. Furthermore, the application processor or the communication processor may store data received from or generated by at least one of the other components in non-volatile memory. The application processor, in conjunction with the communication processor, may control service provision operations through service modules (160a, 160b) through communication with other computing devices connected to a network. The application processor may also be implemented to be included as at least a part of the communication processor.
[0108] When loaded into the memory (120a, 120b), the computer program may include one or more instructions that cause the processor (110a, 110b) to perform a method / operation according to various embodiments of the present invention. That is, the processor (110a, 110b) may perform the method / operation according to various embodiments of the present invention by executing the one or more instructions. In one embodiment, the computer program may include one or more instructions that cause the processor to perform a remixing, editing, and / or generating operation of music according to one embodiment of the present invention.
[0109] According to one embodiment of the present invention, the processor (110a, 110b) may be configured with one or more cores and may include a processor for data analysis and deep learning, such as a central processing unit (CPU), a general purpose graphics processing unit (GPGPU), and a tensor processing unit (TPU).
[0110] The processors (110a, 110b) can read a computer program stored in the memory (120a, 120b) and perform data processing for machine learning according to an embodiment of the present invention. According to an embodiment of the present invention, the processors (110a, 110b) can perform operations for learning a neural network. The processors (110a, 110b) can perform calculations for learning a neural network, such as processing input data for learning in deep learning (DL), extracting features from input data, calculating errors, and updating weights of the neural network using backpropagation.
[0111] Additionally, at least one of the CPU, GPGPU, and TPU of the processors (110a, 110b) can process neural network training. For example, the CPU and GPGPU can jointly process neural network training and data classification using the neural network. Additionally, in one embodiment of the present invention, processors of multiple computing devices can be used together to process neural network training and data classification using the neural network. Additionally, a computer program executed in a computing device according to one embodiment of the present invention can be a CPU, GPGPU, or TPU executable program.
[0112] In this specification, a model may include a neural network. In this specification, the term "neural network" may be used interchangeably with "artificial neural network" and "neural network." In this specification, a neural network may include one or more neural networks, in which case the output of the model and / or the neural network may be an ensemble of the outputs of one or more neural networks.
[0113] According to one embodiment of the present invention, the processor (110a, 110b) can typically process the overall operation of the user terminal (100a) and / or the local terminal (100b). The processor (110a, 110b) can process signals, data, information, etc. input or output through the components described above, or can run application programs stored in the memory (120a, 120b) to provide or process appropriate information or functions to the user terminal.
[0114] Memory (120a, 120b)
[0115] According to one embodiment of the present invention, the memory (120a, 120b) may include built-in memory or external memory. The built-in memory may include at least one of volatile memory (e.g., dynamic RAM (DRAM), static RAM (SRAM), synchronous dynamic RAM (SDRAM), etc.) or non-volatile memory (e.g., one time programmable ROM (OTPROM), programmable ROM (PROM), erasable and programmable ROM (EPROM), electrically erasable and programmable ROM (EEPROM), mask ROM, flash ROM, NAND flash memory, NOR flash memory, etc.). According to one embodiment, the built-in memory may take the form of a solid state drive (SSD). The external memory may further include a flash drive, for example, a compact flash (CF), a secure digital (SD), a micro secure digital (Micro-SD), a mini secure digital (Mini-SD), an extreme digital (xD), or a memory stick.
[0116] According to one embodiment of the present invention, the memory (120a, 120b) can store a computer program for performing a method of remixing, editing, and / or generating music according to one embodiment of the present invention, and the stored computer program can be read and executed by the processor (110a, 110b). In addition, the memory (120a, 120b) can store any type of information generated or determined by the processor (110a, 110b) and any type of information received through the communication module (170a, 170b). In addition, the memory (120a, 120b) can store data regarding a user prompt. In addition, the memory (120a, 120b) can store data regarding remixing, editing, and / or generating music. For example, the memory (120a, 120b) can temporarily or permanently store input / output data.
[0117] According to one embodiment of the present invention, the memory (120a, 120b) may include at least one type of storage medium among a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, and an optical disk. The user terminal (100a) and / or the local terminal (100b) may operate in relation to a web storage that performs the storage function of the memory (120a, 120b) on the Internet. The description of the above-described memory is merely an example, and the present invention is not limited thereto.
[0118] Output section (130a, 130b)
[0119] An output unit (130a, 130b) according to one embodiment of the present invention is for generating output related to vision, hearing, or tactile sensations, and may include at least one of a display unit, an audio output unit, a haptic module, and an optical output unit.
[0120] According to one embodiment of the present invention, a display unit may be formed as a touch screen by forming a mutual layer structure with a touch sensor or by forming an integral structure. Such a touch screen may function as a user input unit that provides an input interface between a user terminal (100a) and / or a local terminal (100b) and a user, and at the same time, may provide an output interface between the user terminal (100a) and / or a local terminal (100b) and the user.
[0121] An audio output unit according to one embodiment of the present invention includes a device such as a speaker, through which audio signals can be output. According to one embodiment of the present invention, remixed, edited, and / or generated music can be output through the audio output unit.
[0122] Input section (140a, 140b)
[0123] An input unit (140a, 140b) according to one embodiment of the present invention may include at least one of a camera or an image input unit for inputting an image signal, a microphone or an audio input unit for inputting an audio signal, and a user input unit (e.g., an input interface, etc.) for receiving information from a user. Voice data or image data collected by the input unit (140a, 140b) may be analyzed and processed into a user control command. A user may input one or more user prompts (e.g., contextual information through text, image information, video information, music information, etc.) related to the service of the present invention through the user input unit according to one embodiment of the present invention.
[0124] Interface (150a, 150b)
[0125] An interface (150a, 150b) according to one embodiment of the present invention serves as a passageway for various types of external devices connected to a user terminal (100a) and / or a local terminal (100b). The interface (150a, 150b) may include at least one of a wired / wireless headset port, an external charger port, a wired / wireless data port, a memory card port, a port for connecting a device equipped with an identification module, an audio I / O (Input / Output) port, a video I / O (Input / Output) port, and an earphone port. In the user terminal (100a) and / or the local terminal (100b), appropriate control related to the connected external device can be performed in response to the external device (e.g., a cloud server (200a) and / or a local server (200b)) being connected to the interface (150a, 150b).
[0126] In addition, the interface (150a, 150b) according to one embodiment of the present invention can perform a data transfer operation between at least one of a processor (110a, 110b), a memory (120a, 120b), an output unit (130a, 130b), an input unit (140a, 140b), a service module (160a, 160b), and a communication module (170a, 170b). The interface (150a, 150b) according to one embodiment of the present invention can transfer a command or data input from a user through the input unit (140a, 140b) or the output unit (130a, 130b) to the processor (110a, 110b), the memory (120a, 120b), the communication module (170a, 170b), etc. through a bus (not shown). For example, the interface (150a, 150b) may provide data regarding a user's touch input input through a touch panel to the processor (110a, 110b). For example, the interface (150a, 150b) may output commands or data received from a processor (110a, 110b), a memory (120a, 120b), a communication module (170a, 170b), etc. via a bus through an output unit (130a, 130b). For example, the interface (150a, 150b) may output voice data processed by the processor (110a, 110b) to the user through a speaker. Meanwhile, the interface (150a, 150b) according to one embodiment of the present invention may be driven based on a user input through an input unit (140a, 140b).
[0127] Service module (160a, 160b)
[0128] The service modules (160a, 160b) according to one embodiment of the present invention can provide various application program services based on the output results of the artificial intelligence model of the present invention. Specifically, the service modules (160a, 160b) can compare, classify, and analyze data on various music files and specific information stored in the memory of at least one of the user terminal (100a), the local terminal (100b), the cloud server (200a), and the local server (200b) to provide various services that meet the needs of the user. For example, the service modules (160a, 160b) can provide a singer identification service, a similar music search service, a similar music search service based on a specific item, a vocal tagging service, a melody extraction service, a humming query service, etc. based on the output results of the artificial intelligence model according to one embodiment of the present invention.
[0129] According to one embodiment of the present invention, the service modules (160a, 160b) may be implemented in the form of an application program on a server. The application program implemented on the server may be provided externally in the form of an application programming interface (API). Furthermore, the service modules (160a, 160b) may be implemented in the form of an application program on a user terminal (100a) and / or a local terminal (100b). The user may receive the services of the application program through a user interface output through the output unit (130a, 130b) of the user terminal (100a) and / or the local terminal (100b).
[0130] A service according to one embodiment of the present invention may include various services, such as an artist search service for input music, a search service for other music composed or performed / sung by the input artist, a vocal tag-based music search service or artist search service, a service for searching for various music based on input of a user prompt, as well as a service for remixing, editing, and / or generating music based on input of a user prompt.
[0131] Communication module (170a, 170b)
[0132] According to one embodiment of the present invention, the communication module (170a, 170b) may include a wireless communication module or an RF module. The wireless communication module may include, for example, Wi-Fi, BT, GPS, or NFC.
[0133] According to one embodiment of the present invention, the communication module (170a, 170b) may provide a wireless communication function using a radio frequency. Additionally or alternatively, the wireless communication module may include a network interface or modem, etc. for connecting the user terminal (100a) and / or the local terminal (100b) to a network (e.g., the Internet, LAN, WAN, telecommunication network, cellular network, satellite network, POTS, or 5G network, etc.).
[0134] According to one embodiment of the present invention, the RF module may be responsible for transmitting and receiving data, for example, transmitting and receiving RF signals or called electronic signals. For example, the RF module may include a transceiver, a power amp module (PAM), a frequency filter, or a low noise amplifier (LNA). In addition, the RF module may further include components for transmitting and receiving electromagnetic waves in free space in wireless communication, for example, a conductor or a wire.
[0135] The communication module (170a, 170b) according to one embodiment of the present invention can use various wired communication systems such as a public switched telephone network (PSTN), xDSL (x Digital Subscriber Line), RADSL (Rate Adaptive DSL), MDSL (Multi Rate DSL), VDSL (Very High Speed DSL), UADSL (Universal Asymmetric DSL), HDSL (High Bit Rate DSL), and a local area network (LAN).
[0136] In addition, the communication module (170a, 170b) presented in this specification can use various wireless communication systems that can be realized now and in the future, such as mobile communication systems such as 4G and 5G (LTE), and satellite communication systems such as Starlink.
[0137] According to one embodiment of the present invention, the communication modules (170a, 170b) can be configured regardless of the communication mode, such as wired or wireless, and can be configured as various communication networks, such as a personal area network (PAN) and a wide area network (WAN). In addition, the network may be the well-known World Wide Web (WWW), and may also utilize a wireless transmission technology used for short-distance communication, such as infrared (IrDA: Infrared Data Association) or Bluetooth. The technologies described herein can be used not only in the networks mentioned above, but also in other networks.
[0138] Cloud server (200a) and local server (200b)
[0139] FIG. 2B is a diagram illustrating the configuration of a cloud server (200a) and a local server (200b) according to an embodiment of the present invention. Referring to FIG. 2B, the cloud server (200a) may include at least one of one or more processors (210a), memories (220a), interfaces (250a), and communication modules (270a). Referring to FIG. 2B, the local server (200b) may include at least one of one or more processors (210b), memories (220b), interfaces (250b), and communication modules (270b). The configuration of the cloud server (200a) and / or local server (200b) of the present invention is not limited to the components illustrated in FIG. 2B. That is, additional components may be included or some of the components illustrated in FIG. 2B may be omitted depending on the implementation aspect of the embodiments of the cloud server (200a) and / or local server (200b) according to the present invention.
[0140] According to one embodiment of the present invention, the cloud server (200a) and / or local server (200b) may be a digital device equipped with a processor, memory, and computing power, such as a laptop computer, notebook computer, desktop computer, web pad, or mobile phone. The cloud server (200a) may be a web server that processes services. The types of servers described above are merely examples, and the present invention is not limited thereto.
[0141] According to one embodiment of the present invention, the cloud server (200a) may be a server that provides a cloud computing service. More specifically, the cloud server (200a) may be a server that provides a cloud computing service, a type of Internet-based computing that processes information using another computer connected to the Internet rather than the user's computer. The cloud computing service may be a service that stores data on the Internet and allows users to access it anytime and anywhere via an Internet connection without having to install necessary data or programs on their own computers. Furthermore, data stored on the Internet can be easily shared and transmitted with simple operations and clicks. Furthermore, the cloud computing service may be a service that not only stores data on a server on the Internet, but also performs desired tasks using the functions of application programs provided on the Web without the need for separate program installation. Furthermore, the cloud computing service may be a service that allows multiple people to simultaneously share and work on documents. Furthermore, the cloud computing service may be implemented in at least one form among Infrastructure as a Service (IaaS), Platform as a Service (PaaS), Software as a Service (SaaS), a virtual machine-based cloud server, and a container-based cloud server. That is, the cloud server (200a) according to one embodiment of the present invention may be implemented in the form of at least one of the aforementioned cloud computing services. The specific description of the aforementioned cloud computing services is merely exemplary, and may include any platform that constructs the cloud computing environment of the present invention.
[0142] Meanwhile, according to embodiments of the present invention, the cloud server (200a) and / or the local server (200b) may be configured as a single server or may be configured as a plurality of servers.
[0143] According to one embodiment of the present invention, the cloud server (200a) and / or the local server (200b) may be servers that store one or more training data sets for training a neural network. The one or more training data sets may include at least one of various information and files, such as raw music files, component music files, stem files, music caption information, music basic information, tag information, and embeddings, according to embodiments of the present invention. The information stored in the cloud server (200a) and / or the local server (200b) may be utilized as training data, verification data, and test data for training the neural network in the present invention.
[0144] At least one of the above-described cloud servers (200a) and / or local servers (200b) can exchange data with at least one of the user terminals (100a) and / or local terminals (100b) via communication modules (270a, 270b). Alternatively, data can be exchanged with at least one of the user terminals (100a) and / or local terminals (100b) based on a data exchange tool (Tool) such as an API. Accordingly, services according to embodiments of the present invention can be provided via the user terminals (100a) and / or local terminals (100b).
[0145] Processor (210a, 210b)
[0146] According to one embodiment of the present invention, the processor (210a, 210b) controls the cloud server (200a) and / or the local server (200b) as a whole. According to one embodiment of the present invention, the processor (210a, 210b) may be a general-purpose processor (e.g., CPU), but may also be an AI-specific processor (e.g., GPU, TPU) for learning and / or executing an artificial intelligence model.
[0147] According to one embodiment of the present invention, the processor (210a, 210b) may include an AI processor. According to one embodiment of the present invention, the AI processor may learn a neural network using a program stored in the memory (220a, 220b). According to one embodiment of the present invention, the AI processor may adjust and learn parameters of an AI model for processing data related to the remixing, editing, and / or creation operations of music according to the present invention. Alternatively, the AI processor according to one embodiment of the present invention may perform an operation to improve the output by feeding back the output result of the AI model according to the present invention.
[0148] According to one embodiment of the present invention, the AI processor may include a data learning unit (not shown) that trains a neural network for data classification / recognition. The data learning unit may learn criteria regarding which learning data to use to determine data classification / recognition and how to classify and recognize data using the learning data. The data learning unit may acquire learning data to be used for learning and train the deep learning model by applying the acquired learning data to the deep learning model.
[0149] According to one embodiment of the present invention, the data learning unit (not shown) may be manufactured in the form of at least one hardware chip and mounted on a cloud server (200a) and / or a local server (200b). For example, the data learning unit may be manufactured in the form of a dedicated hardware chip for artificial intelligence, and may be manufactured as a part of a general-purpose processor (CPU) or a graphics processor (GPU) and mounted on the cloud server (200a) and / or the local server (200b). In addition, the data learning unit may be implemented as a software module. When implemented as a software module (or a program module including instructions), the software module may be stored on a non-transitory computer-readable medium that can be read by a computer. In this case, at least one software module may be provided to an operating system (OS) or provided by an application.
[0150] According to one embodiment of the present invention, a data learning unit (not shown) can learn to use acquired learning data to enable a neural network model to have judgment criteria regarding how to classify / recognize certain data. At this time, the learning method by the model learning unit can be classified into supervised learning, unsupervised learning, and reinforcement learning. Here, supervised learning refers to a method of training an artificial neural network in a state where labels for learning data are given, and the labels can mean the correct answer (or result value) that the artificial neural network should infer when learning data is input to the artificial neural network.
[0151] According to one embodiment of the present invention, unsupervised learning may refer to a method for training an artificial neural network without providing labels for training data. Reinforcement learning may refer to a method for training an agent defined within a specific environment to select actions or action sequences that maximize cumulative rewards in each state. In addition, the model training unit may train an artificial intelligence model using a learning algorithm including error backpropagation or gradient descent. Once the artificial neural network is trained, the trained artificial neural network may be referred to as an artificial intelligence model. The artificial intelligence model may be stored in memory (220a, 220b) as described below and used to infer results for new input data other than training data.
[0152] According to one embodiment of the present invention, the AI processor may further include a data preprocessing unit (not shown) and / or a data selection unit (not shown) to improve analysis results using an artificial intelligence model or to save resources or time required for generating an artificial intelligence model.
[0153] According to one embodiment of the present invention, a data preprocessing unit (not shown) may preprocess acquired data so that the acquired data can be used for learning / inference for situational judgment. For example, the data preprocessing unit may extract feature information as preprocessing for input data received via the communication modules (270a, 270b), and the feature information may be extracted in the form of a feature vector, feature point, or feature map.
[0154] According to one embodiment of the present invention, a data selection unit (not shown) can select data required for learning from among learning data or learning data preprocessed in a preprocessing unit. The selected learning data can be provided to a model learning unit. For example, when data acquired through the input unit (140a, 140b) of a user terminal (100a) and / or a local terminal (100b) is received through a communication module (270a, 270b), the data selection unit can detect a specific area among the received data, thereby selecting only data for objects included in the specific area as learning data. In addition, the data selection unit can select data required for inference from among input data acquired through the input unit (140a, 140b) and received through the communication module (270a, 270b) or input data preprocessed in a preprocessing unit.
[0155] In addition, according to one embodiment of the present invention, the AI processor may further include a model evaluation unit (not shown) to improve the analysis results of the neural network model. The model evaluation unit (not shown) inputs evaluation data into the neural network model, and if the analysis results output from the evaluation data do not satisfy a predetermined standard, it may cause the model learning unit to relearn. In this case, the evaluation data may be preset data for evaluating the artificial intelligence model. For example, the model evaluation unit may evaluate that the predetermined standard is not satisfied if the number or ratio of evaluation data with inaccurate analysis results among the analysis results of the trained neural network model for the evaluation data exceeds a preset threshold. According to one embodiment of the present invention, the model evaluation unit may also improve the model's inference ability through feedback.
[0156] Memory (220a, 220b)
[0157] According to one embodiment of the present invention, the memories (220a, 220b) can store various files and information necessary to provide the services of the present invention. The memories (220a, 220b) are accessed by the processors (210a, 210b), and data can be read / recorded / modified / deleted / updated by the processors (210a, 210b). In addition, the memories (220a, 220b) can store neural network models (e.g., deep learning models) generated through learning algorithms for data classification / recognition. Furthermore, the memories (220a, 220b) can store not only artificial intelligence models, but also input data, learning data, learning history, etc.
[0158] According to one embodiment of the present invention, the memory (220a, 220b) may include built-in memory or external memory. The built-in memory may include at least one of volatile memory (e.g., dynamic RAM (DRAM), static RAM (SRAM), synchronous dynamic RAM (SDRAM), etc.) or non-volatile memory (e.g., one time programmable ROM (OTPROM), programmable ROM (PROM), erasable and programmable ROM (EPROM), electrically erasable and programmable ROM (EEPROM), mask ROM, flash ROM, NAND flash memory, NOR flash memory, etc.). According to one embodiment, the built-in memory may take the form of a solid state drive (SSD). The external memory may further include a flash drive, for example, a compact flash (CF), a secure digital (SD), a micro secure digital (Micro-SD), a mini secure digital (Mini-SD), an extreme digital (xD), or a memory stick.
[0159] According to one embodiment of the present invention, the memory (220a, 220b) can store a computer program for performing a method of remixing, editing, and / or generating music according to one embodiment of the present invention, and the stored computer program can be read and executed by the processor (210a, 210b). In addition, the memory (220a, 220b) can store any type of information generated or determined by the processor (210a, 210b) and any type of information received through the communication module (270a, 270b). In addition, the memory (220a, 220b) can store data regarding a user prompt. In addition, the memory (220a, 220b) can store data regarding remixing, editing, and / or generating music. For example, the memory (220a, 220b) can temporarily or permanently store input / output data.
[0160] According to one embodiment of the present invention, the memory (220a, 220b) may include at least one type of storage medium among a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, and an optical disk. The cloud server (200a) and / or the local server (200b) may operate in relation to web storage that performs the storage function of the memory (220a, 220b) on the Internet. The description of the above-described memory is merely an example, and the present invention is not limited thereto.
[0161] Interface (250a, 250b)
[0162] An interface (250a, 250b) according to one embodiment of the present invention serves as a passageway for various types of external devices connected to a cloud server (200a) and / or a local server (200b). Such an interface (150a, 150b) may include at least one of a wired / wireless headset port, an external charger port, a wired / wireless data port, a memory card port, a port for connecting a device equipped with an identification module, an audio I / O (Input / Output) port, a video I / O (Input / Output) port, and an earphone port. In the cloud server (200a) and / or the local server (200b), appropriate control related to the connected external device (e.g., a user terminal (100a) and / or a local terminal (100b)) may be performed in response to the connection of the external device to the interface (250a, 250b).
[0163] In addition, the interface (250a, 250b) according to one embodiment of the present invention can perform a data transfer operation between at least one of the processor (210a, 210b), the memory (220a, 220b), and the communication module (270a, 270b). For example, the interface (250a, 250b) can provide data regarding a user's input received through the communication module (270a, 270b) to the processor (210a, 210b). For example, the interface (250a, 250b) can transmit a command or data received from the processor (210a, 210b) and the memory (220a, 220b) through a bus to another electronic device through the communication module (270a, 270b).
[0164] Communication module (270a, 270b)
[0165] According to one embodiment of the present invention, the communication module (270a, 270b) may include a wireless communication module or an RF module. The wireless communication module may include, for example, Wi-Fi, BT, GPS, or NFC.
[0166] According to one embodiment of the present invention, the communication modules (270a, 270b) may provide wireless communication functions using radio frequencies. Additionally or alternatively, the wireless communication modules may include a network interface or modem, etc., for connecting the cloud server (200a) and / or the local server (200b) to a network (e.g., the Internet, LAN, WAN, telecommunication network, cellular network, satellite network, POTS, or 5G network, etc.).
[0167] According to one embodiment of the present invention, the RF module may be responsible for transmitting and receiving data, for example, transmitting and receiving RF signals or called electronic signals. For example, the RF module may include a transceiver, a power amp module (PAM), a frequency filter, or a low noise amplifier (LNA). In addition, the RF module may further include components for transmitting and receiving electromagnetic waves in free space in wireless communication, for example, a conductor or a wire.
[0168] The communication module (270a, 270b) according to one embodiment of the present invention can use various wired communication systems such as a public switched telephone network (PSTN), xDSL (x Digital Subscriber Line), RADSL (Rate Adaptive DSL), MDSL (Multi Rate DSL), VDSL (Very High Speed DSL), UADSL (Universal Asymmetric DSL), HDSL (High Bit Rate DSL), and a local area network (LAN).
[0169] In addition, the communication module (270a, 270b) presented in this specification can use various wireless communication systems that can be realized now and in the future, such as mobile communication systems such as 4G and 5G (LTE), and satellite communication systems such as Starlink.
[0170] According to one embodiment of the present invention, the communication modules (270a, 270b) can be configured regardless of the communication mode, such as wired or wireless, and can be configured as various communication networks, such as a personal area network (PAN) and a wide area network (WAN). In addition, the network may be the well-known World Wide Web (WWW), and may also utilize a wireless transmission technology used for short-distance communication, such as infrared (IrDA: Infrared Data Association) or Bluetooth. The technologies described herein can be used not only in the networks mentioned above, but also in other networks.
[0171] network
[0172] The network according to embodiments of the present invention may use various wired communication systems such as Public Switched Telephone Network (PSTN), xDSL (x Digital Subscriber Line), RADSL (Rate Adaptive DSL), MDSL (Multi Rate DSL), VDSL (Very High Speed DSL), UADSL (Universal Asymmetric DSL), HDSL (High Bit Rate DSL), and local area networks (LANs). In addition, the network presented herein may use various wireless communication systems such as CDMA (Code Division Multi Access), TDMA (Time Division Multi Access), FDMA (Frequency Division Multi Access), OFDMA (Orthogonal Frequency Division Multi Access), SC-FDMA (Single Carrier-FDMA), and other systems.
[0173] The network according to embodiments of the present invention can be configured regardless of the communication mode, such as wired or wireless, and can be configured as various communication networks, such as a personal area network (PAN) and a wide area network (WAN). In addition, the network may be the well-known World Wide Web (WWW), and may also utilize a wireless transmission technology used for short-distance communication, such as infrared (IrDA) or Bluetooth. The technologies described herein can be used not only in the networks mentioned above, but also in other networks.
[0174] Artificial intelligence model based on natural language processing
[0175] According to one embodiment of the present invention, an artificial intelligence (AI) model based on natural language processing (NLP) may be used in the process of providing music remixing, editing, and / or generation services. The NLP-based AI model may be designed to understand, interpret, generate, and manipulate human language. Such a model may utilize various computing technologies to analyze and process text or speech information, thereby enabling the machine to perform tasks such as language translation, sentiment analysis, text summarization, and conversational agents such as chatbots. The first step of an artificial intelligence model based on natural language processing according to one embodiment of the present invention may be to divide a text into smaller units such as words, subwords, or characters (e.g., tokens). For example, each word in a sentence such as "the cat is sleeping" or "it is raining in the forest" may be tokenized and processed. According to one embodiment of the present invention, once the text is tokenized, it must be expressed in a numerical format that the artificial intelligence model can understand. In order to express the text in a numerical format, techniques such as one-hot encoding, word embedding such as Word2Vec, GloVe, or FastText, and contextual embedding may be used.
[0176] An artificial intelligence model based on natural language processing according to one embodiment of the present invention may include a statistical language model that uses a statistical method based on the probability of a word sequence and a neural language model that uses deep learning technology. The neural language model according to one embodiment of the present invention may include an artificial intelligence model based on an RNN, an LSTM, a Transformer, etc. A Transformer architecture according to one embodiment of the present invention may include an encoder for processing an input sequence and converting it into a hidden representation, a decoder for obtaining the hidden representation and generating an output such as a translation or summary, and attention layers for allowing the encoder and decoder models to focus on relevant parts of the input text.
[0177] According to one embodiment of the present invention, powerful language models, such as large-scale language models (LLMs), may be used. A large-scale language model may refer to an artificial intelligence model based on natural language processing with billions of parameters. Large-scale language models exhibit strong performance in both few-shot and zero-shot tasks. Large-scale language models can be trained using text data from various domains, such as Wikipedia, GitHub, chat logs, medical articles, legal articles, books, and crawled web pages. If successfully trained, such models demonstrate the ability to understand words in various domains. According to one embodiment of the present invention, a music remixing, editing, and / or generation service may be provided based on data understanding utilizing a natural language processing-based artificial intelligence model.
[0178] Full service overview
[0179] Figure 3a is a flowchart illustrating an artificial intelligence model for providing a service according to one embodiment of the present invention and an overall method for providing a service according to one embodiment of the present invention. At least one of the models and information illustrated in Figure 3a may be stored in at least one of a user terminal (100a), a local terminal (100b), a cloud server (200a), and a local server (200b).
[0180] At least one of the steps of the method illustrated in FIG. 3a may be performed in at least one of the user terminal (100a), the local terminal (100b), the cloud server (200a), and the local server (200b). That is, although all steps of the overall method illustrated in FIG. 3a may be performed in one of the user terminal (100a), the local terminal (100b), the cloud server (200a), and the local server (200b), it should be understood that some steps of the overall method illustrated in FIG. 3a may be performed in one of the user terminal (100a), the local terminal (100b), the cloud server (200a), and the local server (200b), and at least some of the remaining steps may be divided and performed in the remaining devices among the user terminal (100a), the local terminal (100b), the cloud server (200a), and the local server (200b).
[0181] Referring to FIG. 3a, the process indicated by the solid line corresponds to the data processing process for at least one of music search, remixing, editing, and creation, and the process indicated by the dotted line corresponds to the process of generating one or more embeddings and one or more pieces of information included in the metadata of the component block database. Hereinafter, a method for providing a service according to one embodiment of the present invention will be described with reference to FIG. 3a and other drawings.
[0182] Creation / extraction / storage of data and information related to the service of the present invention
[0183] Figure 3b is a diagram illustrating an artificial intelligence model for obtaining information and / or files according to one embodiment of the present invention. Figure 4a is a table organizing one or more files for providing a service according to one embodiment of the present invention. Figure 4b is a table organizing one or more pieces of information for providing a service according to one embodiment of the present invention.
[0184] Referring to FIG. 3b, a raw music file (30) may be processed as input to a component music file extraction model (333), so that a component music file (31) may be generated and / or extracted. The component music file (31) may be referred to as a component block in this specification. Meanwhile, according to one embodiment of the present invention, one or more component music files (31) may be combined to form a stem file (32).
[0185] According to one embodiment of the present invention, tag information (33) can be extracted by processing a raw music file (30) and / or a component music file (31) as input to a tag information extraction model (333).
[0186] According to one embodiment of the present invention, a raw music file (30) and / or a component music file (31) may be processed as input to a music captioning model (302) to generate music caption information (36).
[0187] According to one embodiment of the present invention, the original music file (30) and / or the component music file (31) can be processed as input to the music basic information extraction model (334) to extract the music basic information (34). The music basic information can include tempo, key, signature, volume, structure information, etc. For a specific example, the music basic information extraction model (334) can extract the tempo of the music by calculating the BPM (beats per minute) of each original music file (30) and / or component music file (31). For example, the music basic information extraction model (334) can analyze the pitch of each original music file (30) and / or component music file (31) to extract in which key the song is being played (for example, the key of C, the key of G, etc.). For example, the music basic information extraction model (334) can analyze the beat of each raw music file (30) and / or component music file (31) to extract what beat structure the music follows, such as 4 / 4 beat, 3 / 4 beat, etc. For example, the music basic information extraction model (334) analyzes the loudness of each raw music file (30) and / or component music file (31) to extract volume information for each raw music file (30) and / or component music file (31). According to one embodiment of the present invention, the music basic information extraction model (334) can also extract structural information, such as which part of the song the component music file (31) belongs to (e.g., verse, chorus, etc.). The music basic information extraction model (334) for extracting structural information may include the hierarchical structure analysis model described herein.
[0188] Raw music files (30)
[0189] Referring to FIG. 4A, a raw music file (30) according to one embodiment of the present invention is a file in which raw music is recorded. The raw music file (30) may be a music file recorded or generated outside of a music-related generation model or application according to the present invention. The raw music file (30) may be human-created music, augmented music, licensed music (human-created music), etc. The raw music file (30) may also include a music file generated through a music-related generation model or application according to the present invention. The format of the raw music file (30) may be mp3, wav, aac, midi, etc. In this detailed description, it is assumed that there are a total of N raw music files by default. The raw music file (30) can be recorded in at least one of the memory (220a) of the cloud server (200a), the memory (220b) of the local server (200b), the memory (120b) of the local terminal (100b), and the memory (120a) of the user terminal (100a).
[0190] Component music files (31)
[0191] Referring to FIG. 4a, a component music file (31) is a file in which music extracted from a raw music file (30) is recorded, and in this specification, the component music file (31) may be referred to as a component block. The component music file (31) is a music file in which a component of the raw music file (30) is recorded. The component music file (31) may already exist together with the raw music file (30), or may be extracted by a generation model or application according to the present invention. The extraction of the component music file (31) may be performed in a local server (200b). The extraction of the component music file (31) may utilize a conventionally known algorithm. Alternatively, according to one embodiment of the present invention, the extraction of the component music file (31) may be performed through a component music file extraction model (331). The format of the component music file (31) is the same as the format of the raw music file (30), and may be mp3, wav, aac, midi, etc. There may be multiple component music files (31) for each raw music file (30). In this detailed description, it is assumed that there are i component music files (31) for each raw music file (30). The number of component music files (31) per raw music file (30) may vary depending on the raw music file (30). The component music file (31) may be recorded in the same location as the raw music file (30), and specifically, may be recorded in at least one of the memory (220a) of the cloud server (200a), the memory (220b) of the local server (200b), the memory (120b) of the local terminal (100b), and the memory (120a) of the user terminal (100a).
[0192] Stem file (32)
[0193] Referring to FIG. 4a, a plurality of component music files (31) may be combined and provided to the user in the form of a stem file (32). The stem file (32) may temporarily exist in the user terminal (100a) on which the application program is installed. If the original music file (30) is a licensed music file, it may be stored in the local server (200b) in the form of a stem file (32). The stem file (32) may refer to a file in which only a specific item is separated from the original music file (30). For example, if the stem is vocal, the stem file (32) may refer to a file in which only the vocal is extracted from the original music file (30) and stored separately. The stem file (32) takes into account the characteristics of music and may include vocal items, bass items, guitar items, MID items, rhythm items, etc. The stem file (32) may be a file in which a specific attribute is separated for a certain period of time from the original music file (30). Specifically, the sound source that constitutes a music file is composed of a human vocal and the sounds of various instruments combined to form a single result, and a stem file (32) may refer to a file in which data for a single attribute that constitutes a sound source is stored. For example, types of stems may include vocal, bass, piano, accompaniment, beat, melody, etc. The stem file (32) may be a file having the same time as the original music file (30), or may be data having only a portion of the entire time of the original music file (30), and may also be divided into several parts according to the range or function for a specific attribute. The stem file (32) of the local server (200b) may be divided into a plurality of component music files (31) in the local server (200b) by the deep learning model according to the present invention.
[0194] Tag information (33)
[0195] Referring to FIG. 4b, tag information (33) is information extracted from a local server (200b) based on a raw music file (30) or a component music file (31). The tag information (33) can be extracted by a deep learning model according to an embodiment. The tag information (33) can be extracted from the local server (200b). The format of the tag information (33) may be json, and other formats are also possible as long as they are in text format. There may be multiple component music files (31) for each raw music file (30). In this detailed description, it is assumed that there are j component music files (31) for the nth raw music file (30) among N raw music files (30). The number of tag information (33) per raw music file (30) may vary depending on the raw music file (30). The tag information (33) may have a hierarchical structure. Tag information (33) can be recorded in at least one of the memory (220a) of the cloud server (200a) and the memory (220b) of the local server (200b).
[0196] According to the present invention, tag information (33) may refer to information that tags the characteristics of music. The tagging may include information on the genre, mood, instrument, and creation period of the music. Specifically, music genres can include rock, alternative rock, hard rock, hip-hop, soul, classic, jazz, punk, pop, dance, progressive rock, electronic, indie, blues, country, metal, indie rock, indie pop, folk, acoustic, ambient, R&B, heavy metal, electronica, funk, and house. Moods in music can include sad, happy, mellow, chill, easy listening, catchy, sexy, chillout (beautiful), and party. Instruments in music can include guitar, male vocalist, female vocalist, and instrumental music. The creation period of music can include information about the era in which the music was created, such as the 1960s, 1970s, 1980s, 1990s, 2000s, 2010s, and 2020s.
[0197] Music Basics (34)
[0198] Referring to FIG. 4b, this is information extracted from a local server (200b) based on a raw music file (30). The basic music information may already exist with the raw music file (30), or may be extracted using a conventionally known method or a deep learning model according to an embodiment. The extraction of the basic music information may be performed on the local server (200b). The format of the basic music information (34) may be json, and other formats may also be used as long as they are in text format. m pieces of basic music information (34) may be extracted per component music file (31). The basic music information (34) may be recorded in at least one of the memory (220a) of the cloud server (200a) and the memory (220b) of the local server (200b). Specifically, the basic music information (34) is important information for the service and thus may exist in the memory (220a) of the cloud server (200a).
[0199] Music Caption Information (36)
[0200] Referring to FIG. 4b, music caption information (36) is information describing music in natural language used in a music-related model according to an embodiment of the present invention. Music caption information (36) may also be referred to as description information in terms of service. It is information extracted from a local server (200b) based on a raw music file (30) or a component music file (31). Music caption information (36) can be extracted using a deep learning model according to an embodiment. Music caption information (36) can be extracted from the local server (200b). The format of music caption information (36) may be json, but other formats are also possible as long as they are in text format.
[0201] Music caption information is extracted as a number (k) related to the length of music per component music file (31), which may be referred to as intermediate music caption information. According to one embodiment of the present invention, intermediate music caption information may be collected based on a music caption information collection model (304). Based on the k pieces of intermediate music caption information, ultimately, one piece of collected music caption information exists per component music file (31), which may be referred to as final music caption information. In summary, k pieces of intermediate music caption information may be collected by the music caption information collection model (304) according to one embodiment of the present invention, and final music caption information may be generated. The intermediate and / or final music caption information may be recorded in at least one of the memory (220a) of the cloud server (200a) and the memory (220b) of the local server (200b). Specifically, the final music caption information is important information for the service and may therefore exist in the memory (220a) of the cloud server (200a). Meanwhile, the intermediate music caption information may only exist in the local server (200b) during the process of being extracted by a deep learning model according to an embodiment.
[0202] Create component music files from raw music files
[0203] FIG. 5A is a conceptual diagram illustrating a process of obtaining a component music file from a raw music file according to an embodiment of the present invention. FIG. 5B is a conceptual diagram illustrating a database of component music files according to an embodiment of the present invention. FIG. 6 is a conceptual diagram illustrating a configuration of a hierarchical structure analysis model according to an embodiment of the present invention. FIG. 7 is a conceptual diagram illustrating a configuration of a transformer-based artificial intelligence model included in a hierarchical structure analysis model according to an embodiment of the present invention. FIG. 8 is a conceptual diagram illustrating a neighboring attention model included in a hierarchical structure analysis model according to an embodiment of the present invention. Hereinafter, a process of obtaining a component music file (31) from a raw music file (30) according to an embodiment of the present invention will be described in detail with reference to FIGS. 5A to 8.
[0204] As illustrated in FIG. 5A, according to one embodiment of the present invention, a process of blocking music information included in a raw music file (30) into component music files (31) is disclosed. The process of blocking into component music files (31) refers to dividing and analyzing music information included in a raw music file (30) according to various characteristics and storing them in block units. In the present invention, a component music file (31) may be referred to as a component block, and a process of generating a component music file (31) may be referred to as component blocking.
[0205] Component blocking according to one embodiment of the present invention can be performed by the component music file extraction model (333) illustrated in FIG. 3B. The component music file extraction model (333) can generate a component music file (31) for a raw music file (30) through structural analysis based on a hierarchical structure analysis model and source separation based on a source separation algorithm. Specifically, according to the present invention, music information included in the raw music file (30) can be separated according to temporal structure and instrumental elements. The process of blocking into component music files (31) according to one embodiment of the present invention can be performed by analyzing rhythmic and structural information of music by utilizing a deep learning-based model and a hierarchical structure analysis model based on a neighboring attention mechanism discussed with reference to FIGS. 6 to 8, and then dividing the information into component blocks and storing the information in a database. Hereinafter, the component blocking process will be described in detail with reference to the drawings.
[0206] Preprocessing and source separation of raw music files
[0207] Referring to FIG. 5A, when a raw music file (30) according to one embodiment of the present invention is input, a source separation process may be performed to separate the input raw music file (30) into one or more instrument tracks. Since the input raw music file (30) according to one embodiment of the present invention is music in which one or more instruments are mixed, the roles of individual instruments cannot be distinguished. Therefore, by extracting each instrument track through a source separation algorithm, analysis and processing of individual tracks in a subsequent processing process become possible.
[0208] In a source separation process according to one embodiment of the present invention, an artificial intelligence-based source separation algorithm may be used. For example, instruments such as vocals, drums, bass, guitar, and rhythm may be separated, allowing each instrument track to be processed independently. The instrument tracks according to one embodiment of the present invention may be referred to as component tracks. In this way, each instrument track (or component track) is extracted from the input raw music file (30) through a source separation algorithm, enabling analysis and processing of individual tracks in a subsequent processing step. Subsequently, each instrument track (or component track) may be converted into a spectrogram. The converted spectrogram may be in a logarithmic scale, thereby obtaining high-dimensional data including time and frequency information. The spectrogram according to one embodiment of the present invention may be a mel-spectrogram to which a mel-filter bank is applied, but is not limited thereto.
[0209] Structural analysis using artificial intelligence models
[0210] FIG. 6 is a conceptual diagram illustrating the configuration of a hierarchical structure analysis model according to one embodiment of the present invention. FIG. 7 is a conceptual diagram illustrating the configuration of a transformer-based artificial intelligence model included in a hierarchical structure analysis model according to one embodiment of the present invention. FIG. 8 is a conceptual diagram illustrating a neighboring attention model included in a hierarchical structure analysis model according to one embodiment of the present invention.
[0211] Referring to FIG. 6, a process of performing beat and downbeat tracking and detecting segment boundaries and functional labels can be performed using an artificial intelligence model based on a transformer module. According to one embodiment of the present invention, a transformer-based artificial intelligence model can be simplified by referring to the efficiency of a lightweight TCN model. Specifically, one embodiment of the present invention can inherit the settings and pipeline of the TCN model for beat, downbeat, and tempo estimation. That is, the same input spectrogram (610) configuration and feature extraction module (620) settings as those of the CN model can be applied. The feature extraction module (620) can include three convolution (621) and max pooling (623) layers.
[0212] Also, referring to FIG. 6, an artificial intelligence model according to an embodiment of the present invention may be configured by stacking 11 sequence modeling blocks (e.g., a transformer module (630)). In addition, a dynamic Bayesian network (DBN) (640) may be used to post-process beats and downbeats. Since the TCN model focuses only on tracking beats and downbeats, in an embodiment of the present invention, a peak picking (650) method may be applied to post-process segment boundaries and functional labels. The peak picking method normalizes segment boundary probabilities using a sliding window average and selects the highest probability. In an embodiment of the present invention, a threshold may not be applied after normalization. In addition, unlike the 18-second window used in the TCN model, a 24-second window may be used. The specific figures described above are merely examples, and the present invention is not limited thereto.
[0213] Processing input data
[0214] According to one embodiment of the present invention, one or more instrument tracks (or component tracks) can be converted into spectrograms and processed as input to an artificial intelligence model as demixed spectrograms (610).
[0215] A transformer-based artificial intelligence model according to an embodiment of the present invention is excellent at learning long-term temporal dependencies, and thus can effectively analyze changes in rhythmic patterns or melodies that are repeated in music. A transformer-based artificial intelligence model according to an embodiment of the present invention can identify rhythmic patterns by tracking beats and downbeats (640) for a spectrogram of an input component track. In addition, a transformer-based artificial intelligence model according to an embodiment of the present invention can detect segment boundaries by detecting structural boundaries or transition points of music, and divide the music into individual segments. A transformer-based artificial intelligence model according to an embodiment of the present invention can structurally analyze music by assigning functional labels (e.g., start section, intro section, end section, outro section, break section, bridge section, inst section, solo section, verse section, chorus section, etc.) to each segment.
[0216] Meanwhile, according to one embodiment of the present invention, at least one of the assigned functional labels or predicted values may be integrated with each other. For example, a task such as integrating a Start section into an Intro section or an End section into an Outro section may be performed. According to one embodiment of the present invention, through the above integration task, an Intro section, a Verse section, a Chorus section, a Bridge section, an Outro section, etc. may be used as predicted values. However, the specific types of sections are merely examples for explanation, and the present invention is not limited thereto.
[0217] Feature extraction module (620)
[0218] Referring to FIG. 6, a separated spectrogram (610) according to one embodiment of the present invention can be processed as an input to a feature extraction module (620). The feature extraction module (620) according to one embodiment of the present invention can extract features from the input separated spectrogram (610), preprocess them, and transfer them to a transformer module (630), which is the next step. Specifically, the feature extraction module (620) according to one embodiment of the present invention can include convolution (621) and max pooling (623).
[0219] Convolution (621) and Max Pooling (623) play an important role in extracting features and reducing the dimensionality of music data in deep learning models.
[0220] According to one embodiment of the present invention, convolution (621) may refer to an operation that processes input data (e.g., an image or a spectrogram, etc.) and extracts features included therein. Through convolution (621) according to one embodiment of the present invention, local features such as rhythm, frequency, and timbre may be extracted by processing input music data. Convolution (621) according to one embodiment of the present invention applies a small-sized matrix called a filter or kernel to input data to detect a specific pattern (e.g., an edge, a texture, etc.).
[0221] This generates a feature map of the input data, and by gradually abstracting the data through multiple layers, high-level features can be extracted. According to one embodiment of the present invention, local patterns in a small area are detected through a convolution (621) filter, and through this process, important information such as rhythmic patterns and interactions between instruments in the frequency and time axes of the input spectrogram are learned. According to one embodiment of the present invention, nonlinear patterns can be modeled by applying a nonlinear activation function (e.g., ReLU, etc.) after the convolution (621) operation. Through this, the model can learn complex musical features beyond simple linear patterns.
[0222] According to one embodiment of the present invention, Max Pooling (623) refers to a downsampling technique for reducing unnecessary data and maintaining important information in a feature map generated or extracted through convolution (621). By Max Pooling (623) according to one embodiment of the present invention, the dimension of the feature map can be reduced while maintaining the main features extracted through convolution (621) by selecting the largest value to reduce the spatial size. By Max Pooling (623) according to one embodiment of the present invention, a result can be generated by selecting only the maximum value within a small area of given input data.
[0223] That is, max pooling (623) selects the most prominent (largest) feature among the features extracted from convolution (621), thereby removing less important details while maintaining important information in the data. This reduces the spatial dimensionality of the data, thereby reducing memory usage and computational load, strengthening invariance that is less sensitive to changes in the data's position, and reducing noise while retaining important information. This allows the model to better learn key rhythmic patterns or structural features without focusing on unnecessary details. According to one embodiment of the present invention, since dimensionality reduction through max pooling (623) reduces the number of parameters, it also helps prevent model overfitting, which can be particularly advantageous when the data size is limited.
[0224] According to one embodiment of the present invention, convolution (621) and max pooling (623) work complementarily to extract low-level features of input data and compress them so that they can be efficiently used in the learning process. Convolution (621) according to one embodiment of the present invention is used to extract low-level features from the input separated spectrogram (610), and max pooling (623) reduces the size of the feature map obtained from convolution (621) to retain only important information so that it can be passed on to the next stage. This process maximizes data processing efficiency and contributes to increasing the accuracy of model learning. Meanwhile, according to one embodiment of the present invention, as illustrated in FIG. 6, the operation of the feature extraction module (620) may be performed three times, but is not limited thereto.
[0225] Transformer module (630)
[0226] According to one embodiment of the present invention, data preprocessed through the feature extraction module (620) may be processed as input to the transformer module (630). The transformer module (630) according to one embodiment of the present invention may be configured to track beats and downbeats based on a neighborhood attention mechanism. The present invention utilizes the characteristic that music has a hierarchical organization composed of clear hierarchical structural units. The basic level of music is composed of meter elements such as beats, bars, and segments, which can form the basic rhythmic structure of music. As the hierarchy progresses, these meter elements are combined into functional units such as verses and choruses, forming the structure of the entire music. Despite the inherent interdependence of these hierarchical levels, previous studies have ignored this property and have divided the tasks into two: beat / downbeat tracking for low-level structure prediction and segment segmentation and labeling for high-level structure prediction.
[0227] This often leads to the potential benefits of inter-structural interactions being missed. However, jointly learning hierarchical information levels within an integrated model has been a challenge due to the relatively long timescales and high dimensionality of individual songs represented as audio data. Furthermore, songs contain diverse acoustic and musical variations within the beat and functional structure layers. In this invention, a single AI model based on Transformer modules simultaneously predicts beat, downbeat, segment, and functional structure labels, demonstrating the synergy between these tasks in multi-task learning.
[0228] According to one embodiment of the present invention, an efficient model capable of learning large-scale information from a sequence of audio frames covering long time units can be designed. In the task of tracking beats and downbeats, the receptive field of the model must be set large enough to cover a sufficient number of beats and downbeats. According to one embodiment of the present invention, a Temporal Convolutional Network (TCN) can be used, which is a type of convolutional neural network that applies a dilation operation in which the size of the receptive field increases exponentially as the number of layers increases.
[0229] According to one embodiment of the present invention, a SpecTNT-TCN model can be used, which uses a time-frequency transformer (SpecTNT) for efficient long-term representation learning and integrates it with a TCN module to achieve performance improvement.
[0230] According to one embodiment of the present invention, a Beat Transformer model that applies a dilation operation to a self-attention layer and uses source-separated inputs may be used.
[0231] Unlike beat / downbeat, temporal changes in segment boundaries and functional structural labels can be much rarer. Therefore, some previous studies have viewed segmentation as a self-similarity-based problem of local audio features within a song, exploiting temporal similarity, semantic labels, and structural labels to develop better audio features or embeddings. Other studies have explored segmentation algorithms that leverage the principles of uniformity, repetition, and novelty at the segment level.
[0232] However, embodiments of the present invention may employ at least one of the following three key features. First, one embodiment of the present invention utilizes a neighborhood attention mechanism to allow the model to generate an attention window that captures the closest possible neighbors without unnecessary computation. This can contribute to expanding the model's receptive area without padding.
[0233] Second, according to one embodiment of the present invention, the model can be configured to directly predict not only beats and downbeats, but also segment boundaries and functional structural labels from audio input. According to one embodiment of the present invention, performance interactions can be investigated when learning all tasks through comprehensive ablation studies.
[0234] Third, according to one embodiment of the present invention, the model size can be significantly reduced while following the settings of the TCN model. The model according to one embodiment of the present invention can be evaluated on the Harmonix Set, which includes all bits and functional structure labels. The model according to one embodiment of the present invention can be shown to outperform state-of-the-art performance on four tasks while maintaining a relatively small number of parameters (approximately 300,000).
[0235] According to one embodiment of the present invention, novel performance can be achieved based on Convolutional Neural Networks (CNNs) or Transformers that directly predict "boundaryness" or "chorusness" in audio. For example, according to one embodiment of the present invention, a dilated self-attention layer and source-separated input can be utilized in a bit-transformer model.
[0236] Referring to FIG. 6, a transformer module (630) according to one embodiment of the present invention may be composed of two different blocks based on two neighboring attention mechanisms: a dilated 1D neighboring attention (DiNA) block (631) and a 2D neighboring attention (NA) block (633). The dilated 1D neighboring attention (DiNA) block (631) according to one embodiment of the present invention models long-term temporal dependencies through dilation operations. The 2D neighboring attention (NA) block (633) according to one embodiment of the present invention maintains locality while modeling dependencies between instruments by focusing on local neighbors.
[0237] Details of the Neighborhood Attention Mechanism
[0238] FIG. 7 is a conceptual diagram illustrating the configuration of a transformer-based artificial intelligence model included in a hierarchical structure analysis model according to one embodiment of the present invention. Referring to FIG. 7, a 1D DiNA block (631) according to one embodiment of the present invention may include two DiNA modules (701, 703) with reference to a TCN model. According to one embodiment of the present invention, the second DiNA module (703) of the two DiNA modules may double the expansion. Accordingly, it is possible to learn musical properties at various levels of integer multiples. The outputs of the two DiNA modules (701, 703) are added to a skip connection (705), then combined (707), and then passed to the next layer.
[0239] A 1D DiNA block (631) according to one embodiment of the present invention may include a multilayer perceptron (MLP) (708) consisting of two fully connected layers. The multilayer perceptron (708) according to one embodiment of the present invention may initially increase the embedding dimension to 8C and then reduce the embedding size to the original size C to maintain consistency.
[0240] A 2D NA block (633) according to one embodiment of the present invention may include a 2D NA (709). Here, the 2D NA (709) is identical to the original, uninflated NA.
[0241] According to one embodiment of the present invention, the transformer module (604) consisting of the expanded 1D neighboring attention block (631) and the 2D neighboring attention (NA) block (633) can be performed a total of 11 times, and the size of the receptive area is expanded exponentially in each time. That is, the expansion of the receptive area size starts from 2 times, and progresses to 4 times, 8 times, and so on, so that the receptive area size is at most 2 10 and 2 11 can be increased. Accordingly, the receptive area sizes of the first and second DiNA modules can be set to approximately 41 seconds and 82 seconds, respectively. The embedding dimension C in all transformer blocks can be fixed to 24. The above-described figures are only an embodiment of the present invention, and the present invention is not limited thereto.
[0242] FIG. 8 is a conceptual diagram illustrating a neighboring attention model included in a hierarchical structure analysis model according to an embodiment of the present invention. Specifically, FIG. 8 exemplarily illustrates the application of the neighboring attention model to the last part of a song, and FIG. 8 is a conceptual diagram illustrating an attention window used in the neighboring attention model included in a hierarchical structure analysis model of music according to an embodiment of the present invention. The lower part of FIG. 8 illustrates the attention window of a 1D DiNA (701), and the upper part illustrates the attention window of a 2D NA (633).
[0243] The 1D DiNA (701), illustrated at the bottom of Figure 8, is used to model long-term temporal dependencies. It focuses on learning relationships between temporally distant intervals. According to one embodiment of the present invention, a window can be configured to secure a wider receptive area along the time axis by applying a dilation operation. Specifically, as illustrated at the bottom of Figure 8, a dilated window is a window to which a dilation operation is applied by sampling data at regular intervals along the time axis. DiNA according to one embodiment of the present invention can learn relationships between distant time frames using a dilated window. For example, it can capture long-term dependencies that occur between the chorus and verse of a song. Unlike a typical sliding window, a dilated window selects distant time frames at regular intervals and applies attention to them. This allows for the analysis of rhythmic patterns or structural changes occurring over a long time span. Furthermore, unlike a traditional sliding window, which analyzes only local information in a specific time interval, this allows for the analysis of relationships between distant time intervals.
[0244] As shown in Figure 8, DiNA can operate without padding even at the end of a song. This contributes to reducing computational waste caused by unnecessary padding when processing long-range data. For example, for a large receptive region such as 82 seconds, unlike conventional mechanisms that may require 41 seconds of padding in the worst case, 1D DiNA (701) does not require zero padding, thereby reducing unnecessary computational complexity.
[0245] The 2D NA illustrated in the upper part of Figure 8 can be used to model the interaction between instruments and learn the local interdependence of instruments. According to one embodiment of the present invention, a window can be configured based on a temporal frame and the neighborhood of each instrument. This allows for intensive analysis of data from multiple adjacent instruments within a specific time frame. According to one embodiment of the present invention, a neighboring attention window can be limited to neighboring instruments located around each instrument within a surrounding temporal frame and at that time. According to one embodiment of the present invention, the neighboring attention window can be used to efficiently capture the interaction between instruments and analyze the frequency band and timbre of each instrument. This allows the 2D NA to maintain the local relationship between instruments and focus on analyzing how the instruments harmonize. Furthermore, the 2D NA can adjust the frequencies so that they do not conflict with each other during simultaneous performances.
[0246] Referring to FIG. 8, one embodiment of the present invention can be designed to simultaneously learn local instrument-to-instrument interactions and long-term temporal relationships in music by using 1D DiNA and 2D NA mechanisms that play complementary roles.
[0247] Meanwhile, according to one embodiment of the present invention, the 1D DiNA block and the 2D NA block constituting the transformer module have a structure that performs multiple layers that exponentially expand the size of the dilated window of the DiNA block. The exponential expansion according to one embodiment of the present invention is such that the window increases by 2 over time. l (2 of l This can mean that the expansion becomes increasingly wider in a squared manner. Here, the exponent l represents the block number or layer number. That is, as each layer progresses, the expansion rate doubles, allowing for capturing a wider range of temporal data. Meanwhile, as illustrated in Fig. 8, the kernel size may be 5, but is not limited thereto.
[0248] In the present invention, blocks or layers are performed multiple times while exponentially expanding the window. According to one embodiment of the present invention, this multiple execution may mean that multiple transformer modules are sequentially stacked. By applying the expanded window to each module, an increasingly wider time range can be analyzed. In one embodiment of the present invention, this may be performed 11 times, taking into account the length of the music, but is not limited thereto.
[0249] For example, the first few modules can capture local dependencies within short time spans, while progressively deeper modules can capture dependencies over longer time spans. Early layers can learn detailed, local patterns in music, while later layers can contribute to learning the overall structure of a song or rhythmic changes over longer time spans. Subsequently, the Transformer module can capture musical patterns or structural changes over longer time spans (e.g., rhythmic changes between verses and choruses, repetitive patterns in a song, etc.) by using exponentially dilated windows. This plays a crucial role in learning long-term temporal dependencies. For example, early layers with a small dilation rate can learn detailed rhythms or instrumental interactions within short time spans in a song, while later layers with a large dilation rate can learn structural changes over longer time spans, such as the transition from verse to chorus.
[0250] According to one embodiment of the present invention, a method utilizing exponentially expanded windows can process data over a long time range simultaneously, resulting in higher computational efficiency than conventional sliding window methods. In particular, unnecessary redundant calculations can be reduced when analyzing long songs. Furthermore, an expanded attention mechanism, such as DiNA, according to one embodiment of the present invention can efficiently compute attention without padding even at the end of a song, significantly reducing the amount of computation.
[0251] Since the 1D DiNA block (631) and the 2D NA block (633) included in the transformer module (630) according to one embodiment of the present invention can analyze different time ranges, information can be extracted at various time scales, from local patterns in a short time range to structural patterns in a long time range. Through this, the complex rhythmic and structural characteristics of music data can be better captured. In addition, as the operation of the transformer module (630) is performed multiple times, the information learned in each module is gradually integrated, and through this, the overall structural information of music and detailed rhythmic patterns, which can be very important information in the process of music creation and remixing, can be learned simultaneously.
[0252] Post-processing
[0253] As illustrated in FIG. 6, according to one embodiment of the present invention, the beat and downbeat information are post-processed and optimized using a Dynamic Bayesian Network (DBN) (64). Through this post-processing, the rhythm pattern between each block can be adjusted to continue naturally.
[0254] Additionally, according to one embodiment of the present invention, a peak picking method based on a peak-picking model (650) can be applied to normalize the probabilities of segment boundaries and functional labels, and then select the boundary with the highest probability. This post-processing allows for smooth transitions between blocks.
[0255] Separation of sections through structural analysis along the time axis
[0256] Again, referring to FIG. 5A, FIG. 5A is a conceptual diagram illustrating a process of obtaining component music files from a raw music file according to an embodiment of the present invention. As illustrated in FIG. 5A, the raw music file (30) may be divided along the time axis according to its structure based on a hierarchical structure analysis model according to an embodiment of the present invention. The raw music file (30) may be divided into sections along the time axis. As illustrated in FIG. 5A, when the raw music file (30) according to an embodiment of the present invention is divided along the time axis, it may be divided into an intro section, a verse section, and a chorus section. The intro section corresponds to the introduction of a song, and may be a section that includes musical elements for setting the mood of the song and attracting the attention of the listener.
[0257] The verse is the section where the main narrative content of the song is conveyed. It can be a section that develops the story or deepens the theme through lyrics and melody. The chorus is the section where the central theme of the song is repeatedly expressed. It can be a section that provides an emotional impact to the listener through a memorable melody and strong rhythm and serves to reinforce the song's identity. Meanwhile, the description below is based on the original music file (30) being divided into the intro, verse, and chorus sections. However, the present invention is not limited to the division into the aforementioned sections.
[0258] That is, according to embodiments of the present invention, there are a Bridge section, which is a musical transition section that is distinguished from other main sections of the song, an Outro section, which is a section that serves to organize the mood of the song or to wrap it up gently similar to an intro as the concluding part of the song, a Pre-Chorus section, which is a transition section located between the verse and the chorus and plays a role in increasing tension or anticipation leading to the chorus, a Post-Chorus section, which is a section that follows the chorus and plays a role in emphasizing the highlight of the song by expanding or repeating the melody or theme of the chorus, an Interlude section, which is a short performance section inserted in the middle of the song and plays a role in raising the mood or diversifying the structure of the song through instrumental performance without vocals, a Drop section, which is a section mainly used in electronic music and where a strong beat and melody explode after a tense build-up, an Extended Chorus section, which is a section that is longer than a general chorus and plays a role in further maximizing the emotion of the song or extending the climax, and a melody or lyrics of the existing verse that are slightly It should be understood that it can be divided into various sections, such as the Variation Verse section, which is a section that changes the progression of the song by modifying it.
[0259] Creating component music files
[0260] As described above, according to one embodiment of the present invention, a source separation process may be performed to separate a raw music file (30) into one or more instrument tracks using a source separation algorithm. In addition, a structure analysis process may be performed to separate the raw music file (30) along the time axis using a hierarchical structure analysis model. Thereafter, as illustrated in FIG. 5A, one or more temporal and instrument-specific component music files (31) may be generated based on the separated sections of one raw music file. Through this, in each temporally separated session, music files related to tracks separated by instrument may be stored as independent component music files.
[0261] For example, each independent component music file can be created, such as a drum track in a chorus section, a guitar track in a verse section, etc. In this specification, a component music file may be referred to as a "component block," and the process of creating a component music file may be referred to as "component blockization."
[0262] Each component music file (or component block) may contain the rhythm or melody of one instrument within a certain time range, and according to one embodiment of the present invention, based on one or more component blocks separated in this way, an operation of reconstructing, remixing, editing, or generating music may be performed by selecting a block suitable for a specific instrument and a specific section.
[0263] According to one embodiment of the present invention, each section (e.g., verse, chorus, intro, etc.) can be further subdivided along the time axis and divided into component blocks. For example, in a 4 / 4 time signature song, a component track for a single measure can be created as a single component block. This method may involve analyzing long, complex musical data by dividing it into short time units using the DiNA block.
[0264] According to one embodiment of the present invention, each component block includes various musical characteristics, such as tempo, beat, downbeat information, segment boundaries and labels, and inter-instrument interaction. The information in each block can be refined and optimized through the aforementioned peak-picking and DBN post-processing methods.
[0265] Database of component music files
[0266] Figure 5b is a conceptual diagram illustrating a database of component music files according to one embodiment of the present invention. According to one embodiment of the present invention, each component block can be stored in the database along with associated metadata, such as tempo, beat, instrument composition, and segment information. This component block database is designed to be used later for remixing or music creation. For example, blocks matching a specific tempo and key can be searched for.
[0267] Furthermore, according to one embodiment of the present invention, each component block includes information of various dimensions. For example, one block may simultaneously store not only the tempo and tonality of the corresponding section, but also the instrument composition, frequency band, rhythm pattern, functional label, etc. Accordingly, one or more embeddings may be extracted for each component block based on one or more embedding models, and each component block may be stored in a database together with the extracted embeddings. For example, according to one embodiment of the present invention, the component block database may include at least one of a text embedding, a music text embedding, a music feature embedding, a chord progression embedding, a frequency band embedding, and basic music information (e.g., beat, tempo, tonality, volume, length, structure, etc.) extracted from the corresponding component block. Based on this, a search and filtering of the component block may be performed based on a similarity judgment between the input embedding and the embedding of the component block previously stored in the database.
[0268] That is, according to one embodiment of the present invention, a user can easily search for a component block that matches a functional label, tempo, instrument composition, etc. based on various data such as text data input through a user prompt, tag data input through a user prompt, image data, video data, music data, audio data, etc., and embeddings extracted in relation thereto, based on the similarity judgment with the embedding of the component block. The embedding model according to one embodiment of the present invention may be an embedding model based on a neighborhood attention mechanism. Similarity calculation based on embedding and search, filtering, etc. of component blocks will be described in detail later.
[0269] Data entry process through prompts
[0270] Referring to FIG. 3a, a user terminal (100a) and / or a local terminal (100b) according to an embodiment of the present invention may receive one or more pieces of information, such as tag information (33), image information, video information, music information, audio information, text information (e.g., natural language-based text, computer language-based text, etc.), through a user input unit (e.g., a user interface, a prompt, etc.). The input information may be transmitted to a cloud server (200a) and / or a local server (200b) according to an embodiment of the present invention, and processed or stored to provide a service of the present invention.
[0271] For example, text information according to one embodiment of the present invention may include natural language-based text such as "a funky 2 minute 30 second song with strong, danceable drum beats and a catchy bassline", "a scene of rain in a calm forest", etc. For example, image information and / or video information according to one embodiment of the present invention may be images and / or videos that can induce mental images such as "a club atmosphere", "a calm forest atmosphere", etc. Music and / or audio information according to one embodiment of the present invention may be music and / or audio having an atmosphere such as "a funky song", "a calm song", etc. It should be understood that the examples of the above-described information are for illustrative purposes only and that the present invention is not limited thereto.
[0272] Captioning process for input data
[0273] According to one embodiment of the present invention, when image and / or video information is input, caption information for the input image and / or video information may be generated through an image captioning model (301). According to one embodiment of the present invention, when music and / or audio information is input, caption information for the input music and / or audio information may be generated through a music captioning model (302). Here, at least one of the image and / or video information and the music and / or audio information may be input by a user, but at least one of data previously stored in a database according to an embodiment of the present invention (e.g., a video file stored in the database, a component music file, etc.) may be input.
[0274] According to one embodiment of the present invention, caption information may be natural language-based information generated for various types of information, such as input images, videos, music, and audio. The caption information according to one embodiment of the present invention may be information that expresses the characteristics of the input information using natural language-based text information. Alternatively, the caption information according to one embodiment of the present invention may be information that explains the input information using natural language-based text information.
[0275] Image caption information according to one embodiment of the present invention may refer to caption information based on image information and / or video information. Music caption information according to one embodiment of the present invention may refer to caption information based on music information and / or audio information. Hereinafter, captioning models (e.g., image captioning model (301), music captioning model (302), etc.) according to one embodiment of the present invention will be described in detail.
[0276] Image captioning model (301)
[0277] According to one embodiment of the present invention, image information inputted through a user input unit (e.g., a user interface, a prompt, etc.) and / or from a database is processed as input to an image captioning model (301), and the image captioning model (301) can output image caption information, which is text information about the input image. For example, the output image caption information can include text describing the input image information, such as 'a night city view with several buildings and several cars on the street.'
[0278] Music Captioning Model (302)
[0279] A music captioning model (302) according to one embodiment of the present invention is an artificial intelligence model capable of describing musical characteristics, such as the genre or mood of music, and the instruments used, in natural language. According to one embodiment of the present invention, music information is processed into natural language data through the music captioning model (302), and a music file can be searched based on information input by a user (e.g., based on similarity with text data input by the user).
[0280] FIG. 9 is a diagram for explaining a process for generating a mock caption according to one embodiment of the present invention. FIG. 10 is a diagram for explaining an example of an instruction input to generate a caption according to one embodiment of the present invention. FIG. 11 is a diagram for explaining an example of an instruction input to generate a caption according to one embodiment of the present invention. FIG. 12 is a conceptual diagram for explaining the configuration of a cross-modal encoder-decoder transformer according to one embodiment of the present invention. FIG. 13 is a diagram for explaining a process for obtaining a music captioning model (302) according to one embodiment of the present invention. Hereinafter, a music captioning model (302) according to one embodiment of the present invention will be described with reference to FIGS. 9 to 13.
[0281] According to one embodiment of the present invention, music information input through a user input unit (e.g., a user interface, a prompt, etc.) is processed as input to a music captioning model (302), and the music captioning model (302) can output music caption information, which is text information about the input music.
[0282] According to one embodiment of the present invention, the process of generating and / or outputting music caption information can be understood as a task for searching music information by generating natural language information related to input music information. The natural language information related to music information can be mainly provided in the form of sentences, and is distinguished from other music meaning understanding tasks, such as tagging tasks that assign tags to music information. Conventionally, there has been a technology for generating music caption information using a deep encoder-decoder framework developed for deep learning machine translation. In addition, among the conventional technologies, a pre-trained music tagging model was used as a music encoder, and an RNN layer initialized with pre-trained word embeddings was utilized for text generation.
[0283] Other prior art techniques combined pre-trained harmonic CNN encoders with LSTM layers to introduce a temporal attention mechanism for audio-text alignment. Another prior art technique used pre-trained GPT-2 to generate playlist titles and descriptions. However, a challenge facing these prior art in music caption generation is the lack of large-scale public datasets. Previous techniques used privately produced music datasets (e.g., private data containing 44 million music-text pairs from video platforms), but these approaches were not reproducible or usable by other developers. Community-driven data collection initiatives have been proposed to address this data challenge, but the scale of existing datasets remains challenging for successfully training music caption generation models.
[0284] To address this issue, according to one embodiment of the present invention, a method for generating a dataset of music caption information by applying a natural language processing-based artificial intelligence model to a tagging dataset can be provided. According to one embodiment of the present invention, music caption information that is semantically consistent with the provided tags, grammatically accurate, clear, and has a rich vocabulary can be obtained. The above-described method, which approaches the dataset level according to one embodiment of the present invention, is simple and practical, and has the effect of resolving the difficulty of music captioning work with data. Once the dataset is generated according to one embodiment of the present invention, a music caption model can be easily trained through supervised learning.
[0285] Referring to FIG. 13, a method for obtaining a music captioning model (302) according to an embodiment of the present invention may include a step of generating a simulated caption using an artificial intelligence model based on natural language processing (S1300), a step of training a music captioning model (302) based on the generated simulated caption data set (S1320), and a step of transfer training a music captioning model (302) pre-trained based on the simulated caption data set based on label data (S1340).
[0286] Creating a mock caption dataset
[0287] FIG. 9 is a diagram for explaining a process of generating pseudo captions (or pseudo music captions) according to one embodiment of the present invention. FIG. 10 is a diagram for explaining an example of an instruction input to generate a caption according to one embodiment of the present invention. FIG. 11 is a diagram for explaining an example of an instruction input to generate a caption according to one embodiment of the present invention. Referring to FIGS. 9 to 11, a step (S1300) of generating a pseudo music caption using an artificial intelligence model based on natural language processing illustrated in FIG. 13 will be described.
[0288] Referring to FIG. 9, a process for generating a mock caption according to one embodiment of the present invention is illustrated. According to one embodiment of the present invention, each music file (e.g., a raw music file, etc.) may have a tag data set designated for the corresponding music stored in advance. The tag data set according to one embodiment of the present invention may be label data, but may also be tag data extracted from music information through a tag information extraction model (333).
[0289] Referring to FIGS. 10 and 11, tag datasets and instruction data input to a natural language processing-based artificial intelligence model in the process of generating a mock caption according to one embodiment of the present invention are illustrated. According to one embodiment of the present invention, a mock caption for music can be generated by inputting natural language-based instruction data into a natural language processing-based artificial intelligence model (e.g., LLM, etc.) together with a tag dataset specified for a music file. For example, a tag dataset specified for music information according to one embodiment of the present invention may include tags such as 'guitar, piano, percussion, calming melody, slow tempo'. For example, an instruction according to one embodiment of the present invention may include at least one of an instruction to output a long sentence (e.g., a writing instruction), an instruction to output a summary, an instruction to express and output a tag using a new term (e.g., a paraphrase instruction), and an instruction to predict a characteristic of music.
[0290] According to one embodiment of the present invention, an artificial intelligence model based on natural language processing generates and returns a sentence that can describe tag information of music according to a method given as a condition by the instruction, based on tag data and instruction data input through a prompt.
[0291] According to one embodiment of the present invention, the AI model is initially trained using a large corpus and extensive computing resources, and then fine-tuned using Reinforcement Learning with Human Feedback (RLHF) to achieve even better performance in interactions with instructions. According to one embodiment of the present invention, user feedback can be utilized to verify a tag data set designated for each song (e.g., pre-entered label tags or tags extracted using a tag information extraction model (333)).
[0292] Specifically, the user inputs at least one of text data (e.g., natural language information such as "make music suitable for a rainy day"), image data, and music data through a prompt to generate desired music, and accordingly, a mock caption generation process according to an embodiment of the present invention is performed to generate music caption information, and music generated based on the generated music caption information (the generated music may be the entire song, or may be a part of the song) is output.
[0293] If the user provides positive feedback, such as downloading the output music, this may mean that the tag data set specified in the simulated caption generation process according to one embodiment of the present invention is well specified for the music. Therefore, based on the user's feedback, in the process of learning the music captioning model (302) according to one embodiment of the present invention, information such as "this tag data set is a data set that has received positive feedback from the user," "image information input by the user," "music information input by the user," and "prompt information input by the user" is additionally input, thereby improving the output of the music captioning model (302).
[0294] Meanwhile, as illustrated in FIG. 10, according to one embodiment of the present invention, an artificial intelligence model based on natural language processing can be trained by utilizing ideal data or reference data (e.g., Ground Truth) that the model hopes to output as a simulated music caption.
[0295] Meanwhile, according to one embodiment of the present invention, various system prompts may be utilized, and are not limited to the examples of instructions illustrated in FIGS. 10 and 11. According to one embodiment of the present invention, when at least one of tag information, image caption information, music caption information, and text information is input, a system prompt is utilized to specify the output format and examples of a large-scale language model, respectively, thereby improving the output results of a simulated caption according to the input data.
[0296] Pre-training of a music captioning model (302) based on a simulated caption dataset
[0297] Referring to FIG. 13, a method for obtaining a music captioning model (302) according to an embodiment of the present invention may include a step (S1320) of training the music captioning model (302) based on a set of simulated caption data generated using an artificial intelligence model based on natural language processing. That is, a set of simulated caption data and music information generated according to an embodiment of the present invention may be designated as pseudo labels to pre-train the music captioning model (302). As illustrated in FIG. 13, a large number of simulated caption data sets may be generated through a step (S1300) of generating simulated captions using an artificial intelligence model based on natural language processing. In the step (S1320) of training the music captioning model (302) based on the generated mock caption data set, the mock caption data set generated in step S1300 is designated as a pseudo label, and the music captioning model (302) can be trained so that when a music file (e.g., a raw music file, a component music file, etc.) is input to the music captioning model (302), mock caption data designated as a pseudo label is output.
[0298] Creation of a music captioning model (302) through transfer learning
[0299] Referring to FIG. 13, a method for obtaining a music captioning model (302) according to one embodiment of the present invention may include a step (S1340) of transfer-learning a music captioning model (302) pre-trained based on a simulated caption data set based on a label data set. The music captioning model (302) according to one embodiment of the present invention may include an artificial intelligence model based on natural language processing.
[0300] The music captioning model (302) according to one embodiment of the present invention is an artificial intelligence model based on natural language processing, which can be implemented as a large-scale language model, and a system prompt can be utilized to instruct how the large-scale language model operates. The system prompt according to one embodiment of the present invention may include a system prompt for specifying the purpose and role of the music captioning model (302), a system prompt for specifying matters to be considered for the output of the music captioning model (302), a system prompt for specifying an output format or example of the music captioning model (302), a system prompt for indicating an example of user input for the music captioning model (302), etc.
[0301] For example, a system prompt for specifying the purpose and role of a music captioning model (302) according to one embodiment of the present invention may include instructions such as, "Given the given input and individual request, create the most appropriate music description."
[0302] System prompts for specifying considerations for the output of a music captioning model (302) according to one embodiment of the present invention may include instructions such as, "Add a musical description that matches the prompt entered by the user, but do not use expressions that are not related to musical characteristics.", "Never use sentences, words, or expressions used in image captions to describe music. Instead, only consider additional musical characteristics that go well with the image caption.", "The content of the music caption should only be referenced after all other attributes have been considered.", "After all other attributes have been considered, the tag content, such as genre / mood, extracted by the tag information extraction model should be referenced."
[0303] A system prompt for specifying an output format or example of a music captioning model (302) according to one embodiment of the present invention may include a directive such as "Please comprehensively describe the contents of each item regardless of the order of the input items. Please explain concisely and clearly."
[0304] Meanwhile, according to one embodiment of the present invention, the present invention is not limited to the examples of the above-described instructions, and various system prompts may be utilized. According to one embodiment of the present invention, when at least one of tag information, image caption information, music caption information, and text information is input, a system prompt is utilized to specify the output format and examples of the music captioning model (302), thereby improving the output results of music caption information according to the input data.
[0305] Text preprocessing model for input data (303)
[0306] Referring to FIG. 3A, a process of preprocessing at least one of input text information, image caption information, and music caption information may be performed. The text preprocessing process according to one embodiment of the present invention may be performed using a text preprocessing model (303). The text preprocessing process according to one embodiment of the present invention may be a process of processing information (or data) to improve the output result of the corresponding artificial intelligence model before processing at least one of the input text information, image caption information, and music caption information as input to an artificial intelligence model in a subsequent step. The text preprocessing model (303) according to one embodiment of the present invention may be an artificial intelligence model based on natural language processing.
[0307] A preprocessing step according to an embodiment of the present invention may be a step of processing at least one of input text information, image caption information, and music caption information. Specifically, the preprocessing may be performed by adjusting the amount of text included in at least one of the input text information, image caption information, and music caption information, or by adjusting keywords, etc., to increase the similarity of the input information with a music file (e.g., a component music file, etc.), thereby improving the output result. According to an embodiment of the present invention, music caption information may also be generated for a music file (e.g., a component music file) stored in a database. The music caption information for the music file stored in the database may be generated based on a music captioning model (302), or may be previously generated and stored as label data. Through the preprocessing step according to an embodiment of the present invention, by processing at least one of the input text information, image caption information, and music caption information, the similarity with the generated music caption information for the music file stored in the database may be increased, thereby improving the output result of the model (e.g., an embedding extraction result, an extraction result of generated music, etc.).
[0308] A text preprocessing model (303) according to one embodiment of the present invention is an artificial intelligence model based on natural language processing, which can be implemented as a large-scale language model, and a system prompt can be utilized to indicate how the large-scale language model operates. The system prompt according to one embodiment of the present invention may include a system prompt for specifying the purpose and role of the text preprocessing model (303), a system prompt for specifying information to be included in the output value of the text preprocessing model (303), a system prompt for specifying the output format or example of the text preprocessing model (303), a system prompt for indicating an example of user input for the text preprocessing model (303), etc.
[0309] For example, a system prompt for specifying the purpose and role of the text preprocessing model (303) according to one embodiment of the present invention may include instructions such as "You are a music expert. Based on the text data entered by the user, tag data of the music, image caption information, music caption information, etc., recommend and describe appropriate music."
[0310] A system prompt for specifying information to be included in the output of a text preprocessing model (303) according to one embodiment of the present invention may include instructions such as "Please provide a detailed description of the recommended music. At least one of the following information must be included: 1) genre, mood, theme and style, 2) characteristics of the main instrument or sound, 3) rhythmic pattern, 4) mood or emotional tone of the song."
[0311] System prompts for specifying the format or example of a text preprocessing model (303) according to one embodiment of the present invention may include instructions such as "Do not use sensitive language", "Limit your answers to the maximum number of tokens, but always answer in complete sentences", etc.
[0312] A system prompt for indicating an example of user input for a text preprocessing model (303) according to one embodiment of the present invention may include an example sentence such as, "User input may include at least one of text, an image caption, an image tag, a music caption, and a music tag."
[0313] According to one embodiment of the present invention, when the text "Music that is good to listen to in a cafe" and the information "Lo-Fi, Low Fidelity, Jazz, Beautiful" as music tags are input through a user prompt, the text preprocessing model (303) may process or generate text information such as "Music that goes well with the following description: [User prompt: Music that is good to listen to in a cafe, Music tags: Lo-fi, Jazz, Beautiful, Music description: We recommend Lo-fi Jazz, which combines calm melodies with a laid-back atmosphere. This genre creates a laid-back atmosphere that is perfect for a cafe. This style is characterized by soft and mellow instrumentation including piano, saxophone, and light beats. Expect laid-back rhythms with syncopated patterns that add to the laid-back atmosphere. The intimate and nostalgic atmosphere encourages comfort and creativity, and is ideal for conversation or alone thinking while drinking coffee.]"
[0314] The text data processed or generated in this way can be processed as input to a text embedding extraction model (311) and / or a music text embedding extraction model (312) according to an embodiment of the present invention to extract an embedding, and based on the extracted embedding, subsequent steps such as searching for component music files can be performed based on a determination of similarity with embeddings related to component music files previously stored in a database. In the present invention, the similarity between pieces of information can be calculated by extracting an embedding for each piece of information and based on the similarity between the extracted embeddings, and a detailed method will be discussed later.
[0315] Extracting embeddings for input data
[0316] Music data and information recorded in the form of files according to one embodiment of the present invention can be calculated and stored as embeddings using an artificial intelligence-based embedding extraction model. Embeddings can be expressed in the form of vectors representing data characteristics. Each embedding can express characteristics of music or sentences in vector form based on music files, caption information, text information, etc. According to one embodiment of the present invention, embeddings can be calculated and recorded using a deep learning model for music files, particularly component music files. Embeddings based on component music files can be recorded on a local server (200b) and a cloud server (200a). From a service perspective, recording on the cloud server (200a) may be preferable, but is not limited thereto. For example, embeddings based on component music files may be stored on at least one of a user terminal (100a), a local terminal (100b), or a local server (200b), thereby providing a service according to one embodiment of the present invention.
[0317] According to one embodiment of the present invention, an embedding process may be performed to express input text information (e.g., text information entered as a prompt, image caption information, music caption information, etc.) in a numerical format that an artificial intelligence model can understand. According to one embodiment of the present invention, as illustrated in FIG. 3A, a process of extracting embeddings for input data may be performed through a text embedding extraction model (311), a music text embedding extraction model (312), a music feature embedding extraction model (313), etc.
[0318] According to one embodiment of the present invention, as illustrated in FIG. 3A, a process of extracting an embedding for input data (e.g., a component music file) may be performed using various embedding models, such as a chord progression embedding extraction model (322) and a frequency band embedding extraction model (323). According to the present invention, one or more features of music may be analyzed and expressed based on one or more embeddings extracted using the embedding models. These embeddings are intended to quantify one or more features of music and express them as vectors through an artificial intelligence model, and play a key role in the process of remixing, editing, and / or generating music according to embodiments of the present invention. According to one embodiment, various services, such as a task of finding a specific piece of music or classifying similar pieces of music, may be performed based on these embeddings by determining the similarity between pieces of music or between pieces of music and sentences. A method of determining the mutual similarity between embeddings will be described later with reference to a similarity calculation model.
[0319] Text embedding
[0320] According to one embodiment of the present invention, the text embedding extraction model (311) can perform embedding on input text information (e.g., text information entered as a prompt, image caption information, music caption information, etc.). The text embedding extraction model (311) can generate an embedding that matches the input text information. According to one embodiment of the present invention, when a user inputs the text "a lively rhythm and a gentle piano," the input text can be embedded, and a component music file can be searched for based on the input text data by judging the similarity of the embedding. According to one embodiment of the present invention, when a user inputs information such as an image, video, music, or audio that can bring to mind the image of "a gentle forest with rain falling gently," caption information for the input data can be generated through a captioning model, the caption information can be embedded, and a component music file having an embedding similar to the input information can be searched for through the embedding.
[0321] According to one embodiment of the present invention, at least one of the input text information, music tag information, image caption information, and music caption information may be preprocessed based on a natural language processing-based artificial intelligence model and then input to a text embedding extraction model (311) to be output as a text embedding. Alternatively, according to one embodiment of the present invention, at least one of the input text information, music tag information, image caption information, and music caption information may not be preprocessed based on a natural language processing-based artificial intelligence model, but may be processed as an input to the text embedding extraction model (311). According to one embodiment of the present invention, at least one of the input text information, music tag information, image caption information, and music caption information may be integrated through a text preprocessing model (303), and the text embedding extraction model (311) may extract a text embedding corresponding to the integrated text information based on the integrated text information. Alternatively, even if not preprocessed, at least one of the input text information, music tag information, image caption information, and music caption information may be integrated, and the text embedding extraction model (311) may extract a text embedding corresponding to the integrated text information. The text embedding extraction model according to one embodiment of the present invention may utilize a text embedding API known to a person skilled in the art.
[0322] Music text embedding
[0323] FIG. 14 is a diagram illustrating a music text embedding extraction model according to one embodiment of the present invention. According to one embodiment of the present invention, a model can be designed so that music (e.g., component music files, etc.) and text information can be learned in the same embedding space. Specifically, music text embedding is a method of expressing music and text in the same embedding space. Music information can be embedded using an artificial intelligence model, and this can be mapped with a text representation to learn the relationship between music and text, thereby extracting an embedding. Music text embeddings extracted through a music text embedding model can reflect the correlation between music features and natural language descriptions.
[0324] Because music text embeddings reflect these relationships, they can enhance the search and matching effect between music and text. For example, if a user describes a specific mood in music using natural language, the embedding can be used to search for music blocks with similar characteristics. In this way, information about the style of a specific singer, information about a specific instrument, etc. can be learned more effectively. For example, if the input text information or tag information includes information about a specific instrument, and the input music also includes a performance by that instrument, the music text embedding model according to one embodiment of the present invention can more effectively learn information about the instrument by extracting a music text embedding based on the input music information and text information.
[0325] According to one embodiment of the present invention, at least one of the input text information, music tag information, image caption information, and music caption information may be preprocessed based on a natural language processing-based artificial intelligence model and then input to a music text embedding extraction model (312) to be output as a music text embedding. Alternatively, according to one embodiment of the present invention, at least one of the input text information, music tag information, image caption information, and music caption information may not be preprocessed based on a natural language processing-based artificial intelligence model, but may be processed as an input to a music text embedding extraction model (312).
[0326] According to one embodiment of the present invention, at least one of input text information, music tag information, image caption information, and music caption information is integrated through a text preprocessing model (303), and a music text embedding extraction model (312) can extract a music text embedding corresponding to the integrated text information based on the integrated text information. Alternatively, even in the case where no preprocessing is performed, at least one of input text information, music tag information, image caption information, and music caption information is integrated, and a music text embedding extraction model (312) can extract a music text embedding corresponding to the integrated text information.
[0327] Music feature embedding
[0328] Figures 15a to 15c are diagrams illustrating a music feature embedding extraction model according to embodiments of the present invention. A music feature embedding according to one embodiment of the present invention is an embedding that expresses the overall characteristics of music as a vector. A music feature embedding according to one embodiment of the present invention may be an embedding output by a model trained to output a vector based on features contained in music, such as genre, mood, tempo, and instrument composition.
[0329] Music feature embeddings allow deep learning models to learn the overall characteristics of music and map them to a vector space to calculate similarity between similar pieces of music. Music feature embeddings can be utilized when users search for music of a specific style or mood. For example, when a user selects a specific song, the embeddings can be used to find music blocks with similar musical characteristics, such as genre, mood, tempo, instrumentation, timbre, and rhythm, for remixing.
[0330] The music feature embedding extraction model according to an embodiment of the present invention may be a model trained to map various features of a sound source into a multidimensional embedding space. The music feature embedding extraction model converts the features of an input sound source into a mel-spectrogram format and represents them as independent axes, such as a genre axis and a mood axis, in the multidimensional space, enabling search and analysis tasks based on these axes.
[0331] FIG. 18 is a diagram for explaining a similarity determination method based on a music feature embedding space according to one embodiment of the present invention.
[0332] According to one embodiment of the present invention, the embedding space of the music feature embedding extraction model can support multidimensional retrieval. Furthermore, within the embedding space of the music feature embedding extraction model, example-based retrieval and tag-based retrieval can be used to evaluate the similarity of specific sound sources or to search for sound sources matching specific attributes.
[0333] Multidimensional retrieval, according to one embodiment of the present invention, may be a method for selecting sound sources suitable for multiple attributes by simultaneously considering multiple axes (e.g., mood axis, genre axis, etc.) of the embedding space. The present invention utilizes cosine similarity to provide efficient and accurate search results in the multidimensionally analyzed embedding space, and enables effective search for sound sources based on specific attributes according to user needs.
[0334] Referring to FIG. 18, the embedding-based similarity search according to an embodiment of the present invention can operate based on cosine similarity by evaluating and searching for the similarity of sound sources by utilizing vectors generated in the embedding space.
[0335] According to one embodiment of the present invention, an input query vector calculates a cosine similarity with other sound source vectors within a multidimensional embedding space, and based on this similarity value, a search method such as example-based retrieval, tag-based retrieval, and multidimensional retrieval can be executed. For example, when searching for a sound source including the mood tag "Sad" and the genre tag "Rock," data corresponding to each axis is independently mapped to a separate representation space, so that a sound source most suitable for the query input by the user can be effectively searched for in various ways.
[0336] Example-based retrieval according to one embodiment of the present invention may be a method of providing search results based on overall similarity by selecting a vector with the highest similarity to an input query vector.
[0337] Tag-based retrieval according to one embodiment of the present invention may be a method of searching for a sound source associated with a specific tag (e.g., “Rock”) input by a user by comparing the cosine similarity with a reference vector corresponding to the tag.
[0338] A music feature embedding extraction model according to one embodiment of the present invention can independently process various musical attributes by converting the features of sound sources into separate multidimensional embeddings, thereby achieving high accuracy in search and classification tasks focused on specific attributes. Furthermore, the separate representation space can be flexibly used in various search scenarios and offers the advantage of evaluating data similarity by focusing on specific musical attributes.
[0339] Figure 15a is a diagram illustrating a triplet-based music feature embedding extraction model according to one embodiment of the present invention. The triplet-based music feature embedding extraction model uses a triplet learning method to represent the features of music data in a multidimensional embedding space. This model operates based on three data relationships: anchor, positive, and negative samples, and utilizes hinge loss to optimize the relative positions of data within the embedding space. Specifically, the model can be trained to minimize the distance between anchors and positive samples and increase the distance between anchors and negative samples. Input data is provided to the model by converting the digital signal of the sound source into a mel-spectrogram format, which can be converted into a low-dimensional vector using the embedding function f(x). The converted vector represents the position in the embedding space, and the model learns and expresses the musical characteristics (e.g., genre, mood, etc.) of each sample. The loss function is defined based on hinge loss and is mathematically expressed as [Equation 1] below:
[0340] [Mathematical Formula 1]
[0341]
[0342] In mathematical expression (1), D(f(x a ),f(x n )), D(f(x a ),f(x p )) represents the distance (usually the cosine distance) between two vectors, and Δ is a margin value. This function can be optimized so that the distance between the anchor and the positive sample is smaller than the distance between the anchor and the negative sample by a certain margin value.
[0343] A triplet-based music feature extraction model according to one embodiment of the present invention can apply a masking function to achieve a disentangled representation. The masking function is designed to activate only specific dimensions for each musical attribute (e.g., genre, mood, etc.), thereby maintaining independence between attributes in a multidimensional space. During the learning process, each triplet sample operates only in a specific subspace of the corresponding attribute according to the activated masking function, and the model learns disentangled embeddings for each attribute.
[0344] Therefore, a triplet-based music feature extraction model according to one embodiment of the present invention independently learns various attributes of music data and creates an embedding space capable of evaluating or searching for similarity based on these attributes. This model can provide a structure that flexibly performs search tasks based on specific musical features and can be optimized for multidimensional music information retrieval and classification tasks.
[0345] FIG. 15b is a diagram for explaining a proxy-based music feature embedding extraction model according to an embodiment of the present invention. The proxy-based music feature embedding extraction model is a model that uses a proxy-based learning method to express the features of music data in a multidimensional embedding space. This model generates a proxy vector corresponding to each class (e.g., genre, mood, etc.) and performs learning based on the similarity between the input data and the proxy vector. The proxy acts as a center point of the class, and calculations are performed based on the proxy instead of directly comparing with data samples, which greatly improves learning efficiency. The input data is provided to the model by converting the digital signal of the sound source into a mel-spectrogram, and the embedding function f(x) is i ) is mapped to a high-dimensional vector. This vector represents the location of the sample within the multidimensional embedding space, and the loss function can be optimized by calculating the cosine similarity with the proxy vector of each class.
[0346] A proxy-based music feature embedding extraction model according to one embodiment of the present invention can use binary cross-entropy loss to process multi-label data. The predicted probability for each class is calculated by passing the cosine similarity with the proxy vector of each class through a sigmoid activation function, and the loss function can be mathematically expressed as shown in [Mathematical Equation 2] below.
[0347]
[0348] In [Equation 2], y z is the actual label, y z hat is the predicted probability. This allows the model to learn how close the input vector is to each proxy vector, and can be optimized to reduce the distance between the proxy center and the sample for each class.
[0349] Meanwhile, a proxy-based music feature embedding extraction model according to one embodiment of the present invention can utilize a masking function to achieve a disentangled representation. The masking function determines the subspaces activated based on specific attributes (e.g., genre, mood, etc.), and the masking process is applied when calculating the similarity between the proxy vector and the input vector. This maintains the independence between each attribute within the multidimensional space and generates an optimized embedding for each attribute.
[0350] Therefore, a proxy-based music feature embedding extraction model according to one embodiment of the present invention can enhance learning efficiency by utilizing proxy vectors, ensure independence between attributes in a multidimensional embedding space, and support search and classification tasks based on specific attributes of music data. A proxy-based music feature embedding extraction model according to one embodiment of the present invention can have a structure that provides high accuracy while minimizing direct comparisons between data.
[0351] FIG. 15c is a diagram for explaining a classification-based music feature embedding extraction model according to an embodiment of the present invention. The classification-based music feature embedding extraction model according to an embodiment of the present invention is a model that uses a classification-based learning method to express the features of music data in a multidimensional embedding space. This model processes multi-label data and focuses on learning an embedding space in which each class can be linearly separated. In addition, to achieve a disentangled representation, a masking function and a sub-dense layer are utilized to construct an independent subspace for each attribute. The input data is input to the model by converting a digital signal of a sound source into a mel-spectrogram, and the embedding function f(x) is i) is converted into a vector representation. This vector represents the location of the sample within the multidimensional embedding space, and the similarity can be calculated through the inner product with each class centroid.
[0352] Meanwhile, for separate representation, the classification-based music feature embedding extraction model according to one embodiment of the present invention is designed to operate in a specific subspace corresponding to each attribute (e.g., genre, mood, etc.). To this end, a masking function can be applied, which can activate only the dimensions corresponding to each attribute. In the activated subspace, a sub-dense layer is added to enable independent learning for each attribute. The sub-dense layer can optimize the attribute-specific embedding by calculating the inner product between the input vector and the class center point within the subspace.
[0353] Meanwhile, in the classification-based music feature embedding extraction model according to one embodiment of the present invention, the loss function uses binary cross-entropy loss to process multi-label data, and the loss function is as described through [Mathematical Formula 2] above in relation to the proxy-based music feature embedding extraction model. However, although the structure of the loss function is the same in the two models, the input vector and the comparison target are different in that the proxy-based model uses a proxy vector, and the classification Gibbon model uses a class center point.
[0354] Code progression embedding
[0355] Figures 16a and 16b are diagrams illustrating a chord progression embedding extraction model according to one embodiment of the present invention. According to one embodiment of the present invention, chord progression patterns in music can be extracted and embedded. Chord progression refers to the harmonic progression of a song and is used to find blocks with similar chord progressions when remixing, editing, and / or creating music.
[0356] According to one embodiment of the present invention, a chord progression embedding quantifies and represents a musical chord progression as a vector. It uses a deep learning model to learn musical chord progression patterns and maps them to a vector space, enabling the discovery of music with similar chord progressions. This embedding can be utilized when a user selects a chord progression for remixing. For example, it can be used to search for music with a specific chord progression or to identify suitable blocks for remixing.
[0357] Referring to Fig. 16a, a chord progression embedding according to an embodiment of the present invention may include a process of extracting and vectorizing a chord progression pattern of music. Fig. 16a is a diagram visually representing such a chord progression. Referring to Fig. 16a, chroma features can be expressed as two-dimensional data consisting of a time axis and 12 pitch classes. Chroma features are mainly used to express the harmonic characteristics of music, and according to chroma features, energy distribution divided into 12 pitch classes (e.g., C, C#, D, etc.) in each time frame can be visualized. In Fig. 16a, the vertical axis represents 12 different pitch classes, and the horizontal axis represents the time frame. The color intensity of each cell represents the relative intensity of a specific pitch class in the corresponding time frame, which is used to capture the characteristics of the chord progression.
[0358] A chord progression embedding extraction model according to one embodiment of the present invention learns chord progression patterns of music using chroma features as input, and maps them to a vector space to effectively compare or search for music with similar chord progressions.
[0359] Referring to Fig. 16b, a chord progression embedding extraction model according to an embodiment of the present invention may include an autoencoder for embedding the chord progression of music. Fig. 16b is a diagram visually illustrating the process of generating a chord progression embedding from input data and then restoring it. As illustrated in Fig. 16b, an audio signal is first converted into chroma features. Chroma is used to extract the harmonic characteristics of music and is input in an abbreviated form for each measure. The chroma data according to an embodiment of the present invention is used as input data and passed to the autoencoder model.
[0360] An autoencoder according to one embodiment of the present invention comprises an encoder and a decoder. As illustrated in Fig. 16b, the encoder compresses input chroma data into a low-dimensional embedding space, and in the process, generates a chord progression embedding, which is a vector that efficiently represents the chord progression pattern of the music. This embedding represents the harmonic structure of the song in a compressed form. Furthermore, as illustrated in Fig. 16b, the decoder receives the chord progression embedding generated by the encoder and restores it to the original chroma data. The restored data has the same form as the input data, indicating that the model has learned the important features (chord progressions) of the input without loss.
[0361] Based on the vectorized chord progression embedding according to one embodiment of the present invention, illustrated with reference to FIGS. 16A and 16B, the generated chord progression embedding can be used to search for or recommend songs with similar chord progressions. Furthermore, music component blocks with similar chord progressions can be identified and utilized for remixing or editing. Furthermore, new music can be created or existing songs can be expanded while maintaining harmony based on the chord progression embedding.
[0362] Frequency band embedding
[0363] Figures 17a and 17b are diagrams illustrating a frequency band embedding extraction model according to one embodiment of the present invention. Frequency band embedding according to one embodiment of the present invention is an embedding that analyzes the frequency distribution of music and quantifies the characteristics of a specific frequency range. The frequency bands can be analyzed and embedded so that the instrumental elements that make up the music (e.g., component tracks, instrument tracks) can be arranged harmoniously without overlapping each other. Frequency band embedding can be used to ensure that the frequency ranges of the instrumental elements do not overlap each other during the remixing or editing process, thereby achieving harmony.
[0364] For example, frequency-band embedding can be used to find appropriate instrumental elements or adjust the corresponding frequency ranges to prevent mid- and high-frequency instruments from clashing with each other. Furthermore, frequency-band embedding can be used to find appropriate instrumental elements or adjust the corresponding frequency ranges to prevent low-frequency instruments, such as bass and drums, from clashing within the same frequency range. Furthermore, embedding-based harmonization can be used to ensure that mid- and high-frequency instruments do not invade each other's frequency ranges. This approach minimizes sonic conflicts between instrumental elements (e.g., component tracks, instrument tracks) in music editing, ensuring optimal sound quality and harmony.
[0365] Figure 17a is a graph visually representing the frequency distribution of music based on frequency band embedding. The graph depicted in Figure 17a represents the energy distribution of the frequency bands of component blocks, each divided into low, mid, and high ranges, as curves. In the graph depicted in Figure 17a, the horizontal axis represents the frequency band in units of note pitch, and the vertical axis represents the energy magnitude in each frequency band.
[0366] Figure 17a is an example of frequency-band embedding, which represents the energy magnitude by frequency band in specific music data. Frequency-band embedding calculates the magnitude of each frequency band over time, then calculates the average over time and quantifies it into a vector, allowing for quantitative analysis of various musical ranges. This analysis can ensure that musical components (e.g., instrumental elements) are harmoniously arranged across different frequency bands.
[0367] Figure 17b is a diagram illustrating the frequency analysis and embedding process of music data. Figure 17b visually illustrates the process of analyzing input music data in the frequency domain through Constant-Q Transform (CQT), extracting energy for each frequency band, and generating an embedding based on this analysis.
[0368] The left side of Figure 17b shows the results of a Constant-Q Transform (CQT), representing time along the horizontal axis and note pitch along the vertical axis. Colors represent the energy level at the corresponding time and note pitch, with brighter colors indicating stronger energy. This transformation transforms music data into the time-frequency domain and is used to analyze the energy distribution for each note pitch.
[0369] The right portion of Figure 17b graphically displays the results of calculating the sum of energy in a specific frequency band based on the energy extracted through CQT analysis. In the graph on the right side of Figure 17b, the horizontal axis represents the frequency band in notes, while the vertical axis represents the energy value in decibels (dB). This graph visually demonstrates energy changes across frequency bands over time.
[0370] According to one embodiment of the present invention, this analysis can be used to quantify music data by frequency band and generate an embedding based on this quantification. The generated embedding quantitatively represents the frequency characteristics of the music data and can be utilized for tasks such as music similarity search, remixing, and editing. For example, excessive overlapping instrumental elements in a specific frequency band can be identified and adjusted, or music with specific frequency characteristics can be searched for.
[0371] Storage of embedding data
[0372] FIG. 4c is a table that organizes one or more embeddings for providing a service according to one embodiment of the present invention.
[0373] Referring to FIG. 4c, a chord progression embedding can be extracted from a component music file (31) based on a deep learning model according to an embodiment of the present invention in a local server (200b). One chord progression embedding can be extracted per component music file (31) and can have a binary format. The chord progression embedding according to an embodiment of the present invention can be stored in at least one of the memory (220a) of the cloud server (200a) and the memory (220b) of the local server (200b).
[0374] Referring to FIG. 4c, music feature embeddings can be extracted from component music files (31) based on a deep learning model according to an embodiment of the present invention in a local server (200b). One music feature embedding can be extracted per component music file (31) and can have a binary format. Music feature embeddings according to an embodiment of the present invention can be stored in at least one of the memory (220a) of the cloud server (200a) and the memory (220b) of the local server (200b).
[0375] Referring to FIG. 4c, text embeddings can be extracted based on a deep learning model according to an embodiment of the present invention in a local server (200b) based on text data such as text information, image caption information, and music caption information input as a prompt. Meanwhile, according to an embodiment of the present invention, when music caption information (36) is generated based on a component music file (31) and text embeddings are extracted based on the generated music caption information (36), one text embedding can be extracted per component music file (31) and can have a binary format. Text embeddings according to an embodiment of the present invention can be stored in at least one of a memory (220a) of a cloud server (200a) and a memory (220b) of a local server (200b).
[0376] Referring to FIG. 4c, a music text embedding can be extracted based on a deep learning model according to an embodiment of the present invention in a local server (200b) based on at least one of text data such as text information input as a prompt, image caption information, music caption information, and music files (e.g., a raw music file (30), a component music file (31), etc.). When a music text embedding is extracted based on a component music file (31), one music text embedding can be extracted per component music file (31) and can have a binary format. A music text embedding according to an embodiment of the present invention can be stored in at least one of a memory (220a) of a cloud server (200a) and a memory (220b) of a local server (200b).
[0377] Referring to FIG. 4c, frequency band embeddings can be extracted from component music files (31) based on a deep learning model according to an embodiment of the present invention in a local server (200b). One frequency band embedding can be extracted per component music file (31) and can have a binary format. Frequency band embeddings according to an embodiment of the present invention can be stored in at least one of the memory (220a) of the cloud server (200a) and the memory (220b) of the local server (200b).
[0378] Similarity calculation model based on extracted embeddings
[0379] According to one embodiment of the present invention, at least one of the characteristics of each music file (e.g., a component music file) in an embedding space can be expressed as a unique embedding. According to one embodiment of the present invention, various services can be performed, such as finding specific music or classifying similar music by determining similarity between music or between music and sentences based on these embeddings.
[0380] According to one embodiment of the present invention, a similarity calculation model can determine the mutual similarity between embeddings. That is, whether or not embeddings are similar to each other can be determined using the cosine similarity (Cosine Analogy) as shown in [Mathematical Formula 3] below.
[0381] [Equation 3]
[0382]
[0383] In the above [Equation 3], A and B are embeddings, A i and B irepresents each component of each embedding. For example, if one of the embeddings to be compared is A and the other is B, and the calculation is performed using the above mathematical formula, if the result is high, A and B can be judged to be relatively similar. Conversely, if the result of the calculation is low, A and B can be judged to be relatively dissimilar. Therefore, the mutual similarity of embeddings can be effectively judged based on these values.
[0384] According to one embodiment of the present invention, one or more embeddings can be used to search for component blocks. At this time, the component blocks can be searched based on the sum of the similarity scores between one or more embeddings. Depending on the type of input data and / or the search step, the similarity scores between one or more embeddings can be weighted the same or different. For example, when text embeddings, music text embeddings, music feature embeddings, chord progression embeddings, and frequency embeddings are all used to search for component blocks, the sum of the similarity scores can be calculated using the following [Mathematical Formula 4].
[0385] [Equation 4]
[0386] Similarity total = A×Similarity text +B×Similaritymusic-text+C×Similarity music +D×Similarity chord +E×Similarity frequency
[0387] In the above [Mathematical Formula 4], A, B, C, D, and E represent weights, and Similarity text , Similaritymusic-text, Similarity music , Similarity chord , Similarity frequencymeans similarity based on text embedding, similarity based on music text embedding, similarity based on music feature embedding, similarity based on chord progression embedding, and similarity based on frequency band embedding, respectively. total refers to the sum of one or more weighted similarities.
[0388] Meanwhile, according to one embodiment of the present invention, as described with reference to the above mathematical expression (4), a weight may be applied to at least one of the similarity based on text embedding, the similarity based on music text embedding, the similarity based on music feature embedding, the similarity based on chord progression embedding, and the similarity based on frequency band embedding, and a sum of one or more similarities to which the weights have been applied may be calculated. Based on the summed similarity score, selection of component blocks, search, editing, remixing, generation, etc. of music according to embodiments of the present invention may be performed.
[0389] The process of remixing, editing, or creating music based on component blocks.
[0390] FIG. 3A is a flowchart illustrating an artificial intelligence model for providing a service according to one embodiment of the present invention and an overall method for providing a service according to one embodiment of the present invention. FIG. 19A is a flowchart illustrating a component block filtering process according to one embodiment of the present invention. FIG. 19B is a flowchart illustrating a component block search process according to one embodiment of the present invention. FIG. 20 is a flowchart illustrating a section expansion process according to one embodiment of the present invention. FIG. 21 is a flowchart illustrating a component block editing process according to one embodiment of the present invention. FIGS. 22A and 22B are flowcharts illustrating a method for generating music based on multi-modality according to one embodiment of the present invention.
[0391] Referring to FIGS. 22A and 22B , a method for generating music based on a multi-modal basis according to one embodiment of the present invention is disclosed. As illustrated in FIG. 22A , the method for generating music based on a multi-modal basis according to one embodiment of the present invention may include a step (S110) of receiving at least one of text information, image information, music information, and video information.
[0392] Thereafter, a multi-modal based automatic music generation method according to one embodiment of the disclosed invention may include a step (S120) of extracting an embedding based on received information.
[0393] More specifically, the step of extracting an embedding (S120) may include at least one of a step of generating a text embedding from text information, a step of generating caption information for at least one of image information and music information and generating a text embedding and / or a music text embedding based on the caption information, a step of generating a music feature embedding from the music information, a step of generating a chord progression embedding from the music information, and a step of generating a frequency band embedding from the music information.
[0394] Thereafter, a multi-modal based automatic music generation method according to one embodiment of the disclosed invention may include a step (S130) of searching and selecting a reference component block based on the similarity between embeddings of component blocks of a database based on the extracted embedding.
[0395] Thereafter, a multi-modal based automatic music generation method according to one embodiment of the disclosed invention may include a step (S140) of filtering a database based on a selected reference component block.
[0396] Thereafter, a method for automatically generating music based on multi-modality according to one embodiment of the disclosed invention may include a step (S150) of selecting one or more component blocks from the database based on the similarity between the embeddings of component blocks included in the filtered database and the extracted embeddings.
[0397] Thereafter, a method for automatically generating music based on a multi-modal system according to one embodiment of the disclosed invention may include a step (S160) of generating a reference section based on a reference component block and one or more component blocks selected from a database.
[0398] Thereafter, a multi-modal based automatic music generation method according to one embodiment of the disclosed invention may include a step (S170) of generating one or more audio sections based on a reference section.
[0399] Thereafter, a multi-modal based automatic music generation method according to one embodiment of the disclosed invention may include a step (S180) of generating a music file in a song unit by listing a reference section and one or more audio sections in time series.
[0400] Therefore, the method and device for automatically generating music based on multi-modality according to the disclosed invention have the advantage of being able to generate various types of embeddings from input information including text information, image information, music information, and video information from a user, and to generate music in song units by combining multiple component blocks from an audio database based on the embeddings.
[0401] Referring to FIG. 22b, the step (S120) of extracting an embedding based on received information according to one embodiment of the present invention may include at least one of a step (S121) of generating a text embedding from text information, a step (S122-1) of generating caption information for at least one of image information and music information, a step (S122-2) of generating a text embedding and / or a music text embedding based on the caption information, a step (S123) of generating a music feature embedding from music information, a step (S124) of generating a chord progression embedding from music information, and a step (S125) of generating a frequency band embedding from music information.
[0402] Specifically, when one or more pieces of information, such as tag information, image information, video information, music information, audio information, and text information (e.g., natural language-based text, code-based text, etc.), are input through a user input unit (e.g., a user interface, a prompt, etc.), an embedding can be extracted or generated for at least one of the input pieces of information. Based on a similarity determination between an embedding extracted from a component block included in a database and an embedding extracted from information input through a user input unit (e.g., a user interface, a prompt, etc.), one or more component blocks, such as a reference component block, are selected, and when one or more audio sections are generated based on the selected component blocks, a music file can be generated based on the one or more audio sections.
[0403] In this way, music can be generated on a multi-modal basis based on various information (e.g., tag information, image information, video information, music information, audio information, text information, etc.) input through a user input section (e.g., user interface, prompt, etc.).
[0404] In addition, the method and device for automatically generating music based on multi-modality according to the disclosed invention have the advantage of being able to generate more natural and complete music by selecting a reference component block from a database based on embedding and then searching and selecting a component block based on the selected reference component block.
[0405] In addition, the method and device for automatically generating music based on multi-modality according to the disclosed invention automatically generates music in the form of a completed song reflecting the mood desired by the user based on data such as text, images, videos, and music input by the user, and updates the same in a database, thereby enabling the user to continuously generate and enjoy sound sources of a similar style. Referring to FIG. 3A, the method according to one embodiment of the present invention may include a component block filtering and search step (S300), a component block selection step (S320), a section expansion step (S340), and a component block editing and mixing step (S360).
[0406] As illustrated in FIG. 19a, according to one embodiment of the present invention, the filtering and searching step (S300) of the component block may include a step of selecting a reference component block from among one or more component blocks included in the database based on at least one of text data input through a user prompt, tag information input through a user prompt, image caption information, music information, and music caption information.
[0407] According to one embodiment of the present invention, a reference component block may be selected from a database based on a comparison between an embedding based on at least one of text data input through a user prompt, tag information input through a user prompt, image caption information, music information, and music caption information, and an embedding extracted from a component block in a database (e.g., at least one of a text embedding, a music text embedding, a music feature embedding, a chord progression embedding, and a frequency band embedding) and at least one of basic music information (e.g., beat, tempo, tonality, instrument information, etc.).
[0408] Thereafter, filtering of the component block database may be performed based on basic music information (e.g., beat, tempo, key, instrument information, etc.) related to the selected reference component block. After filtering of the component block database is performed, a component block may be searched and selected based on a comparison between at least one of a text embedding, a music text embedding, a music feature embedding, a chord progression embedding, and a frequency band embedding based on the component block included in the filtered database and an embedding based on at least one of text data input via a user prompt, tag information input via a user prompt, image caption information, music information, and music caption information.
[0409] According to one embodiment of the present invention, after one or more component blocks have been filtered, the filtered one or more component blocks can be retrieved. According to one embodiment of the present invention, based on each component block, the temporal order or instrument combination can be changed to remix or create new music. At this time, the interaction between instruments can be considered to ensure natural connections between blocks.
[0410] According to the present invention, based on the hierarchical structure analysis model of music described above, music files are divided into component blocks using neighborhood attention and transformer-based DiNA and NA blocks, and these blocks are stored in a database, thereby obtaining precise music information necessary for music remixing and creation. This allows blocks to be combined or newly created based on various musical needs of users (e.g., data input via a user input unit).
[0411] Music is created section by section based on selected component blocks, and the music can be remixed, edited, and / or created later through section expansion. Filtering and searching of component blocks will be described below with reference to the drawings.
[0412] Select a reference component block
[0413] According to one embodiment of the present invention, a reference component block may be selected from a database based on a comparison between an embedding based on at least one of text data input through a user prompt, tag information input through a user prompt, image caption information, music information, and music caption information, and an embedding extracted from a component block in a database (e.g., at least one of a text embedding, a music text embedding, a music feature embedding, a chord progression embedding, and a frequency band embedding) and at least one of basic music information (e.g., beat, tempo, tonality, instrument information, etc.).
[0414] As illustrated in FIG. 5b, the reference component block included in the database may have at least one of text embedding, music text embedding, music feature embedding, chord progression embedding, frequency band embedding, and basic music information (e.g., beat, tempo, composition, instrument information, etc.) as metadata, and in FIG. 19a, such metadata is illustrated as 'reference component block information'.
[0415] Component block filtering
[0416] Figure 19a is a flowchart illustrating a component block filtering process according to one embodiment of the present invention. Based on basic music information (e.g., beat, tempo, composition, instrument information, etc.) of a selected reference component block, one or more component blocks included in the database can be filtered.
[0417] According to one embodiment of the present invention, a database can be constructed by filtering component blocks having the same beat information as the beat information of a reference component block, filtering component blocks having tonality information that differs from the tonality information of the reference component block by within 2 Keys, and filtering component blocks having tempo information that differs from the tempo information of the reference component block by within 20 BPM, using a filtering algorithm. The above-described reference component block information and specific numbers are merely examples for explanation, and the present invention is not limited thereto. A component block search process can be performed based on the database filtered in this way.
[0418] Search component blocks
[0419] Fig. 19b is a flowchart illustrating a component block search process according to an embodiment of the present invention. Based on a comparison between at least one of the above-described embeddings and music basic information and data input through a user input unit (e.g., a user interface, a prompt, etc.), a search and selection for one or more component blocks may be performed from a filtered database. As illustrated in Fig. 5b, each component block included in the database may have at least one of text embedding, music text embedding, music feature embedding, chord progression embedding, frequency band embedding, and music basic information (e.g., beat, tempo, key, instrument information, etc.) as metadata.
[0420] Referring to FIG. 19B, an embedding extraction process may be performed on at least one of text information input through a user prompt, tag information input through a user prompt, image caption information, and music caption information. According to one embodiment of the present invention, at least one of text embedding, music text embedding, music feature embedding, chord progression embedding, and frequency band embedding may be extracted based on user input data. A component block search model based on a similarity calculation model may search and select component blocks based on a similarity comparison between one or more embeddings based on user input data and one or more embeddings of component blocks included in a filtered database.
[0421] For example, according to one embodiment of the present invention, searching and selecting component blocks can be performed based on a similarity comparison between the text embeddings of user input data and the text embeddings of component blocks included in a filtered database. For example, the similarity between text embeddings can be determined, and scores can be assigned to each component block in order of increasing similarity.
[0422] According to one embodiment of the present invention, component block search and selection can be performed based on a similarity comparison between the music text embeddings of user input data and the music text embeddings of component blocks included in a filtered database. For example, the similarity between music text embeddings can be determined, and scores can be assigned to each component block in order of increasing similarity.
[0423] According to one embodiment of the present invention, component block search and selection can be performed based on a similarity comparison between the music embeddings of user input data and the music embeddings of component blocks included in a filtered database. For example, the similarity between music embeddings can be determined, and scores can be assigned to each component block in order of increasing similarity.
[0424] According to one embodiment of the present invention, component block search and selection can be performed based on a similarity comparison between the code progression embeddings of user input data and the code progression embeddings of component blocks included in a filtered database. For example, the similarity between code progression embeddings can be determined, and scores can be assigned to each component block in order of increasing similarity.
[0425] According to one embodiment of the present invention, component block search and selection can be performed based on a similarity comparison between the frequency band embeddings of user input data and the frequency band embeddings of component blocks included in a filtered database. For example, the similarity between frequency band embeddings can be determined, and scores can be assigned to each component block in order of increasing similarity.
[0426] That is, according to one embodiment of the present invention, as discussed above with reference to the mathematical expression (4), a weight may be applied to at least one of the similarity based on text embedding, the similarity based on music text embedding, the similarity based on music feature embedding, the similarity based on chord progression embedding, and the similarity based on frequency band embedding, and a sum of one or more similarities to which the weights have been applied may be calculated. Based on the summed similarity score, selection of component blocks, search, editing, remixing, generation, etc. of music according to embodiments of the present invention may be performed.
[0427] Meanwhile, according to one embodiment of the present invention, as illustrated in FIG. 19B, search information may be updated based on the selected component block. That is, based on at least one of text embedding, music text embedding, music feature embedding, chord progression embedding, and frequency band embedding related to the selected component block, another component block in the filtered database may be searched for, and based on a similarity judgment of at least one of text embedding, music text embedding, music feature embedding, chord progression embedding, and frequency band embedding between the selected component block and another component block included in the filtered database, search and selection of an additional component block may be performed.
[0428] Create a baseline section and expand the section
[0429] Fig. 20 is a flowchart illustrating a section expansion process according to an embodiment of the present invention. As illustrated in Fig. 20, based on a reference component block, component blocks of other instrument elements included in a database can be searched and selected to create a reference section. Specifically, by a component block search model according to an embodiment of the present invention, one or more component blocks can be searched and selected from a database based on a determination of similarity with one or more embeddings of the reference component block. According to an embodiment of the present invention, a process of providing feedback based on the reference component block for the searched and selected one or more component blocks can be performed, and a component block for creating a reference section can be searched and selected as a result of the feedback. A music generation model according to an embodiment of the present invention can create a reference section by editing and mixing the reference component block and one or more selected component blocks.
[0430] The reference section may include a plurality of component blocks whose frequency bands do not overlap with each other. For example, the plurality of component blocks refer to data for a single item constituting a sound source. For example, the types of component blocks may include a Rhythm component block, a Bass component block, a Low component block, a Mid component block, a High component block, an FX component block, and a Melody component block. For example, the reference section may include a Rhythm component block, a Low component block, a Mid component block, a High component block, an FX component block, and the like. Here, Rhythm, Bass, Mid, High, FX, and Melody may be terms referring to a sound range, an instrument element, or a type of instrument track (component track). According to one embodiment of the present invention, a Bass component block may be included in a Low component block. In addition, Low may mean a Low-Range component block, Mid may mean a Mid-Range component block, and High may mean a High-Range component block.
[0431] According to one embodiment of the present invention, the music generation model can perform editing such as adjusting the positions of a plurality of component blocks, adjusting the tempo, adjusting the composition, or adjusting the length, and mix the plurality of component blocks for which editing has been completed to generate a reference section.
[0432] Referring to FIG. 20, a reference section according to one embodiment of the present invention may include a reference component block corresponding to Mid and a plurality of component blocks selected based on the reference component block.
[0433] Referring to FIG. 20, a music generation model according to an embodiment of the present invention can select a mid component block of another section as a reference component block of each audio section from a raw music file (30) including a reference component block of a reference section. At this time, the music generation model can sequentially arrange the mid component blocks in each audio section in the progression order of the raw music file (30), or can arrange the mid component blocks in a free order considering the progression of the song regardless of this. The mid component blocks arranged in this way can serve as reference component blocks of each audio section, and each audio section can be generated in the same way as the process of generating the reference section.
[0434] Specifically, a music generation model according to an embodiment of the present invention can select at least one of component blocks (e.g., a Rhythm component block, a Low component block, a High component block, an FX component block) corresponding to the remaining instrument elements, excluding a component block identical to the reference component block (in this embodiment, a Mid component block), from a database after a reference component block of an audio section is determined, and include it in the corresponding audio section. That is, based on a reference component block selected and arranged from a raw music file (30), a component block of another instrument element included in the database is searched, and the searched component block is selected and included in the corresponding audio section, thereby generating a new audio section. In this case, the search and selection of the component block can be performed based on a similarity judgment between one or more embeddings of the reference component block and one or more embeddings of the component blocks included in the database, as discussed above. In this manner, audio sections can be generated based on a reference section, and section expansion can be achieved by listing the generated audio sections chronologically with the reference section, resulting in remixing, editing, and / or generating music files.
[0435] Meanwhile, according to one embodiment of the present invention, the reference component block is not always composed of a mid component block, and the reference component block may be composed of other component blocks such as a low component block or a high component block.
[0436] Additionally, at least one of the reference section and the audio section generated by the music generation module according to one embodiment of the present invention may be composed of a Rhythm component block, a Low component block, a Mid component block, a High component block, and an FX component block, but each section may not include at least one of the Rhythm component block, the Low component block, the Mid component block, the High component block, and the FX component block.
[0437] Meanwhile, the subscript of each component block illustrated in Fig. 20 is information indicating the name of the section, and the superscript is information indicating the raw music file containing the corresponding component block. Referring to Fig. 20, each component block of each audio section is composed of component blocks extracted from the raw music file (30) containing each component block of the reference section.
[0438] For example, as illustrated in FIG. 20, the Mid component block of the reference section is a component block extracted from the first raw music file, and therefore, each Mid component block included in one or more audio sections generated based thereon can be selected from among the Mid component blocks extracted from the first raw music file.
[0439] In addition, as illustrated in FIG. 20, since the low component block of the reference section is a component block extracted from the second raw music file, each low component block included in one or more audio sections generated based thereon can be selected from among the low component blocks extracted from the second raw music file.
[0440] In addition, as illustrated in FIG. 20, since the high component block of the reference component music file is a component block extracted from the first raw music file, each high component block included in one or more audio sections generated based thereon can be selected from among the high component blocks extracted from the first raw music file.
[0441] If, among the high component blocks extracted from the first raw music file, there is no appropriate high component block based on the calculation of the embedding similarity with the reference component block of the corresponding audio section, the corresponding audio section may be configured without the high component block.
[0442] As described above, by generating audio sections based on component blocks (component music files) extracted from each raw music file (30), there is a technical effect that allows for the generation of more natural and complete music. In addition, since the music generation model according to one embodiment of the present invention generates audio sections in the above manner, reverberation can be arranged so that the musical connection between multiple section audio can be naturally formed.
[0443] Editing and remixing component blocks
[0444] FIG. 21 is a flowchart illustrating a component block editing and remixing process according to one embodiment of the present invention. As illustrated in FIG. 21, a process for editing and remixing one or more component blocks included in an audio section according to one embodiment of the present invention may be performed.
[0445] Referring to FIG. 21, an audio section according to an embodiment of the present invention is configured by selecting various component blocks such as Mid, High, Low, Rhythm, and FX, and each component block may have unique musical properties (e.g., basic music information) such as tempo (BPM), key, and range. However, if the unique musical properties included in one or more selected component blocks do not match each other, there is a concern that the music may feel awkward. Therefore, as illustrated in FIG. 21, a process for editing and remixing one or more component blocks included in the audio section may be required.
[0446] Referring to FIG. 21, various editing operations can be applied to component blocks, and each component block having musical properties adjusted by the editing operations can be newly arranged within an audio section. For example, the lengths of one or more component blocks included in an audio section can be adjusted to be the same. Furthermore, the musical properties of each component block can be edited to create harmony, for example, the tempo (BPM) of the High component block can be adjusted to -10 and the key can be changed to -2, and the number of repetitions of the Rhythm component block can be doubled. Furthermore, the volume of one or more component blocks can be individually adjusted to maintain a harmonious volume balance. Based on one or more component blocks edited in this way, a new audio section can be created, thereby improving the overall expression of the music and designing it to achieve harmony. By efficiently adjusting and optimizing the properties of component blocks during the music creation and editing process according to one embodiment of the present invention, the productivity of remixing work can be increased and new musical possibilities can be provided.
[0447] Each method according to the present invention can be implemented as a computer program. The computer program can be stored on a recording medium to execute each method according to the present invention.
[0448] In addition, each method according to the present invention may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable recording medium may include program commands, data files, data structures, etc., alone or in combination. The program commands recorded on the medium may be those specially designed and configured for the present invention or may be those known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands, such as ROMs, RAMs, and flash memories. Examples of program commands include not only machine language codes generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc. The above-mentioned hardware devices may be configured to operate as one or more software modules to perform the operations of the present invention, and vice versa.
[0449] In addition, although the above description focuses on examples, these are merely examples and do not limit the present invention. Those skilled in the art to which the present invention pertains will appreciate that various modifications and applications not exemplified above are possible without departing from the essential characteristics of the present invention. For example, each component specifically shown in the examples can be modified and implemented. In addition, differences related to such modifications and applications should be interpreted as being included within the scope of the present invention defined in the appended claims.
Claims
1. A method for providing music generated based on artificial intelligence, A step of receiving at least one of text information, tag information, image information, and music information input from a user terminal; A step of extracting one or more embeddings based on the received information through one or more embedding extraction models; A step of calculating a similarity between one or more embeddings based on the received information and one or more embeddings extracted from component blocks included in a database, the component blocks being extracted from a raw music file, through a similarity calculation model; A step of selecting a first reference component block from the database based on the calculated similarity; A step of filtering the database based on the above-mentioned selected first reference component block; selecting one or more component blocks from the filtered database; and A step of generating a reference section based on the first reference component and the one or more selected component blocks; Including, A method for providing music generated based on artificial intelligence.
2. In paragraph 1, A step of selecting a second reference component block having the same instrument elements as the first reference component block from the above raw music file; A step of selecting one or more additional component blocks from the filtered database based on the second reference component block; generating an audio section based on the second reference component block and the one or more additional component blocks; and A step of generating a music file by arranging the above reference section and the above audio section in time series; comprising; A method for providing music generated based on artificial intelligence.
3. In paragraph 1, At least one embedding extracted based on the received information and at least one embedding extracted from a component block included in the database, each of which includes at least one of a text embedding, a music text embedding, a music feature embedding, a chord progression embedding, and a frequency band embedding. A method for providing music generated based on artificial intelligence.
4. In paragraph 3, The step of calculating the similarity between one or more embeddings extracted based on the received information and one or more embeddings extracted from component blocks included in the database through a similarity calculation model is as follows. A step of calculating one or more similarity scores based on the text embedding, the music text embedding, the music feature embedding, the chord progression embedding, and the frequency band embedding; and A step of calculating the similarity by reflecting a weight according to the type of the embedding for the one or more similarity scores and adding the one or more similarity scores to which the weight is reflected; further comprising; A method for providing music generated based on artificial intelligence.
5. In any one of paragraphs 1 to 4, The step of filtering the database based on the above-mentioned selected first reference component block is: further comprising a step of filtering the database based on music basic information related to the first reference component block; A method for providing music generated based on artificial intelligence.
6. In a server for providing music generated based on artificial intelligence, Communication module; Memory for storing the database; and a processor; including; The above processor, At least one of text information, tag information, image information and music information input from a user terminal is received through the communication module, Extracting one or more embeddings based on the received information through one or more embedding extraction models stored in the memory, Computing the similarity between one or more embeddings based on the received information and one or more embeddings extracted from component blocks included in the database, wherein the component blocks are extracted from raw music files, through a similarity calculation model stored in the memory, Selecting a first reference component block from the database based on the calculated similarity, Filtering the database based on the first selected reference component block, Select one or more component blocks from the above filtered database, Generating a reference section based on the first reference component and one or more selected component blocks; A server for providing music generated based on artificial intelligence.
7. In paragraph 6, The above processor, From the above raw music file, a second reference component block having the same instrument elements as the first reference component block is selected, Selecting one or more additional component blocks from the filtered database based on the second reference component block, Generating an audio section based on the second reference component block and the one or more additional component blocks, Generating a music file by arranging the above reference section and the above audio section in time series, A server for providing music generated based on artificial intelligence.
8. In paragraph 6, At least one embedding extracted based on the received information and at least one embedding extracted from a component block included in the database, each of which includes at least one of a text embedding, a music text embedding, a music feature embedding, a chord progression embedding, and a frequency band embedding. A server for providing music generated based on artificial intelligence.
9. In paragraph 8, The above processor, Through the similarity calculation model stored in the above memory, Compute one or more similarity scores based on the text embedding, the music text embedding, the music feature embedding, the chord progression embedding, and the frequency band embedding, Reflecting a weight according to the type of the embedding for the one or more similarity scores above, and calculating the similarity by adding up the one or more similarity scores to which the weight is reflected. A server for providing music generated based on artificial intelligence.
10. In any one of paragraphs 6 to 9, The above processor, Filtering the database based on the music basic information related to the first reference component block, A server for providing music generated based on artificial intelligence.
11. A method for providing music generated based on artificial intelligence, A step of receiving text information, tag information, image information and music information input from a user terminal; A step of generating image caption information for the image information through an image captioning model; A step of generating music caption information for the above music information through a music captioning model; A step of performing preprocessing on the text information, the tag information, the image caption information, and the music caption information through a text preprocessing model; A step of generating at least one of a text embedding and a music text embedding based on the preprocessed information through one or more embedding extraction models; Including, A method for providing music generated based on artificial intelligence.
12. In paragraph 11, The above music captioning model is an artificial intelligence model that is pre-trained on a mock caption data set generated based on an artificial intelligence model utilizing a system prompt including one or more tags and one or more instructions, and is an artificial intelligence model that is transferred based on a label data set. A method for providing music generated based on artificial intelligence.
13. In paragraph 12, The one or more directives include at least one of a directive for specifying a purpose or role of the music captioning model, and a directive for specifying an output format of the simulated caption. A method for providing music generated based on artificial intelligence.
14. In any one of paragraphs 11 to 13, A step of calculating the similarity between at least one of the text embedding and the music text embedding based on the above preprocessed information and at least one of the text embedding and the music text embedding extracted from the component block included in the database, through a similarity calculation model; A method for providing music generated based on artificial intelligence.
15. In paragraph 14, The steps for calculating the above similarity are: Calculating at least one similarity score between at least one of a text embedding and a music text embedding based on the above preprocessed information and at least one of a text embedding and a music text embedding extracted from a component block included in a database, A step of calculating the similarity by reflecting a weight according to the type of the embedding for the one or more similarity scores and adding the one or more similarity scores to which the weight is reflected; further comprising; A method for providing music generated based on artificial intelligence.
16. In a server for providing music generated based on artificial intelligence, Communication module; Memory for storing the database; and a processor; including; The above processor, Receive text information, tag information, image information and music information input from the user terminal through the communication module, Generate image caption information for the image information through the image captioning model stored in the above memory, Generate music captioning information for the music information through a music captioning model stored in the above memory, Preprocessing is performed on the text information, the tag information, the image caption information, and the music caption information through the text preprocessing model stored in the memory. Generating at least one of a text embedding and a music text embedding based on the preprocessed information through one or more embedding extraction models stored in the memory. A server for providing music generated based on artificial intelligence.
17. In paragraph 16, The above music captioning model is an artificial intelligence model that is pre-trained on a mock caption data set generated based on an artificial intelligence model utilizing a system prompt including one or more tags and one or more instructions, and is an artificial intelligence model that is transferred based on a label data set. A server for providing music generated based on artificial intelligence.
18. In paragraph 16, The above music captioning model is a natural language processing-based artificial intelligence model that is pre-trained on a simulated caption data set generated based on a natural language processing-based artificial intelligence model utilizing a system prompt including one or more tags and one or more instructions, and is a natural language processing-based artificial intelligence model that is transferred based on a label data set. A server for providing music generated based on artificial intelligence.
19. In paragraph 18, The one or more directives include at least one of a directive for specifying a purpose or role of the music captioning model, and a directive for specifying an output format of the simulated caption. A server for providing music generated based on artificial intelligence.
20. In any one of paragraphs 16 to 19, The above processor, Through the similarity calculation model stored in the above memory, Computing the similarity between at least one of the text embedding and the music text embedding based on the above preprocessed information and at least one of the text embedding and the music text embedding extracted from the component block included in the database. A server for providing music generated based on artificial intelligence.
Citation Information
Patent Citations
Machine learning model trained based on music and decisions generated by expert system
WO2024178038A1
Hinge assembly and electronic device including the same
KR1020250018050A
Tape gripping device
KR1020250057360A
Method and apparatus for providing an audio mixing interface using a plurality of audio stems
KR102534870B1
Method and Apparatus for Searching Similar Music Based on Music Attributes Using Artificial Neural Network
KR102538680B1
Cited By
Film and television content label processing method and terminal based on large model
CN120873232A
All-in-One video restoration system, method and device based on expert system
CN121073836A