Artificial intelligence generated music streaming platform system
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- 박종민
- Filing Date
- 2025-12-31
- Publication Date
- 2026-08-03
Smart Images

Figure 112025149681482-PAT00003_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a streaming platform system dedicated to AI-generated sound sources, and more specifically, to a music streaming platform system that exclusively handles sound sources generated by artificial intelligence (AI) and a method of operating the same. Background Technology
[0002] Recently, music streaming services have become widely available to the general public through various platforms such as Melon, Bugs, and Genie. These existing music streaming platforms primarily provide services based on general music created by singers and distributed through record companies, and their copyright management systems are also optimized for traditional music distribution methods.
[0003] Meanwhile, driven by recent advancements in artificial intelligence, music generation technology utilizing AI is developing rapidly. Various AI music generation models have been developed, reaching a level where music can be automatically generated when a user inputs text prompts or specific conditions. Consequently, while the demand and supply for AI-generated audio are increasing, it remains virtually impossible for existing music streaming platforms to systematically register or manage such content.
[0004] Specifically, existing platforms mix AI-generated and standard audio without distinction, making it difficult to satisfy the demand of users who prefer only AI music; furthermore, there is a complete lack of functionality to systematically manage metadata unique to AI audio, such as generation model information, prompts, and generation conditions. Additionally, despite the fact that the copyright attribution and revenue distribution structures for AI audio differ from those of standard audio, no system has been established to reflect these differences.
[0005] Therefore, there is a growing need for a new type of music streaming platform that exclusively handles AI-generated audio and provides metadata management, search, recommendation, and revenue sharing functions optimized for the characteristics of AI audio. The problem to be solved
[0006] The purpose of the present invention is to solve the problems of the aforementioned prior art.
[0007] The objective of the present invention is to solve the problem that existing music streaming platforms are designed around human singers and copyrighted music, making it difficult to systematically register, manage, and distribute AI-generated music. In existing platforms, the registration of AI music is either impossible or is processed in the same way as general music, failing to reflect the unique characteristics of AI music.
[0008] The objective of the present invention is to solve the problem that it is difficult to satisfy the demand of users who wish to listen only to AI music because AI sound sources and general sound sources are mixed. It is necessary to provide a dedicated environment where users can select, listen to, and explore only AI-generated music.
[0009] The objective of the present invention is to solve the problem that existing platforms lack the functionality to systematically manage and utilize metadata unique to AI audio, such as copyright management, generation model information, prompt information, and generation conditions. Although the type and version of the model used during the generation process, as well as the content of the prompts, are important information for AI audio, existing systems do not have a structure to store or utilize such information. means of solving the problem
[0010] According to one embodiment of the present invention for achieving the above-mentioned purpose, a streaming platform system dedicated to AI-generated audio is provided, comprising: an AI audio database configured to register and store only audio generated by artificial intelligence; a format verification unit that verifies the format and quality of an audio file uploaded from a user terminal; an AI verification unit that determines whether the audio is generated by AI through metadata included in the audio file or an audio analysis algorithm; a registration processing unit that registers the audio file in the AI audio database when it is determined to be an AI-generated audio; and a streaming processing module that plays the audio stored in the AI audio database in a streaming manner.
[0011] The above system may further include a revenue distribution module that calculates cumulative usage by linking AI model information, prompt information, and creator identification information to the playback log of each sound source, and distributes revenue to the sound source creator and the AI model provider based on the cumulative usage.
[0012] The above system further includes a search and recommendation module that recommends audio tracks similar to a specific audio track selected by the user or entered in the form of a prompt, and the similarity between audio tracks can be calculated based on keyword similarity, genre similarity, and AI model match in the prompt. Effects of the invention
[0013] According to a streaming platform system dedicated to AI-generated sound sources according to one embodiment of the present invention, by providing an environment in which users can selectively listen to and explore only AI music through a catalog structure that exclusively handles AI-generated sound sources, it is possible to effectively satisfy the demands of AI sound source enthusiasts.
[0014] In addition, according to the present invention, by systematically managing metadata unique to AI sound sources, such as generation model information, prompt information, and generation conditions, it is possible to provide AI-specialized search and recommendation functions that were impossible in existing general sound source-centric platforms, thereby having the effect of significantly improving the user experience.
[0015] In addition, according to the present invention, by blocking the inflow of general sound sources through a function that verifies whether AI was generated and maintaining a pure catalog consisting only of AI sound sources, it is possible to secure the identity and reliability of the platform.
[0016] In addition, according to the present invention, by linking the playback logs and metadata of AI sound sources to establish a foundation for fair revenue distribution to sound source creators and AI model providers, it is possible to promote the healthy development of the AI music ecosystem. Brief explanation of the drawing
[0017] FIG. 1 is a schematic diagram showing the overall configuration of a streaming platform system dedicated to AI-generated sound sources according to one embodiment of the present invention. FIG. 2 is a block diagram showing the main components inside a server according to one embodiment of the present invention. FIG. 3 is a flowchart illustrating a sound source verification and registration process according to an embodiment of the present invention. FIG. 4 is an exemplary diagram showing a user interface of an AI sound source specialized search and recommendation function according to one embodiment of the present invention. Specific details for implementing the invention
[0018] The following detailed description of the invention refers to the accompanying drawings, which illustrate specific embodiments in which the invention may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the invention. It should be understood that various embodiments of the invention are different but need not be mutually exclusive. For example, specific shapes, structures, and characteristics described herein with respect to one embodiment may be implemented in other embodiments without departing from the spirit and scope of the invention. It should also be understood that the location or arrangement of individual components within each disclosed embodiment may be changed without departing from the spirit and scope of the invention. Accordingly, the following detailed description is not intended to be limiting, and the scope of the invention is limited only by the appended claims, including all equivalents to those claimed therein, provided appropriately described. Similar reference numerals in the drawings refer to the same or similar functions across various aspects.
[0019] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the attached drawings in order to enable a person skilled in the art to easily practice the present invention.
[0020] FIG. 1 is a schematic diagram showing the overall configuration of a streaming platform system dedicated to AI-generated sound sources according to one embodiment of the present invention.
[0021] Referring to FIG. 1, the system of the present invention is largely configured to include a user terminal (100) and a service server (300).
[0022] A user terminal (100) is a device that provides an interface for a user to listen to, explore, and upload AI-generated sound sources, and may include various types of electronic devices such as smartphones, tablets, desktop computers, and laptop computers. A dedicated application that provides the AI sound source-specific streaming service of the present invention is installed and executed on the user terminal (100). The application may be provided in the form of a mobile app or a web browser-based web application.
[0023] Data communication between the user terminal (100) and the service server (300) is performed via a network. The network is an intermediary communication network and may include various wired and wireless communication technologies such as the Internet, mobile communication networks, and wireless LANs (WLAN). The user terminal (100) receives audio streaming data from the service server (300) via the network and transmits commands, such as audio upload requests and search requests, to the service server (300).
[0024] The service server (300) is a server system that performs the core functions of the present invention and can be implemented as a plurality of physical servers or cloud-based virtual servers. The service server (300) stores and manages only AI-generated sound sources, transmits sound sources via streaming in response to a request from a user terminal (100), and provides AI sound source specialized search and recommendation functions. The detailed internal configuration of the service server (300) will be described later with reference to FIG. 2.
[0025] FIG. 2 is a block diagram showing the main components inside a service server (300) according to one embodiment of the present invention.
[0026] Referring to FIG. 2, the service server (300) is configured to include an AI sound source database (310), a sound source verification and registration module (320), a streaming processing module (330), a search and recommendation module (340), and a revenue distribution module (350).
[0027] The AI sound source database (310) is a database specifically designed to register and store only sound sources generated by artificial intelligence. For each sound source, the AI sound source database (310) stores metadata unique to the AI sound source as required fields, in addition to general sound source information such as title, genre, and playback time. Specifically, for each sound source, the AI sound source database (310) stores the name and version information (311) of the AI model used to generate the sound source, prompt information (312) entered by the user when generating the sound source, information on the conditions for generating the sound source (313), identification information of the sound source creator (314), and information on the license and revenue sharing rules (315) of the sound source.
[0028] The name and version information (311) of the AI model is information indicating the name and version of the AI model that generated the sound source, and may be stored in the form of, for example, "MusicGen", "Stable Audio", etc. The name of the AI model may be the name of a publicly announced AI service platform. The prompt information (312) is a field that stores text prompts or generation instructions entered by the user into the AI model, and may be stored in the form of, for example, "Exciting electronic dance music, BPM 128", "Calm piano ballad, emotional atmosphere", etc. The generation condition information (313) is various parameter values applied when generating the sound source, and may include the sampling rate, bit rate, length, type of instrument used, etc.
[0029] The creator identification information (314) stores a unique identifier of the user or service that created and uploaded the sound source, and is used to identify the creator when distributing revenue. The license and revenue distribution rule information (315) is information that defines the copyright license type of the sound source and rules on how playback revenue will be distributed among the creator, the AI model provider, and the platform operator.
[0030] The sound source verification and registration module (320) performs the function of verifying a sound source file uploaded from a user terminal (100) and registering it in the AI sound source database (310) only if it is confirmed to be an AI-generated sound source. The sound source verification and registration module (320) includes a format verification unit (321), an AI verification unit (322), and a registration processing unit (323).
[0031] The format verification unit (321) performs the function of verifying whether the file format of the uploaded audio file is a supported format and whether the sound quality is above a standard. For example, the format verification unit (321) can verify whether formats such as MP3, WAV, FLAC are supported and whether the bitrate is at least 128kbps.
[0032] The AI verification unit (322) is a key component that determines whether the uploaded audio is actually an audio generated by AI. The AI verification unit (322) can determine whether the AI generated the audio by analyzing metadata included in the audio file to check if it contains signature information or generation information of the AI model, or by using a deep learning-based classification algorithm that analyzes the acoustic characteristics of the audio. In some embodiments, a method may be used in which the user is required to submit AI model information and a prompt together when uploading, and this is verified.
[0033] The registration processing unit (323) registers the sound source identified as an AI-generated sound source by the AI verification unit (322) into the AI sound source database (310) and stores the necessary metadata together. Conversely, for sound sources determined not to be AI-generated, registration is blocked and the user is notified of the reason for rejection. Through this, the platform of the present invention can maintain a pure catalog consisting only of AI-generated sound sources.
[0034] The streaming processing module (330) performs the function of transmitting audio stored in the AI audio database (310) in real-time via streaming in response to a playback request from the user terminal (100). The streaming processing module (330) may apply adaptive streaming technology that adaptively adjusts the bitrate according to the user's network bandwidth, and records playback logs of the audio to be used for revenue distribution and statistical analysis.
[0035] The search and recommendation module (340) utilizes AI sound source-specific metadata to provide differentiated search and recommendation functions to the user. The search and recommendation module (340) provides a search interface that allows the user to search for sound sources based on various conditions such as the type of AI model, prompt keywords, mood of the music, beats per minute (BPM), and genre.
[0036] For example, the user can search by complex conditions such as "electronic music with a BPM of 120 to 130 generated by the MusicGen model," and the search and recommendation module (340) returns a list of music that satisfies these conditions. Additionally, the search and recommendation module (340) can provide functions such as filtering only music generated by a specific AI model, automatically adding music generated with similar prompts to a playlist, and recommending similar music by analyzing the prompt patterns of music that the user has liked.
[0037] The search and recommendation module (340) provides a recommendation function based particularly on prompt similarity, which utilizes the prompt similarity score calculated by the following mathematical formula 1.
[0038] [Mathematical Formula 1]
[0039]
[0040] In the above Equation 1, S(Pi, Pj) represents the prompt similarity score between source i and source j, and Pi and Pj represent the prompts of source i and source j, respectively. The first term represents the weighted sum of TF-IDF (Term Frequency-Inverse Document Frequency) for n keywords included in the prompt. That is, N is the number of keywords included in the prompts of source i and source j, Wk is a binary weight (0 or 1) indicating whether the k-th keyword matches, and TF-IDF k represents the TF-IDF score of the k-th keyword.
[0041] The second term, M(mi, mj), is a function representing the degree of agreement between the AI models that generated source i and source j, returning a value of 1 if they are the same model and 0 if they are different models. Here, mi and mj represent the identifiers of the AI models that generated source i and source j, respectively.
[0042] The third term, G(gi, gj), is a function representing the genre similarity between sound source i and sound source j, which returns a value of 1 if the genres match perfectly, a value of 0.5 if they are similar genres, and a value of 0 if they are completely different genres. Here, gi and gj represent the genre information of sound source i and sound source j, respectively. The similarity between multiple genres may be stored in the database (310).
[0043] α, β, and γ are weight parameters that adjust the relative importance of each term and are set to satisfy the condition α + β + γ = 1. In one embodiment, α can be set to 0.6, β to 0.25, and γ to 0.15, which means considering prompt keyword similarity as the most important factor, AI model match as the second most important factor, and genre similarity as the third most important factor.
[0044] The user can select a specific sound source i and send a request to the server (300) to recommend other similar sound sources. Alternatively, the user can send a request to recommend other similar sound sources by entering a specific prompt. In this case, the specific prompt becomes sound source i.
[0045] The search and recommendation module (340) calculates a prompt similarity score between the sound source currently being played by the user, the selected sound source, or the sound source entered in the form of a prompt and all other sound sources stored in the database using the above mathematical formula 1, and recommends the sound sources to the user by sorting them in order of highest similarity score. For example, sound sources with a similarity score of 0.7 or higher can be automatically added to a playback queue, or the top 10 similar sound sources can be presented as a recommendation list.
[0046] Additionally, the search and recommendation module (340) can provide a function to search by complex conditions such as "electronic music with a BPM of 120 to 130 generated by the MusicGen model," a function to filter only music generated by a specific AI model, and a function to recommend similar music by analyzing the prompt pattern of music that the user liked.
[0047] The revenue distribution module (350) performs the function of distributing revenue to the sound source creator and the AI model provider by linking the playback log of each sound source with the metadata stored in the AI sound source database (310). The revenue distribution module (350) calculates the cumulative number of plays and playback time by linking the AI model information (311), prompt information (312), and creator identification information (314) to the playback log of each sound source.
[0048] The revenue distribution module (350) distributes advertising revenue or subscription revenue among the sound source creator, AI model provider, and platform operator according to predefined license and revenue distribution rule information (315) based on the calculated cumulative usage. For example, if the playback revenue of a specific sound source is 100 won and the revenue distribution ratio is set to 50% for the creator, 30% for the AI model provider, and 20% for the platform operator, then 50 won, 30 won, and 20 won are distributed respectively.
[0049] FIG. 3 is a flowchart illustrating a sound source verification and registration process according to an embodiment of the present invention.
[0050] Referring to FIG. 3, first, when a user uploads an audio file through a user terminal (100) (S100), a format verification unit (321) of a service server (300) verifies the format and quality of the audio file (S110). If format verification fails, an error message is sent to the user and the process is terminated (S120).
[0051] If format verification is successful, the AI verification unit (322) determines whether the sound source was generated by AI (S130). That is, it determines whether the sound source was generated by AI. The AI verification unit (322) determines whether the sound source was generated by AI by analyzing the metadata of the sound source file or by executing an acoustic characteristic analysis algorithm. If it is determined that the sound source was not generated by AI, a rejection message is sent to the user and the process is terminated (S140).
[0052] When AI generation is confirmed, metadata such as AI model information, prompt information, and generation condition information is collected from the user (S150). The registration processing unit (323) stores the sound file and metadata in the AI sound database (310) (S160) and sends a registration completion message to the user (S170).
[0053] Meanwhile, a client application installed on a user terminal (100) provides a search screen that allows the user to search for sound sources with various AI-specific conditions.
[0054] The search screen includes an AI model selection menu, a prompt keyword input field, a BPM range setting slider, a genre selection menu, and a mood tag selection area. Users can combine these various filter conditions to precisely search for desired AI audio sources.
[0055] The search results screen displays a list of audio tracks matching the search criteria, providing the title, generation model icon, prompt preview, and play button for each track. When a user selects a specific track, a detailed information screen is displayed, which includes details of the AI model, the full prompt content, generation conditions, and creator information.
[0056] In addition, the client application provides a "similar prompt audio auto-play" feature, allowing users to automatically add other audio tracks generated with prompts similar to the currently playing track to the playback queue. This is implemented by the search and recommendation module analyzing prompt information and calculating keyword similarity.
[0057] As explained above, the AI-generated audio-only streaming platform system of the present invention provides a database structure exclusively for AI audio, a function to verify whether AI was generated, and search and recommendation functions based on AI audio-specific metadata, thereby enabling the provision of differentiated services for AI music enthusiasts that were impossible with existing general audio-centric platforms.
[0058] The foregoing description of the present invention is for illustrative purposes only, and those skilled in the art will understand that other specific forms can be easily modified without altering the technical spirit or essential features of the present invention. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, each component described as a single unit may be implemented in a distributed manner, and components described as distributed may likewise be implemented in a combined form.
[0059] The scope of the present invention is defined by the claims set forth below, and all modifications or variations derived from the meaning and scope of the claims and equivalent concepts thereof should be interpreted as being included within the scope of the present invention.
Claims
Claim 1 An AI audio database configured to register and store only audio generated by artificial intelligence; a format verification unit that verifies the format and quality of an audio file uploaded from a user terminal; an AI verification unit that determines whether the audio is generated by AI through metadata included in the audio file or an audio analysis algorithm; a registration processing unit that registers the audio file in the AI audio database when it is determined to be generated by AI; a streaming processing module that plays the audio stored in the AI audio database in a streaming manner; and a search and recommendation module that recommends audio similar to a specific audio selected by a user or entered in the form of a prompt, wherein the search and recommendation module performs recommendations by sorting audio in order of highest to lowest prompt similarity score calculated by the following mathematical formula. In the above mathematical formula, S(Pi, Pj) represents the prompt similarity score of source i and source j, where Pi and Pj represent the prompts of source i and source j, respectively, and in the above mathematical formula, the first term represents the weighted sum of TF-IDF (Term Frequency-Inverse Document Frequency) for n keywords included in the prompt, N is the number of keywords included in the prompts of source i and source j, Wk is a binary weight indicating whether the k-th keyword matches, and TF-IDF k A streaming platform system dedicated to AI-generated audio, wherein represents the TF-IDF score of the k-th keyword, the second term M(mi, mj) is a function representing the degree of agreement of the AI models that generated audio source i and audio source j, returning a value of 1 if they are the same model and 0 if they are different models, where mi and mj represent the identifiers of the AI models that generated audio source i and audio source j, respectively, the third term G(gi, gj) is a function representing the genre similarity of audio source i and audio source j, where it returns a value of 1 if the genres match completely, a value of 0.5 if they are similar genres, and a value of 0 if they are completely different genres, where gi and gj represent the genre information of audio source i and audio source j, where the similarity between multiple genres is stored in a database, and α, β, and γ are weight parameters. Claim 2 A streaming platform system dedicated to AI-generated music according to claim 1, further comprising a revenue distribution module that calculates cumulative usage by linking AI model information, prompt information, and creator identification information to the playback log of each music source, and distributes revenue to the music source creator and the AI model provider based on the cumulative usage. Claim 3 delete