System

The music streaming platform addresses the issue of fair compensation and recognition for singers by registering and managing AI-generated voice data, enabling efficient revenue distribution and promoting creativity.

JP2026025530APending Publication Date: 2026-02-16SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024128339
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-02
Publication Date
2026-02-16

AI Technical Summary

Technical Problem

The replication of singers' voices using generative AI technology has led to issues with copyright protection and revenue distribution, resulting in inadequate compensation and recognition for original singers and creators.

Method used

A music streaming platform that registers duplicated voice data, manages a music database, streams music data, records play counts, and calculates and distributes revenue using a generative AI model to ensure fair compensation for all parties involved.

Benefits of technology

Ensures fair revenue distribution and appropriate recognition for singers and creators, promoting creativity and innovation in the music industry by efficiently managing and utilizing AI-generated voice data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026025530000001_ABST
    Figure 2026025530000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for registering replicated voice data; means for searching a music database; means for streaming music data; means for recording the number of plays; and means for calculating and distributing revenue.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] With the development of modern generative AI technology, it has become easy to replicate singers' voices. However, this has led to issues regarding copyright protection and revenue distribution for songs that use AI-generated voices. As a result, the original singers and creators are not receiving adequate compensation or recognition. The objective of this invention is to solve this problem. [Means for solving the problem]

[0005] To solve this problem, the present invention provides the following means: A new music streaming platform is provided that allows singers to earn revenue even when their voices are duplicated using AI, through a system that includes a means for registering duplicated voice data, a means for searching a music database, a means for streaming music data, a means for recording the number of plays, and a means for calculating and distributing revenue. Specifically, the system includes a means for integrating an artist's voice data into an AI model and a means for providing an interface for users to search for and play songs, and as a whole, it is a system that achieves fair revenue distribution and appropriate evaluation.

[0006] 1. "Replica Voice Data" means voice data that has been created by replicating the voice of a specific singer using generative AI technology and stored in a digital format.

[0007] 2. A "music database" is a digital information management system in which information on a large number of songs is systematically stored.

[0008] 3. "Music Data" means the audio data of a musical composition stored in digital format and associated metadata.

[0009] 4. "Streaming" is a technology that delivers and plays music data in real time to a user's device via the Internet.

[0010] 5. "Plays" is numerical data that indicates the number of times a particular song has been played by a user.

[0011] 6. "Revenue" means monetary income generated from music streams and purchases.

[0012] 7. "Calculation" is the act of processing numerical data based on specific rules or algorithms to produce a result.

[0013] 8. "Distribution" is the act of allocating the profits earned among the parties based on a specific ratio.

[0014] 9. A "system" is a set of processes or devices with integrated functions that work together to achieve a specific purpose.

[0015] 10. "Interface" means the contact point or operating screen through which a user operates a system and receives responses from the system.

[0016] 11. “AI model” means a computational model trained to perform a specific task using artificial intelligence algorithms.

[0017] 12. "Artist" means an individual or group who creates a musical composition and owns the rights to their voice and composition.

[0018] 13. "Voice Data" means data that is a digital representation of recorded voice.

[0019] 14. "Search" is the act of searching a database to find specific information or data.

[0020] 15. "User" means an individual who uses the System to search for, play, or view revenue from Songs.

[0021] 16. "Song" means a musical composition, which is a collection of sounds including melody, lyrics, and musical accompaniment.

[0022] 17. "Real-time" means processing immediately, without delay.

[0023] 18. “Metadata” means information associated with a song, including the title, artist name, album name, genre, etc.

[0024] 19. "Specific rules or algorithms" means a specific set of mathematical or logical steps or instructions, which are standards or guidelines for performing calculations or data processing.

[0025] 20. "Digital format" means a form in which sound or data is digitally encoded so that it can be stored and reproduced. [Brief explanation of the drawings]

[0026] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0027] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0028] First, the terms used in the following description will be explained.

[0029] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0030] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0031] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0032] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0033] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0034] [First embodiment]

[0035] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0036] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0037] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0038] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0039] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0040] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0041] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0042] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0043] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0044] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0045] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0046] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0047] The present invention relates to a music streaming platform that uses singer voices replicated using generative AI technology. The system has the following components and functions:

[0048] 1. Registering duplicated voice data

[0049] Singers upload their own voice data via a terminal. The server receives this data and stores it in a database. This data is then integrated into an AI model and used as training data to be reflected in the generated music.

[0050] 2. Music database management

[0051] A music database is a system that stores song information in an organized manner. The server searches the music database based on a user's search query, retrieves the relevant songs, and provides them to the user. It also adds new songs and updates existing songs as needed.

[0052] 3. Streaming music data

[0053] When a user searches for and plays a song, the server quickly streams the corresponding music data to the device, allowing the user to enjoy music in real time. The device then plays the received streaming data, providing the user with a music experience.

[0054] 4. Recording Plays and Calculating Revenue

[0055] The server records the number of plays each time a song is played. Based on this play count data, revenue for each song is calculated and distributed appropriately. Specific rules and algorithms are used to calculate revenue, and the revenue is distributed fairly to each party involved (singer, lyricist, composer, etc.).

[0056] Specific examples

[0057] As an example, consider a scenario in which user A wants to listen to a new song that has been created using the voice of a particular singer B.

[0058] 1. Uploading voice data

[0059] Singer B uploads his / her voice data to the server via his / her device, which receives this data and integrates it into the AI ​​model.

[0060] 2. Music Generation

[0061] Composer C creates a new song and requests to use the voice of singer B. The server uses the AI ​​model to generate the new song in singer B's voice.

[0062] 3. Search and play songs

[0063] User A searches for new songs through a device. The device sends a query to the server, which searches the music database and finds the relevant new songs. The server streams the new songs and delivers them to the device. The device plays the received streaming data, allowing User A to listen to the new songs.

[0064] 4. Playback count and revenue sharing

[0065] The server records the number of plays each time User A plays a new song. Based on the play count data, the server calculates and distributes revenue to Singer B, Composer C, and other parties.

[0066] This system will ensure that singers who use AI-generated voices receive fair revenue, and that all creators involved receive appropriate recognition and compensation. This will improve fairness and transparency throughout the music industry, and promote greater creativity and innovation.

[0067] The processing flow will be explained below.

[0068] 1. Registering duplicated voice data

[0069] Processing Steps

[0070] Step 1:

[0071] The terminal accepts input from the user (singer). The singer selects their own voice data through the terminal and performs the upload operation.

[0072] Step 2:

[0073] The device sends voice data to the server, which receives and temporarily stores the data.

[0074] Step 3:

[0075] The server validates the voice data it receives, checking that the data format and quality meet the standards.

[0076] Step 4:

[0077] The server stores the verified voice data in a voice database, where it associates it with the singer's ID.

[0078] Step 5:

[0079] The server integrates the stored voice data into the AI ​​model, adding new voice data as training data for future music generation.

[0080] 2. Music database management

[0081] Processing Steps

[0082] Step 1:

[0083] The device accepts the user's request to upload music. The user (artist) selects their own music data through the device and performs the upload operation.

[0084] Step 2:

[0085] The device sends the music data to the server, which analyzes the received music data and temporarily stores it.

[0086] Step 3:

[0087] The server validates the song data, checking file integrity, format, and data completeness.

[0088] Step 4:

[0089] The server stores the verified music data in a music database, and associates the music data with artist information and metadata.

[0090] Step 5:

[0091] The server regularly backs up the music database to ensure data safety.

[0092] 3. Streaming music data

[0093] Processing Steps

[0094] Step 1:

[0095] The device accepts the user's song search request. The user enters a search query and clicks the search button.

[0096] Step 2:

[0097] The device sends a search query to the server, which receives it and searches its music database.

[0098] Step 3:

[0099] The server generates the search results, creates a list of matching songs, and returns it to the device.

[0100] Step 4:

[0101] The terminal displays the search results to the user.

[0102] Step 5:

[0103] The user selects the song they want to play and clicks the play button. The device sends a playback request to the server.

[0104] Step 6:

[0105] The server generates streaming data for the specified music piece and transmits it to the terminal.

[0106] Step 7:

[0107] The terminal plays the received streaming data and provides music to the user.

[0108] 4. Recording Plays and Calculating Revenue

[0109] Processing Steps

[0110] Step 1:

[0111] The server detects when a song starts playing. Each time a request to start playing is received, it records the play count.

[0112] Step 2:

[0113] The server collects the play count data and stores it in a database.

[0114] Step 3:

[0115] The server periodically aggregates the play count data and calculates revenue, which is then distributed based on specific rules and algorithms.

[0116] Step 4:

[0117] The server reflects the revenue data in the accounts of each party (singer, lyricist, composer, etc.).

[0118] Step 5:

[0119] Each party checks their own earnings through the terminal, which retrieves the earnings data from the server and displays it to the user.

[0120] This detailed processing flow allows the system to efficiently register the duplicated voice data, stream the song, record the number of plays, and calculate and distribute revenue.

[0121] Example 1

[0122] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0123] In conventional music streaming systems, music data is provided only using existing audio data, making it difficult to generate new audio data. Furthermore, there is no system in place for managing the appropriate use of copied audio data or for revenue distribution. Therefore, there is a need to improve fairness and transparency throughout the music industry and promote new creativity and innovation.

[0124] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0125] In this invention, the server includes a means for registering the replicated audio data, a means for managing a music database, a means for streaming the music data, a means for recording the number of plays, a means for calculating and distributing revenue, and a means for using a generative AI model to generate the replicated audio data. This not only allows users to enjoy high-quality audio data, but also enables fair revenue distribution based on the usage. Furthermore, utilizing the generative AI model enables the generation of new audio data and the provision of songs based on that data, promoting new creativity and innovation in the music industry.

[0126] "Duplicated voice data" is new voice data generated based on the original voice data using a generative AI model.

[0127] A "music database" is a database for systematically storing and managing song information and music files.

[0128] "Streaming" is a technology that distributes music data in real time over the Internet, allowing it to be played instantly on the user's device.

[0129] The "number of plays" is data indicating the number of times a particular song has been played by a user.

[0130] "Revenue" refers to income generated based on the playback of music data, which must be distributed appropriately to the parties involved.

[0131] A "generative AI model" is a model that uses machine learning algorithms to generate new data based on specific input data (in this case, voice data).

[0132] "Training data" refers to the dataset used to train a generative AI model, in this case the original audio data needed to generate the replicated audio data.

[0133] "Interface" refers to the operation screen and input means that a user uses to search for and play music.

[0134] The present invention relates to a music streaming platform that uses replicated audio data using generative AI technology. The system has the following components and functions:

[0135] 1. Registering duplicated audio data

[0136] Artist users upload their audio files using their devices. The devices then send the files to the server using HTTP POST requests. The server validates the received audio data and stores it in a secure database. The server then integrates the audio data into a generative AI model (e.g., WaveNet or Tacotron2) and uses it as training data to influence the generated music.

[0137] 2. Music database management

[0138] The server stores song information and music files in a music database in an organized manner. Song information includes title, artist, album, genre, and release year. The server also searches the database based on user search queries to extract and provide matching songs. New songs are added and existing songs are updated as needed.

[0139] 3. Streaming music data

[0140] When a user searches for and plays a song through a device, the server quickly streams the selected song to the device using HTTP or RTMP protocols. The device buffers the received streaming data and plays it using the built-in music player application, providing the user with a real-time music experience.

[0141] 4. Recording Plays and Calculating Revenue

[0142] The server records the number of plays as a log each time a song is played. This log includes information such as the song ID, playback start time, playback end time, and user ID. The log data is sent to a big data platform (e.g., Hadoop or Spark) for analysis. Based on the analysis, the server uses the play count data to calculate revenue and distributes it fairly to each relevant party (artist, lyricist, composer). Pre-set rules and algorithms are applied to calculate the revenue.

[0143] Specific examples

[0144] A specific example is given below: For example, consider a scenario in which user A listens to a new song created using the voice of a particular artist.

[0145] 1. Uploading voice data

[0146] Artists record their own voices and upload the audio data via their devices to a server, which receives the data and integrates it into a generative AI model.

[0147] 2. Music Generation

[0148] A composer creates a new song and requests that the artist's voice be used, and the server uses a generative AI model to generate the new song in the artist's voice.

[0149] 3. Search and play songs

[0150] User A searches for new songs through a device. The device sends a query to the server, which searches the music database to find the new songs. The server streams the new songs, and the device plays the received streaming data, allowing User A to listen to the new songs.

[0151] 4. Playback count and revenue sharing

[0152] The server records the number of plays each time User A plays a new song. Based on this data, the server calculates revenue and distributes it to the artists, composers, and other parties involved.

[0153] This system will enable efficient management and utilization of duplicated audio data using generative AI models, enabling fair revenue distribution, improving fairness and transparency across the music industry and encouraging greater creativity and innovation.

[0154] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0155] Step 1: Upload your voice data

[0156] The user, an artist, uses a device to record his or her own voice and uploads the audio data file to the server. Specifically, the user uses the device's recording function to create an audio file (e.g., WAV or MP3 format) and sends it to the server via an upload interface. The input is the audio data file, and the output is that file being saved on the server.

[0157] Step 2: Saving and integrating audio data

[0158] The server receives the uploaded audio data and verifies its integrity and quality. After verification is complete, it stores it in a database and performs preprocessing for incorporation into the generative AI model. Specifically, it performs noise removal and sampling rate adjustment. The input is the uploaded audio data file, and the output is the preprocessed audio data stored in the database.

[0159] Step 3: Store song information

[0160] The server adds new and updated song information to the music database. Song information includes title, artist, album, genre, and release year. Specifically, it creates database entries based on information provided by administrators and stakeholders. The input is song information and music files, and the output is that they are accurately stored in the database.

[0161] Step 4: Search for songs

[0162] The server uses the search query received from the user to search the music database and extract the corresponding songs. Specifically, it uses a full-text search engine (e.g., ElasticSearch) to quickly search for songs that match the query. The input is the search query, and the output is a list of songs as search results.

[0163] Step 5: Request a song to play

[0164] When a user uses a terminal to select a particular song, the terminal sends the selection to the server. The input is the user's song selection, and the output is a play request sent to the server.

[0165] Step 6: Stream your music

[0166] The server receives playback requests from users and streams the specified music. Specifically, it sends music files to terminals at a certain bit rate via HTTP or RTMP protocol. The input is the playback request, and the output is the streaming data.

[0167] Step 7: Playing a song

[0168] The device buffers the streaming data received from the server and plays it using the built-in music player application. The input is streaming data and the output is audio playback.

[0169] Step 8: Record play counts

[0170] The server records a playback event as a log each time a song is played. This log includes the song ID, playback start time, playback end time, user ID, etc. The input is the playback event information, and the output is the recorded log.

[0171] Step 9: Calculate and distribute revenue

[0172] The server analyzes the play count log and calculates revenue based on the number of plays of a specific song. This calculation is based on pre-defined rules and algorithms. Based on the analysis results, the revenue is distributed to the relevant parties (artist, lyricist, composer). The input is the play count log, and the output is the calculated revenue and its distribution information.

[0173] This enables the system to efficiently manage and utilize duplicated audio data using generative AI models, enabling fair revenue distribution.

[0174] (Application example 1)

[0175] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0176] On today's music distribution platforms, it is difficult for users to create and enjoy songs using the voice of a specific artist, and there is a demand for proper management of the number of plays and distribution of revenue for the created songs. Therefore, it is necessary to create songs using copied voices and provide them to users while realizing fair and transparent revenue distribution.

[0177] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0178] In this invention, the server includes a means for registering the copied audio data, a means for searching a music database, a means for streaming the music data, a means for users to generate music based on specific requests using a generative AI model, a means for recording the number of plays, and a means for calculating and distributing revenue. This allows users to freely generate new music using artists' voices and enjoy them in real time, while also enabling fair and transparent revenue distribution to artists and related parties.

[0179] "Duplicated audio data" refers to audio data that has been duplicated or generated using artificial intelligence technology from data recorded from the original artist's voice.

[0180] A "music database" is a collection of digital data that organizes and stores information about songs, and can be accessed by users through searches.

[0181] "Music Data" means data, including audio and metadata, of a musical composition stored in digital format.

[0182] "Streaming" is a method of playing music data in real time over the Internet.

[0183] "Plays" is a record of the number of times a particular song has been played by a user.

[0184] "Revenue sharing" is the process of fairly distributing revenues earned as a result of the playback of a generated song among the parties involved.

[0185] A "generative AI model" is an algorithm or system that uses artificial intelligence technology to generate new audio data or music.

[0186] The "user interface" is the part of the software that provides an operation screen for the user to search for and play music.

[0187] This invention provides a music streaming platform that uses duplicated audio using generative AI technology. The system includes functions such as registration of duplicated audio data, management of music data, streaming, recording of play counts, and revenue sharing.

[0188] System Configuration

[0189] 1. Server:

[0190] Audio data registration:

[0191] The server receives the voice data uploaded via the device and stores it in a database. Voice providers provide their own voice, and the server integrates this data into the AI ​​model.

[0192] Music database management:

[0193] The server manages a database that systematically stores song information, receives search queries from users, searches for and provides matching songs, and also adds and updates new songs.

[0194] Streaming:

[0195] The server streams the music selected by the user in real time and delivers it to the device, allowing the user to enjoy the music instantly.

[0196] Playback Tracking and Revenue Sharing:

[0197] The server records the number of times a song is played and calculates and distributes revenue based on that number, using specific rules and algorithms to ensure fair distribution among all parties involved.

[0198] Use of generative AI models:

[0199] It receives prompts based on the user's request and uses a generative AI model to generate new music, which is then instantly added to a database and made available for streaming.

[0200] 2. Device (smartphone):

[0201] Upload audio data:

[0202] The terminal provides an interface for the user to record the voice of the voice provider and upload it to the server.

[0203] Search and play songs:

[0204] The device provides an interface for users to search for and play music. Users can search for music within the app and enjoy music by pressing the play button.

[0205] Music generation using prompts:

[0206] The device provides an interface for users to input specific prompts to generate new music, which are then sent to a server and processed by a generative AI model.

[0207] Hardware and software used

[0208] The hardware includes servers (cloud-based servers or dedicated servers) and user devices (smartphones or tablets). As for software, the server side runs a web application based on Flask and a generative AI model using TensorFlow. The client side runs a smartphone application that provides the user interface.

[0209] Specific examples

[0210] For example, if a user wants to generate a new pop song using the voice of a particular singer, they might enter the prompt text as follows:

[0211] "Generate love songs with pop rhythms using singer's voices"

[0212] By passing this prompt to a generative AI model, a new pop song using the specified voice is generated and instantly added to the database, where users can then search for and stream the new song.

[0213] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0214] Step 1:

[0215] The user uses a device to record the voice of the voice provider and uploads the voice data from the device to the server. The server receives this voice data and stores it in a database as training data for the generative AI model. The input is the voice data, and the output is the voice data stored in the database.

[0216] Step 2:

[0217] The server integrates the received voice data into the generative AI model. The data processing performed here involves converting the voice data into an appropriate format and training the model so that it can generate new voices. The input is the voice data, and the output is updating the generative AI model.

[0218] Step 3:

[0219] A user inputs a search query for a song through a terminal. The terminal sends the query to a server, which then searches a music database to find the relevant song. The input is a search query, and the output is a list of songs as search results.

[0220] Step 4:

[0221] The user selects a specific song from the search results and instructs it to be played. The server streams the corresponding music data and delivers it to the device. The device receives this streaming data and plays it for the user. The input is the song ID, and the output is the song data played in real time.

[0222] Step 5:

[0223] The server records the user's playback actions and updates the database with the number of times the song has been played. This process takes the user ID and song ID as input, increments the number of times the song has been played, and updates the database. The input is the user ID and song ID, and the output is the updated number of times the song has been played.

[0224] Step 6:

[0225] The user uses a terminal to input a specific prompt sentence and request the generation of a new song. The server receives this prompt sentence and passes it to the generative AI model, which then generates a new song. The input is the prompt sentence, and the output is the generated song data. As a concrete example, we use a prompt sentence such as "Generate a love song with a pop rhythm and a singer's voice."

[0226] Step 7:

[0227] The server registers the generated new music in a database, allowing users to search and play it. The input is the generated music data, and the output is the new music registered in the database.

[0228] Step 8:

[0229] The server calculates and distributes revenue. It calculates revenue based on the number of views using specific rules and algorithms and distributes it to the parties involved. The input is the number of views, and the output is the calculated revenue and distribution results.

[0230] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0231] The present invention combines an emotion engine with a music streaming platform that uses singer voices replicated using generative AI technology. The system has the following components and functions:

[0232] 1. Registering duplicated voice data

[0233] Singers upload their own voice data via a terminal. The server receives this data and stores it in a database. This data is then integrated into an AI model and used as training data to be reflected in the generated music.

[0234] 2. Music database management

[0235] A music database is a system that stores song information in an organized manner. The server searches the music database based on a user's search query, retrieves the relevant songs, and provides them to the user. It also adds new songs and updates existing songs as needed.

[0236] 3. Streaming music data

[0237] When a user searches for and plays a song, the server quickly streams the corresponding music data to the device, allowing the user to enjoy music in real time. The device then plays the received streaming data, providing the user with a music experience.

[0238] 4. Recording Plays and Calculating Revenue

[0239] The server records the number of plays each time a song is played. Based on this play count data, revenue for each song is calculated and distributed appropriately. Specific rules and algorithms are used to calculate revenue, and the revenue is distributed fairly to each party involved (singer, lyricist, composer, etc.).

[0240] 5. Use of Emotion Engine

[0241] The system is equipped with an emotion engine for recognizing a user's emotions, which analyzes the user's voice, visual data, text input, and other physiological data to identify the user's current emotional state.

[0242] 1. Collecting Emotional Data

[0243] The device collects emotional data while the user listens to the music being played, including facial expression recognition using a camera, voice analysis using a microphone, and physiological data acquisition using sensors.

[0244] 2. Emotion Analysis

[0245] The server analyzes the collected emotion data to identify the user's emotional state. The emotion engine uses machine learning algorithms to identify the user's emotion based on the collected data.

[0246] 3. Song Recommendations

[0247] The server recommends songs to the user based on the identified emotional state, for example, if the user is in a sad state, the emotion engine selects and suggests comforting songs to the user.

[0248] 4. Emotion history storage and analysis

[0249] The server stores the user's emotional history in a database and performs long-term analysis, which allows it to understand the user's emotional patterns and provide a more personalized music experience.

[0250] Specific examples

[0251] 1. Uploading voice data and integrating AI models

[0252] The terminal accepts input for uploading singer's voice data.

[0253] The server receives the voice data, stores it, and integrates it into an AI model, which then uses it to generate new songs.

[0254] 2. Search and play songs

[0255] The user searches for a new song and clicks the play button.

[0256] The server searches for the music, generates streaming data, and delivers it to the device.

[0257] The terminal receives the streaming data and plays it back to the user.

[0258] 3. Recording of play counts and revenue sharing

[0259] The server records the number of times the song is played and calculates the revenue.

[0260] Profits are distributed among the parties according to certain rules.

[0261] 4. Utilizing the Emotion Engine

[0262] While the user is listening to music, the device collects emotion data.

[0263] The server uses an emotion engine to analyze the user's emotions and recommend appropriate songs.

[0264] It analyzes emotional history over the long term to provide users with the optimal music experience.

[0265] This system ensures that singers who use AI-replicated voices receive fair revenue, and that all creators involved receive appropriate recognition and compensation. Furthermore, the emotional engine can provide users with a more personalized music experience. This will improve fairness and transparency throughout the music industry, promoting greater creativity and innovation.

[0266] The processing flow will be explained below.

[0267] 1. Registering duplicated voice data

[0268] Processing Steps

[0269] Step 1:

[0270] The terminal accepts input from the user (singer). The singer selects his / her own voice data via the terminal and performs the upload operation.

[0271] Step 2:

[0272] The device sends voice data to the server, which receives and temporarily stores the data.

[0273] Step 3:

[0274] The server validates the voice data it receives, checking that the format and quality of the voice data meets standards.

[0275] Step 4:

[0276] The server stores the verified voice data in a voice database, and when it is saved, it associates the voice data with the singer's ID.

[0277] Step 5:

[0278] The server integrates the stored voice data into the AI ​​model, adding new voice data as training data for future music generation.

[0279] 2. Music database management

[0280] Processing Steps

[0281] Step 1:

[0282] The device accepts a song upload request from the user (artist). The artist selects their own song data through the device and performs the upload operation.

[0283] Step 2:

[0284] The device sends the music data to the server, which analyzes the received music data and temporarily stores it.

[0285] Step 3:

[0286] The server validates the song data, checking its integrity, format, and quality.

[0287] Step 4:

[0288] The server stores the verified music data in a music database, and associates the music data with artist information and metadata.

[0289] Step 5:

[0290] The server regularly backs up the music database to ensure data safety and availability.

[0291] 3. Streaming music data

[0292] Processing Steps

[0293] Step 1:

[0294] The device accepts the user's song search request. The user enters a search query and clicks the search button.

[0295] Step 2:

[0296] The device sends a search query to the server, which receives it and searches its music database.

[0297] Step 3:

[0298] The server generates the search results, creating a list of matching songs and sending it back to the device.

[0299] Step 4:

[0300] The terminal displays the search results to the user.

[0301] Step 5:

[0302] The user selects the song they want to play and clicks the play button. The device sends a playback request to the server.

[0303] Step 6:

[0304] The server generates streaming data for the specified music piece and transmits it to the terminal.

[0305] Step 7:

[0306] The terminal plays the received streaming data and provides music to the user.

[0307] 4. Recording Plays and Calculating Revenue

[0308] Processing Steps

[0309] Step 1:

[0310] The server detects when a song starts playing. Each time a request to start playing is received, it records the play count.

[0311] Step 2:

[0312] The server collects play count data and stores it in a database.

[0313] Step 3:

[0314] The server periodically aggregates the play count data and calculates revenue, which is then distributed based on rules and algorithms.

[0315] Step 4:

[0316] The server reflects the revenue data in the accounts of each party (singer, lyricist, composer, etc.).

[0317] Step 5:

[0318] Each party checks their own earnings through the terminal, which retrieves the earnings data from the server and displays it to the user.

[0319] 5. Use of Emotion Engine

[0320] Processing Steps

[0321] Step 1:

[0322] The device collects the user's emotional data. While the user is playing music, it uses a camera to recognize facial expressions, a microphone to analyze voices, and sensors to acquire physiological data.

[0323] Step 2:

[0324] The device sends the collected emotion data to a server, which receives it and temporarily stores it.

[0325] Step 3:

[0326] The server analyzes the emotion data, and the emotion engine uses machine learning algorithms to identify the user's emotion based on the collected data.

[0327] Step 4:

[0328] The server recommends songs based on the analyzed emotional data, suggesting songs that best suit the user's state of mind.

[0329] Step 5:

[0330] The server stores the user's emotional history in a database and performs long-term analysis, analyzing the user's emotional patterns to provide a more personalized music experience.

[0331] This detailed processing flow allows the system to efficiently register the duplicated voice data, stream the songs, record the number of plays, calculate and distribute revenue, and even recommend songs using an emotion engine.

[0332] Example 2

[0333] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0334] While conventional music streaming systems can play songs and distribute revenue, they lack the functionality to recommend songs based on user emotions or replicate artists' voices using AI. This makes it difficult to provide a music experience tailored to each user's emotions, and also poses the problem of insufficient fairness in revenue distribution for songs that use replicated voices.

[0335] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0336] In this invention, the server includes means for registering the duplicated voice data, means for searching a music database, means for streaming the music data, means for recording the number of plays, means for calculating and distributing revenue, means for collecting user emotion data, and means for analyzing the collected emotion data and recommending songs to the user. This makes it possible to provide a personalized music experience according to the user's emotion while maintaining fairness in the distribution of revenue from songs based on the duplicated voice using AI.

[0337] "Replica voice data" refers to audio information that has been recorded from an artist's voice and reproduced using a generative AI model.

[0338] A "music database" is an information system that efficiently manages metadata and audio files related to music, allowing them to be searched and stored.

[0339] "Streaming" is a technology that distributes and plays music data in real time to a user's terminal via the Internet.

[0340] "Plays" is an indicator that records the number of times a particular song has been played by a user.

[0341] "Revenue sharing" is the process of fairly distributing revenue from music content to stakeholders based on data such as the number of plays.

[0342] "Emotion data" is data that indicates the user's current emotional state, collected from the user's facial expressions, voice, physiological data, and the like.

[0343] A "generative AI model" is a collection of algorithms that perform specific tasks or generate results based on large amounts of data, and in this case refers specifically to models used to reproduce an artist's voice.

[0344] "Recommendation" refers to the act of selecting the most suitable music piece based on the analyzed emotional data of the user and suggesting it to the user.

[0345] This invention provides a personalized music experience for users by integrating an emotion engine into a music streaming system that uses replicated voice data. The main hardware used includes a database server, a terminal (including a camera, microphone, and sensors) that collects emotion data, and the user's device. The software used includes a generative AI model, an emotion analysis algorithm, streaming server software, and a user interface.

[0346] Singers record their own voice data on the device and upload the data to the server. The server receives this voice data, stores it in a database, and integrates it into the generative AI model. For example, a singer can record themselves singing specific lyrics on the device and send it to the server via a dedicated app. The server receives the voice data and stores it as replicated voice data. When integrated into the generative AI model, it is used as new voice data added to the existing voice dataset.

[0347] The server then manages a music database, which stores music metadata (such as title, artist name, and album name) and audio files in an organized manner. When a user searches for a song, the server searches the database based on the search query and returns the relevant songs. For example, if a user enters a prompt such as "Add a new song to the server," the server adds the new song's metadata and audio files to the database.

[0348] In music streaming, when a user searches for a specific song and clicks the play button, the server prepares the song data in real time and delivers it to the user's device as streaming data. The device receives the streaming data and immediately starts playing it. For example, the server can generate streaming data and provide it to the user by using a prompt such as "Please search for a specific song and play it."

[0349] Regarding the recording of play counts and revenue calculation, the server records the number of plays in real time each time each song is played. Based on this data, revenue is calculated and fairly distributed to the parties according to specific rules. The revenue calculation takes into account the number of plays, song information, related contract information, etc. For example, upon the prompt "Please calculate and distribute revenue based on the number of song plays," the server will perform the revenue calculation and save the result in the database.

[0350] When using an emotion engine, the user's device uses a camera, microphone, and sensors to collect the user's facial, voice, and physiological data. The collected emotion data is sent to a server and analyzed using an emotion analysis algorithm. Based on the analysis results, the server recommends music that best suits the user's emotional state. For example, a prompt such as "Analyze the user's emotion data and recommend music that matches that emotion" allows the server to analyze the emotion data and recommend appropriate music.

[0351] As described above, this invention provides a music streaming system that uses replicated voice data to recommend songs based on the user's emotions, enabling users to enjoy a personalized music experience. Furthermore, fairness in revenue distribution is maintained, making this a system that benefits all parties involved.

[0352] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0353] Step 1:

[0354] The terminal provides a user interface for singers to record their voice data using an input device. The user presses the record button, and after recording is completed, clicks the upload button. The input is the recorded voice data, and the output is an upload request to the server.

[0355] Step 2:

[0356] The server receives the voice data sent from the terminal and performs pre-processing such as data formatting and format conversion.The server then stores the formatted voice data in a voice database.The input is the uploaded voice data, and the output is the voice data stored in the database.

[0357] Step 3:

[0358] The server adds the stored voice data to a training dataset for integration into a generative AI model. This training dataset is used to generate replicated voices. The input is the voice data stored in the database, and the output is the training data integrated into the generative AI model.

[0359] Step 4:

[0360] A user enters a specific song title into the search bar of their device and presses the search button. A song search request is sent from the device to the server. The input is the search query entered by the user, and the output is the song search request.

[0361] Step 5:

[0362] The server searches the music database based on the user's search query, retrieves the corresponding song data, and returns the retrieved song data to the device. The input is the search query, and the output is the corresponding song data.

[0363] Step 6:

[0364] The device receives the music data sent from the server and launches the streaming player. When the user clicks the play button, the music starts playing. The input is the music data from the server, and the output is the music being played.

[0365] Step 7:

[0366] The server records the number of times a song is played in real time. This play count data is later used to calculate revenue. The input is the song play event, and the output is the recorded play count data.

[0367] Step 8:

[0368] The server uses a revenue calculation algorithm to calculate revenue based on the number of times a song is played and distribute it to the parties involved. The input is the recorded number of times a song is played and the revenue distribution rule, and the output is the revenue data distributed to each party.

[0369] Step 9:

[0370] The device uses sensors, cameras, and microphones to collect emotional data while the user is listening to music. The collected emotional data is sent to a server. The input is the user's physiological response, and the output is the collected emotional data.

[0371] Step 10:

[0372] The server analyzes the collected emotional data using an emotion analysis algorithm to identify the user's current emotional state, where the input is the collected emotional data and the output is the analyzed emotional state.

[0373] Step 11:

[0374] The server recommends the best songs to the user based on the analyzed emotional state. The input is the analyzed emotional state, and the output is a list of recommended songs.

[0375] Step 12:

[0376] The server stores the user's emotion history in a database and uses it to analyze long-term emotion patterns. The input is the analyzed emotion data, and the output is the emotion history stored in the database.

[0377] (Application example 2)

[0378] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0379] Traditional music streaming services lack personalized song recommendations based on the user's emotional state and new music experiences using AI-generated singers. This makes it difficult to respond to diverse user emotions and preferences, and has made it difficult to improve user satisfaction and provide novel music experiences. Another problem is the lack of transparency regarding revenue distribution to music creators.

[0380] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0381] In this invention, the server includes means for registering copied voice data, means for searching a music database, means for streaming music data, means for recording the number of plays, means for calculating and distributing revenue, means for collecting user emotion data, means for analyzing the emotion data to identify the user's emotional state, means for recommending songs based on the identified emotional state, and means for storing and analyzing the emotion history over time. This enables personalized song recommendations based on the user's emotional state, improving user satisfaction and providing a novel music experience. It also enables transparent and fair revenue distribution to music creators.

[0382] "Duplicated voice data" refers to voice data that reproduces a singer's voice using generative AI technology.

[0383] A "music database" is a system that systematically stores information about multiple songs.

[0384] "Streaming" is a technology that distributes stored music data in real time over the Internet.

[0385] "Play count" is data that records the number of times a particular song has been played by a user.

[0386] "Revenue sharing" is the process of appropriately distributing revenue earned based on the number of times a song is played among the parties involved.

[0387] "Emotional data" is data that indicates the user's current emotional state, and is mainly collected from voice, visual information, physiological data, and the like.

[0388] "Emotion analysis" is the process of identifying a user's emotional state based on collected emotional data.

[0389] "Music recommendation" is the act of suggesting the most suitable music based on the analyzed emotional state of the user.

[0390] "Emotion history" is data that records changes in the user's emotional state over a long period of time.

[0391] An "AI model" is a data model that is optimized for a specific task using machine learning algorithms.

[0392] The present invention relates to a music streaming platform using replicated voice data and an emotion engine, and can be implemented by applying the following elements and means:

[0393] 1. System Configuration

[0394] Hardware

[0395] Camera: Captures the user's facial expressions.

[0396] Microphone: Collects the user's voice and physiological data.

[0397] Smartphone: Used as an application execution environment.

[0398] software

[0399] OpenCV: Used as an image processing library to acquire camera input and perform facial expression recognition.

[0400] Librosa: A library for music analysis. It extracts features such as tempo from music data.

[0401] TensorFlow / Keras: Load and run the emotion engine model.

[0402] Requests: Used to send HTTP requests to the music recommendation service.

[0403] 2. System Operation

[0404] The server implements various functions using the following means:

[0405] Registering duplicated voice data

[0406] Singers upload their own voice data via a terminal. The server receives this data and stores it in a database. The server then integrates this data into an AI model and generates new songs using generative AI technology.

[0407] Music database management

[0408] The server searches the music database based on the user's search query, retrieves the relevant songs, and provides them to the user. It also adds new songs and updates existing songs.

[0409] Streaming music data

[0410] When a user searches for and plays a song, the server quickly streams the corresponding music data to the device, allowing the user to enjoy music in real time.

[0411] Recording views and calculating revenue

[0412] The server records the number of plays each time a song is played, and calculates revenue based on this data, which is then distributed equally among all parties involved (singers, lyricists, composers, etc.).

[0413] Use of emotion engine

[0414] The server collects users' emotional data and analyzes it using an emotion engine. It recommends songs based on the user's emotional state and analyzes their emotional history over time to provide a more personalized music experience.

[0415] 3. Specific Examples

[0416] Emotion data collection and analysis

[0417] While the user is listening to music, the device's camera and microphone work together to capture the user's facial expressions and vocal characteristics, which are then analyzed by the emotion engine to determine the user's emotional state.

[0418] Song recommendations

[0419] The server recommends songs based on the identified emotional state, for example, if the user is in a sad state, the emotion engine selects and recommends comforting songs.

[0420] Example: Using a prompt statement

[0421] The implicit knowledge is that when the user is in a sad state, the prompt sentence is:

[0422] "The user is in a sad state, so please recommend some songs that are slow and comforting."

[0423] Long-term emotion history analysis

[0424] The server stores the user's emotional history in a database and analyzes it over the long term, allowing it to understand the user's emotional patterns and reflect them in future song recommendations.

[0425] By combining these elements and methods, it becomes possible to provide a personalized music streaming service that responds to the user's emotional state, thereby improving user satisfaction and providing a novel music experience.

[0426] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0427] Step 1:

[0428] The user launches a music application.

[0429] Input: User launches application.

[0430] Output: The application home screen is displayed.

[0431] Step 2:

[0432] The user uploads the singer's voice data.

[0433] Input: User selection and upload of voice data.

[0434] Output: Duplicate voice data stored on the server.

[0435] Specific behavior:

[0436] The terminal receives selected voice data from the user.

[0437] The received data is sent to the server and stored in a database.

[0438] The voice data is used as training data to integrate into generative AI models.

[0439] Step 3:

[0440] The server searches the music database.

[0441] Input: A search query by the user.

[0442] Output: A list of songs as search results.

[0443] Specific behavior:

[0444] The user types a query into the search bar and sends it to the server.

[0445] The server searches the music database based on the received query and retrieves the corresponding songs.

[0446] Step 4:

[0447] The user plays a song.

[0448] Input: User clicks the play button.

[0449] Output: The streaming data to be played.

[0450] Specific behavior:

[0451] The server delivers streaming data of the selected music to the terminal.

[0452] The terminal plays the received streaming data, providing the user with a musical experience.

[0453] Step 5:

[0454] The device collects the user's emotional data.

[0455] Input: User's facial and voice data.

[0456] Output: Collected emotion data.

[0457] Specific behavior:

[0458] The device's camera and microphone are activated to capture the user's facial expressions and voice.

[0459] The acquired data is preprocessed and sent to the server.

[0460] Step 6:

[0461] The server analyzes the emotional data to determine the user's emotional state.

[0462] Input: Collected emotion data.

[0463] Output: The identified emotional state of the user.

[0464] Specific behavior:

[0465] The server analyzes the emotion data using machine learning algorithms.

[0466] An emotion engine identifies the user's current emotional state.

[0467] Step 7:

[0468] The server recommends songs based on the identified emotional state.

[0469] Input: The identified emotional state of the user.

[0470] Output: A list of recommended songs.

[0471] Specific behavior:

[0472] The server searches a music database for songs that correspond to the identified emotional state.

[0473] A list of recommended songs is presented to the user.

[0474] Step 8:

[0475] The server records the number of plays and calculates the revenue.

[0476] Input: Song playback event.

[0477] Output: Calculated revenue and distribution results.

[0478] Specific behavior:

[0479] The server records the number of times a song is played in a database each time it is played.

[0480] Revenue is calculated based on the number of plays and distributed to each party.

[0481] Step 9:

[0482] The server stores emotion history and analyzes it over the long term.

[0483] Input: Identified emotional state.

[0484] Output: The stored emotion history and its analysis results.

[0485] Specific behavior:

[0486] The server stores the history of the user's emotional state in a database.

[0487] Conduct long-term analysis to understand user sentiment patterns.

[0488] Through the above steps, a music streaming experience based on the user's emotional state can be provided.

[0489] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0490] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0491] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0492] [Second embodiment]

[0493] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0494] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0495] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0496] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0497] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0498] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0499] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0500] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0501] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0502] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0503] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0504] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0505] The present invention relates to a music streaming platform that uses singer voices replicated using generative AI technology. The system has the following components and functions:

[0506] 1. Registering duplicated voice data

[0507] Singers upload their own voice data via a terminal. The server receives this data and stores it in a database. This data is then integrated into an AI model and used as training data to be reflected in the generated music.

[0508] 2. Music database management

[0509] A music database is a system that stores song information in an organized manner. The server searches the music database based on a user's search query, retrieves the relevant songs, and provides them to the user. It also adds new songs and updates existing songs as needed.

[0510] 3. Streaming music data

[0511] When a user searches for and plays a song, the server quickly streams the corresponding music data to the device, allowing the user to enjoy music in real time. The device then plays the received streaming data, providing the user with a music experience.

[0512] 4. Recording Plays and Calculating Revenue

[0513] The server records the number of plays each time a song is played. Based on this play count data, revenue for each song is calculated and distributed appropriately. Specific rules and algorithms are used to calculate revenue, and the revenue is distributed fairly to each party involved (singer, lyricist, composer, etc.).

[0514] Specific examples

[0515] As an example, consider a scenario in which user A wants to listen to a new song that has been created using the voice of a particular singer B.

[0516] 1. Uploading voice data

[0517] Singer B uploads his / her voice data to the server via his / her device, which receives this data and integrates it into the AI ​​model.

[0518] 2. Music Generation

[0519] Composer C creates a new song and requests to use the voice of singer B. The server uses the AI ​​model to generate the new song in singer B's voice.

[0520] 3. Search and play songs

[0521] User A searches for new songs through a device. The device sends a query to the server, which searches the music database and finds the relevant new songs. The server streams the new songs and delivers them to the device. The device plays the received streaming data, allowing User A to listen to the new songs.

[0522] 4. Playback count and revenue sharing

[0523] The server records the number of plays each time User A plays a new song. Based on the play count data, the server calculates and distributes revenue to Singer B, Composer C, and other parties.

[0524] This system will ensure that singers who use AI-generated voices receive fair revenue, and that all creators involved receive appropriate recognition and compensation. This will improve fairness and transparency throughout the music industry, and promote greater creativity and innovation.

[0525] The processing flow will be explained below.

[0526] 1. Registering duplicated voice data

[0527] Processing Steps

[0528] Step 1:

[0529] The terminal accepts input from the user (singer). The singer selects their own voice data through the terminal and performs the upload operation.

[0530] Step 2:

[0531] The device sends voice data to the server, which receives and temporarily stores the data.

[0532] Step 3:

[0533] The server validates the voice data it receives, checking that the data format and quality meet the standards.

[0534] Step 4:

[0535] The server stores the verified voice data in a voice database, where it associates it with the singer's ID.

[0536] Step 5:

[0537] The server integrates the stored voice data into the AI ​​model, adding new voice data as training data for future music generation.

[0538] 2. Music database management

[0539] Processing Steps

[0540] Step 1:

[0541] The device accepts the user's request to upload music. The user (artist) selects their own music data through the device and performs the upload operation.

[0542] Step 2:

[0543] The device sends the music data to the server, which analyzes the received music data and temporarily stores it.

[0544] Step 3:

[0545] The server validates the song data, checking file integrity, format, and data completeness.

[0546] Step 4:

[0547] The server stores the verified music data in a music database, and associates the music data with artist information and metadata.

[0548] Step 5:

[0549] The server regularly backs up the music database to ensure data safety.

[0550] 3. Streaming music data

[0551] Processing Steps

[0552] Step 1:

[0553] The device accepts the user's song search request. The user enters a search query and clicks the search button.

[0554] Step 2:

[0555] The device sends a search query to the server, which receives it and searches its music database.

[0556] Step 3:

[0557] The server generates the search results, creates a list of matching songs, and returns it to the device.

[0558] Step 4:

[0559] The terminal displays the search results to the user.

[0560] Step 5:

[0561] The user selects the song they want to play and clicks the play button. The device sends a playback request to the server.

[0562] Step 6:

[0563] The server generates streaming data for the specified music piece and transmits it to the terminal.

[0564] Step 7:

[0565] The terminal plays the received streaming data and provides music to the user.

[0566] 4. Recording Plays and Calculating Revenue

[0567] Processing Steps

[0568] Step 1:

[0569] The server detects when a song starts playing. Each time a request to start playing is received, it records the play count.

[0570] Step 2:

[0571] The server collects the play count data and stores it in a database.

[0572] Step 3:

[0573] The server periodically aggregates the play count data and calculates revenue, which is then distributed based on specific rules and algorithms.

[0574] Step 4:

[0575] The server reflects the revenue data in the accounts of each party (singer, lyricist, composer, etc.).

[0576] Step 5:

[0577] Each party checks their own earnings through the terminal, which retrieves the earnings data from the server and displays it to the user.

[0578] This detailed processing flow allows the system to efficiently register the duplicated voice data, stream the song, record the number of plays, and calculate and distribute revenue.

[0579] Example 1

[0580] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0581] In conventional music streaming systems, music data is provided only using existing audio data, making it difficult to generate new audio data. Furthermore, there is no system in place for managing the appropriate use of copied audio data or for revenue distribution. Therefore, there is a need to improve fairness and transparency throughout the music industry and promote new creativity and innovation.

[0582] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0583] In this invention, the server includes a means for registering the replicated audio data, a means for managing a music database, a means for streaming the music data, a means for recording the number of plays, a means for calculating and distributing revenue, and a means for using a generative AI model to generate the replicated audio data. This not only allows users to enjoy high-quality audio data, but also enables fair revenue distribution based on the usage. Furthermore, utilizing the generative AI model enables the generation of new audio data and the provision of songs based on that data, promoting new creativity and innovation in the music industry.

[0584] "Duplicated voice data" is new voice data generated based on the original voice data using a generative AI model.

[0585] A "music database" is a database for systematically storing and managing song information and music files.

[0586] "Streaming" is a technology that distributes music data in real time over the Internet, allowing it to be played instantly on the user's device.

[0587] The "number of plays" is data indicating the number of times a particular song has been played by a user.

[0588] "Revenue" refers to income generated based on the playback of music data, which must be distributed appropriately to the parties involved.

[0589] A "generative AI model" is a model that uses machine learning algorithms to generate new data based on specific input data (in this case, voice data).

[0590] "Training data" refers to the dataset used to train a generative AI model, in this case the original audio data needed to generate the replicated audio data.

[0591] "Interface" refers to the operation screen and input means that a user uses to search for and play music.

[0592] The present invention relates to a music streaming platform that uses replicated audio data using generative AI technology. The system has the following components and functions:

[0593] 1. Registering duplicated audio data

[0594] Artist users upload their audio files using their devices. The devices then send the files to the server using HTTP POST requests. The server validates the received audio data and stores it in a secure database. The server then integrates the audio data into a generative AI model (e.g., WaveNet or Tacotron2) and uses it as training data to influence the generated music.

[0595] 2. Music database management

[0596] The server stores song information and music files in a music database in an organized manner. Song information includes title, artist, album, genre, and release year. The server also searches the database based on user search queries to extract and provide matching songs. New songs are added and existing songs are updated as needed.

[0597] 3. Streaming music data

[0598] When a user searches for and plays a song through a device, the server quickly streams the selected song to the device using HTTP or RTMP protocols. The device buffers the received streaming data and plays it using the built-in music player application, providing the user with a real-time music experience.

[0599] 4. Recording Plays and Calculating Revenue

[0600] The server records the number of plays as a log each time a song is played. This log includes information such as the song ID, playback start time, playback end time, and user ID. The log data is sent to a big data platform (e.g., Hadoop or Spark) for analysis. Based on the analysis, the server uses the play count data to calculate revenue and distributes it fairly to each relevant party (artist, lyricist, composer). Pre-set rules and algorithms are applied to calculate the revenue.

[0601] Specific examples

[0602] A specific example is given below: For example, consider a scenario in which user A listens to a new song created using the voice of a particular artist.

[0603] 1. Uploading voice data

[0604] Artists record their own voices and upload the audio data via their devices to a server, which receives the data and integrates it into a generative AI model.

[0605] 2. Music Generation

[0606] A composer creates a new song and requests that the artist's voice be used, and the server uses a generative AI model to generate the new song in the artist's voice.

[0607] 3. Search and play songs

[0608] User A searches for new songs through a device. The device sends a query to the server, which searches the music database to find the new songs. The server streams the new songs, and the device plays the received streaming data, allowing User A to listen to the new songs.

[0609] 4. Playback count and revenue sharing

[0610] The server records the number of plays each time User A plays a new song. Based on this data, the server calculates revenue and distributes it to the artists, composers, and other parties involved.

[0611] This system will enable efficient management and utilization of duplicated audio data using generative AI models, enabling fair revenue distribution, improving fairness and transparency across the music industry and encouraging greater creativity and innovation.

[0612] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0613] Step 1: Upload your voice data

[0614] The user, an artist, uses a device to record his or her own voice and uploads the audio data file to the server. Specifically, the user uses the device's recording function to create an audio file (e.g., WAV or MP3 format) and sends it to the server via an upload interface. The input is the audio data file, and the output is that file being saved on the server.

[0615] Step 2: Saving and integrating audio data

[0616] The server receives the uploaded audio data and verifies its integrity and quality. After verification is complete, it stores it in a database and performs preprocessing for incorporation into the generative AI model. Specifically, it performs noise removal and sampling rate adjustment. The input is the uploaded audio data file, and the output is the preprocessed audio data stored in the database.

[0617] Step 3: Store song information

[0618] The server adds new and updated song information to the music database. Song information includes title, artist, album, genre, and release year. Specifically, it creates database entries based on information provided by administrators and stakeholders. The input is song information and music files, and the output is that they are accurately stored in the database.

[0619] Step 4: Search for songs

[0620] The server uses the search query received from the user to search the music database and extract the corresponding songs. Specifically, it uses a full-text search engine (e.g., ElasticSearch) to quickly search for songs that match the query. The input is the search query, and the output is a list of songs as search results.

[0621] Step 5: Request a song to play

[0622] When a user uses a terminal to select a particular song, the terminal sends the selection to the server. The input is the user's song selection, and the output is a play request sent to the server.

[0623] Step 6: Stream your music

[0624] The server receives playback requests from users and streams the specified music. Specifically, it sends music files to terminals at a certain bit rate via HTTP or RTMP protocol. The input is the playback request, and the output is the streaming data.

[0625] Step 7: Playing a song

[0626] The device buffers the streaming data received from the server and plays it using the built-in music player application. The input is streaming data and the output is audio playback.

[0627] Step 8: Record play counts

[0628] The server records a playback event as a log each time a song is played. This log includes the song ID, playback start time, playback end time, user ID, etc. The input is the playback event information, and the output is the recorded log.

[0629] Step 9: Calculate and distribute revenue

[0630] The server analyzes the play count log and calculates revenue based on the number of plays of a specific song. This calculation is based on pre-defined rules and algorithms. Based on the analysis results, the revenue is distributed to the relevant parties (artist, lyricist, composer). The input is the play count log, and the output is the calculated revenue and its distribution information.

[0631] This enables the system to efficiently manage and utilize duplicated audio data using generative AI models, enabling fair revenue distribution.

[0632] (Application example 1)

[0633] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0634] On today's music distribution platforms, it is difficult for users to create and enjoy songs using the voice of a specific artist, and there is a demand for proper management of the number of plays and distribution of revenue for the created songs. Therefore, it is necessary to create songs using copied voices and provide them to users while realizing fair and transparent revenue distribution.

[0635] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0636] In this invention, the server includes a means for registering the copied audio data, a means for searching a music database, a means for streaming the music data, a means for users to generate music based on specific requests using a generative AI model, a means for recording the number of plays, and a means for calculating and distributing revenue. This allows users to freely generate new music using artists' voices and enjoy them in real time, while also enabling fair and transparent revenue distribution to artists and related parties.

[0637] "Duplicated audio data" refers to audio data that has been duplicated or generated using artificial intelligence technology from data recorded from the original artist's voice.

[0638] A "music database" is a collection of digital data that organizes and stores information about songs, and can be accessed by users through searches.

[0639] "Music Data" means data, including audio and metadata, of a musical composition stored in digital format.

[0640] "Streaming" is a method of playing music data in real time over the Internet.

[0641] "Plays" is a record of the number of times a particular song has been played by a user.

[0642] "Revenue sharing" is the process of fairly distributing revenues earned as a result of the playback of a generated song among the parties involved.

[0643] A "generative AI model" is an algorithm or system that uses artificial intelligence technology to generate new audio data or music.

[0644] The "user interface" is the part of the software that provides an operation screen for the user to search for and play music.

[0645] This invention provides a music streaming platform that uses duplicated audio using generative AI technology. The system includes functions such as registration of duplicated audio data, management of music data, streaming, recording of play counts, and revenue sharing.

[0646] System Configuration

[0647] 1. Server:

[0648] Audio data registration:

[0649] The server receives the voice data uploaded via the device and stores it in a database. Voice providers provide their own voice, and the server integrates this data into the AI ​​model.

[0650] Music database management:

[0651] The server manages a database that systematically stores song information, receives search queries from users, searches for and provides matching songs, and also adds and updates new songs.

[0652] Streaming:

[0653] The server streams the music selected by the user in real time and delivers it to the device, allowing the user to enjoy the music instantly.

[0654] Playback Tracking and Revenue Sharing:

[0655] The server records the number of times a song is played and calculates and distributes revenue based on that number, using specific rules and algorithms to ensure fair distribution among all parties involved.

[0656] Use of generative AI models:

[0657] It receives prompts based on the user's request and uses a generative AI model to generate new music, which is then instantly added to a database and made available for streaming.

[0658] 2. Device (smartphone):

[0659] Upload audio data:

[0660] The terminal provides an interface for the user to record the voice of the voice provider and upload it to the server.

[0661] Search and play songs:

[0662] The device provides an interface for users to search for and play music. Users can search for music within the app and enjoy music by pressing the play button.

[0663] Music generation using prompts:

[0664] The device provides an interface for users to input specific prompts to generate new music, which are then sent to a server and processed by a generative AI model.

[0665] Hardware and software used

[0666] The hardware includes servers (cloud-based servers or dedicated servers) and user devices (smartphones or tablets). As for software, the server side runs a web application based on Flask and a generative AI model using TensorFlow. The client side runs a smartphone application that provides the user interface.

[0667] Specific examples

[0668] For example, if a user wants to generate a new pop song using the voice of a particular singer, they might enter the prompt text as follows:

[0669] "Generate love songs with pop rhythms using singer's voices"

[0670] By passing this prompt to a generative AI model, a new pop song using the specified voice is generated and instantly added to the database, where users can then search for and stream the new song.

[0671] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0672] Step 1:

[0673] The user uses a device to record the voice of the voice provider and uploads the voice data from the device to the server. The server receives this voice data and stores it in a database as training data for the generative AI model. The input is the voice data, and the output is the voice data stored in the database.

[0674] Step 2:

[0675] The server integrates the received voice data into the generative AI model. The data processing performed here involves converting the voice data into an appropriate format and training the model so that it can generate new voices. The input is the voice data, and the output is updating the generative AI model.

[0676] Step 3:

[0677] A user inputs a search query for a song through a terminal. The terminal sends the query to a server, which then searches a music database to find the relevant song. The input is a search query, and the output is a list of songs as search results.

[0678] Step 4:

[0679] The user selects a specific song from the search results and instructs it to be played. The server streams the corresponding music data and delivers it to the device. The device receives this streaming data and plays it for the user. The input is the song ID, and the output is the song data played in real time.

[0680] Step 5:

[0681] The server records the user's playback actions and updates the database with the number of times the song has been played. This process takes the user ID and song ID as input, increments the number of times the song has been played, and updates the database. The input is the user ID and song ID, and the output is the updated number of times the song has been played.

[0682] Step 6:

[0683] The user uses a terminal to input a specific prompt sentence and request the generation of a new song. The server receives this prompt sentence and passes it to the generative AI model, which then generates a new song. The input is the prompt sentence, and the output is the generated song data. As a concrete example, we use a prompt sentence such as "Generate a love song with a pop rhythm and a singer's voice."

[0684] Step 7:

[0685] The server registers the generated new music in a database, allowing users to search and play it. The input is the generated music data, and the output is the new music registered in the database.

[0686] Step 8:

[0687] The server calculates and distributes revenue. It calculates revenue based on the number of views using specific rules and algorithms and distributes it to the parties involved. The input is the number of views, and the output is the calculated revenue and distribution results.

[0688] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0689] The present invention combines an emotion engine with a music streaming platform that uses singer voices replicated using generative AI technology. The system has the following components and functions:

[0690] 1. Registering duplicated voice data

[0691] Singers upload their own voice data via a terminal. The server receives this data and stores it in a database. This data is then integrated into an AI model and used as training data to be reflected in the generated music.

[0692] 2. Music database management

[0693] A music database is a system that stores song information in an organized manner. The server searches the music database based on a user's search query, retrieves the relevant songs, and provides them to the user. It also adds new songs and updates existing songs as needed.

[0694] 3. Streaming music data

[0695] When a user searches for and plays a song, the server quickly streams the corresponding music data to the device, allowing the user to enjoy music in real time. The device then plays the received streaming data, providing the user with a music experience.

[0696] 4. Recording Plays and Calculating Revenue

[0697] The server records the number of plays each time a song is played. Based on this play count data, revenue for each song is calculated and distributed appropriately. Specific rules and algorithms are used to calculate revenue, and the revenue is distributed fairly to each party involved (singer, lyricist, composer, etc.).

[0698] 5. Use of Emotion Engine

[0699] The system is equipped with an emotion engine for recognizing a user's emotions, which analyzes the user's voice, visual data, text input, and other physiological data to identify the user's current emotional state.

[0700] 1. Collecting Emotional Data

[0701] The device collects emotional data while the user listens to the music being played, including facial expression recognition using a camera, voice analysis using a microphone, and physiological data acquisition using sensors.

[0702] 2. Emotion Analysis

[0703] The server analyzes the collected emotion data to identify the user's emotional state. The emotion engine uses machine learning algorithms to identify the user's emotion based on the collected data.

[0704] 3. Song Recommendations

[0705] The server recommends songs to the user based on the identified emotional state, for example, if the user is in a sad state, the emotion engine selects and suggests comforting songs to the user.

[0706] 4. Emotion history storage and analysis

[0707] The server stores the user's emotional history in a database and performs long-term analysis, which allows it to understand the user's emotional patterns and provide a more personalized music experience.

[0708] Specific examples

[0709] 1. Uploading voice data and integrating AI models

[0710] The terminal accepts input for uploading singer's voice data.

[0711] The server receives the voice data, stores it, and integrates it into an AI model, which then uses it to generate new songs.

[0712] 2. Search and play songs

[0713] The user searches for a new song and clicks the play button.

[0714] The server searches for the music, generates streaming data, and delivers it to the device.

[0715] The terminal receives the streaming data and plays it back to the user.

[0716] 3. Recording of play counts and revenue sharing

[0717] The server records the number of times the song is played and calculates the revenue.

[0718] Profits are distributed among the parties according to certain rules.

[0719] 4. Utilizing the Emotion Engine

[0720] While the user is listening to music, the device collects emotion data.

[0721] The server uses an emotion engine to analyze the user's emotions and recommend appropriate songs.

[0722] It analyzes emotional history over the long term to provide users with the optimal music experience.

[0723] This system ensures that singers who use AI-replicated voices receive fair revenue, and that all creators involved receive appropriate recognition and compensation. Furthermore, the emotional engine can provide users with a more personalized music experience. This will improve fairness and transparency throughout the music industry, promoting greater creativity and innovation.

[0724] The processing flow will be explained below.

[0725] 1. Registering duplicated voice data

[0726] Processing Steps

[0727] Step 1:

[0728] The terminal accepts input from the user (singer). The singer selects his / her own voice data via the terminal and performs the upload operation.

[0729] Step 2:

[0730] The device sends voice data to the server, which receives and temporarily stores the data.

[0731] Step 3:

[0732] The server validates the voice data it receives, checking that the format and quality of the voice data meets standards.

[0733] Step 4:

[0734] The server stores the verified voice data in a voice database, and when it is saved, it associates the voice data with the singer's ID.

[0735] Step 5:

[0736] The server integrates the stored voice data into the AI ​​model, adding new voice data as training data for future music generation.

[0737] 2. Music database management

[0738] Processing Steps

[0739] Step 1:

[0740] The device accepts a song upload request from the user (artist). The artist selects their own song data through the device and performs the upload operation.

[0741] Step 2:

[0742] The device sends the music data to the server, which analyzes the received music data and temporarily stores it.

[0743] Step 3:

[0744] The server validates the song data, checking its integrity, format, and quality.

[0745] Step 4:

[0746] The server stores the verified music data in a music database, and associates the music data with artist information and metadata.

[0747] Step 5:

[0748] The server regularly backs up the music database to ensure data safety and availability.

[0749] 3. Streaming music data

[0750] Processing Steps

[0751] Step 1:

[0752] The device accepts the user's song search request. The user enters a search query and clicks the search button.

[0753] Step 2:

[0754] The device sends a search query to the server, which receives it and searches its music database.

[0755] Step 3:

[0756] The server generates the search results, creating a list of matching songs and sending it back to the device.

[0757] Step 4:

[0758] The terminal displays the search results to the user.

[0759] Step 5:

[0760] The user selects the song they want to play and clicks the play button. The device sends a playback request to the server.

[0761] Step 6:

[0762] The server generates streaming data for the specified music piece and transmits it to the terminal.

[0763] Step 7:

[0764] The terminal plays the received streaming data and provides music to the user.

[0765] 4. Recording Plays and Calculating Revenue

[0766] Processing Steps

[0767] Step 1:

[0768] The server detects when a song starts playing. Each time a request to start playing is received, it records the play count.

[0769] Step 2:

[0770] The server collects play count data and stores it in a database.

[0771] Step 3:

[0772] The server periodically aggregates the play count data and calculates revenue, which is then distributed based on rules and algorithms.

[0773] Step 4:

[0774] The server reflects the revenue data in the accounts of each party (singer, lyricist, composer, etc.).

[0775] Step 5:

[0776] Each party checks their own earnings through the terminal, which retrieves the earnings data from the server and displays it to the user.

[0777] 5. Use of Emotion Engine

[0778] Processing Steps

[0779] Step 1:

[0780] The device collects the user's emotional data. While the user is playing music, it uses a camera to recognize facial expressions, a microphone to analyze voices, and sensors to acquire physiological data.

[0781] Step 2:

[0782] The device sends the collected emotion data to a server, which receives it and temporarily stores it.

[0783] Step 3:

[0784] The server analyzes the emotion data, and the emotion engine uses machine learning algorithms to identify the user's emotion based on the collected data.

[0785] Step 4:

[0786] The server recommends songs based on the analyzed emotional data, suggesting songs that best suit the user's state of mind.

[0787] Step 5:

[0788] The server stores the user's emotional history in a database and performs long-term analysis, analyzing the user's emotional patterns to provide a more personalized music experience.

[0789] This detailed processing flow allows the system to efficiently register the duplicated voice data, stream the songs, record the number of plays, calculate and distribute revenue, and even recommend songs using an emotion engine.

[0790] Example 2

[0791] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0792] While conventional music streaming systems can play songs and distribute revenue, they lack the functionality to recommend songs based on user emotions or replicate artists' voices using AI. This makes it difficult to provide a music experience tailored to each user's emotions, and also poses the problem of insufficient fairness in revenue distribution for songs that use replicated voices.

[0793] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0794] In this invention, the server includes means for registering the duplicated voice data, means for searching a music database, means for streaming the music data, means for recording the number of plays, means for calculating and distributing revenue, means for collecting user emotion data, and means for analyzing the collected emotion data and recommending songs to the user. This makes it possible to provide a personalized music experience according to the user's emotion while maintaining fairness in the distribution of revenue from songs based on the duplicated voice using AI.

[0795] "Replica voice data" refers to audio information that has been recorded from an artist's voice and reproduced using a generative AI model.

[0796] A "music database" is an information system that efficiently manages metadata and audio files related to music, allowing them to be searched and stored.

[0797] "Streaming" is a technology that distributes and plays music data in real time to a user's terminal via the Internet.

[0798] "Plays" is an indicator that records the number of times a particular song has been played by a user.

[0799] "Revenue sharing" is the process of fairly distributing revenue from music content to stakeholders based on data such as the number of plays.

[0800] "Emotion data" is data that indicates the user's current emotional state, collected from the user's facial expressions, voice, physiological data, and the like.

[0801] A "generative AI model" is a collection of algorithms that perform specific tasks or generate results based on large amounts of data, and in this case refers specifically to models used to reproduce an artist's voice.

[0802] "Recommendation" refers to the act of selecting the most suitable music piece based on the analyzed emotional data of the user and suggesting it to the user.

[0803] This invention provides a personalized music experience for users by integrating an emotion engine into a music streaming system that uses replicated voice data. The main hardware used includes a database server, a terminal (including a camera, microphone, and sensors) that collects emotion data, and the user's device. The software used includes a generative AI model, an emotion analysis algorithm, streaming server software, and a user interface.

[0804] Singers record their own voice data on the device and upload the data to the server. The server receives this voice data, stores it in a database, and integrates it into the generative AI model. For example, a singer can record themselves singing specific lyrics on the device and send it to the server via a dedicated app. The server receives the voice data and stores it as replicated voice data. When integrated into the generative AI model, it is used as new voice data added to the existing voice dataset.

[0805] The server then manages a music database, which stores music metadata (such as title, artist name, and album name) and audio files in an organized manner. When a user searches for a song, the server searches the database based on the search query and returns the relevant songs. For example, if a user enters a prompt such as "Add a new song to the server," the server adds the new song's metadata and audio files to the database.

[0806] In music streaming, when a user searches for a specific song and clicks the play button, the server prepares the song data in real time and delivers it to the user's device as streaming data. The device receives the streaming data and immediately starts playing it. For example, the server can generate streaming data and provide it to the user by using a prompt such as "Please search for a specific song and play it."

[0807] Regarding the recording of play counts and revenue calculation, the server records the number of plays in real time each time each song is played. Based on this data, revenue is calculated and fairly distributed to the parties according to specific rules. The revenue calculation takes into account the number of plays, song information, related contract information, etc. For example, upon the prompt "Please calculate and distribute revenue based on the number of song plays," the server will perform the revenue calculation and save the result in the database.

[0808] When using an emotion engine, the user's device uses a camera, microphone, and sensors to collect the user's facial, voice, and physiological data. The collected emotion data is sent to a server and analyzed using an emotion analysis algorithm. Based on the analysis results, the server recommends music that best suits the user's emotional state. For example, a prompt such as "Analyze the user's emotion data and recommend music that matches that emotion" allows the server to analyze the emotion data and recommend appropriate music.

[0809] As described above, this invention provides a music streaming system that uses replicated voice data to recommend songs based on the user's emotions, enabling users to enjoy a personalized music experience. Furthermore, fairness in revenue distribution is maintained, making this a system that benefits all parties involved.

[0810] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0811] Step 1:

[0812] The terminal provides a user interface for singers to record their voice data using an input device. The user presses the record button, and after recording is completed, clicks the upload button. The input is the recorded voice data, and the output is an upload request to the server.

[0813] Step 2:

[0814] The server receives the voice data sent from the terminal and performs pre-processing such as data formatting and format conversion.The server then stores the formatted voice data in a voice database.The input is the uploaded voice data, and the output is the voice data stored in the database.

[0815] Step 3:

[0816] The server adds the stored voice data to a training dataset for integration into a generative AI model. This training dataset is used to generate replicated voices. The input is the voice data stored in the database, and the output is the training data integrated into the generative AI model.

[0817] Step 4:

[0818] A user enters a specific song title into the search bar of their device and presses the search button. A song search request is sent from the device to the server. The input is the search query entered by the user, and the output is the song search request.

[0819] Step 5:

[0820] The server searches the music database based on the user's search query, retrieves the corresponding song data, and returns the retrieved song data to the device. The input is the search query, and the output is the corresponding song data.

[0821] Step 6:

[0822] The device receives the music data sent from the server and launches the streaming player. When the user clicks the play button, the music starts playing. The input is the music data from the server, and the output is the music being played.

[0823] Step 7:

[0824] The server records the number of times a song is played in real time. This play count data is later used to calculate revenue. The input is the song play event, and the output is the recorded play count data.

[0825] Step 8:

[0826] The server uses a revenue calculation algorithm to calculate revenue based on the number of times a song is played and distribute it to the parties involved. The input is the recorded number of times a song is played and the revenue distribution rule, and the output is the revenue data distributed to each party.

[0827] Step 9:

[0828] The device uses sensors, cameras, and microphones to collect emotional data while the user is listening to music. The collected emotional data is sent to a server. The input is the user's physiological response, and the output is the collected emotional data.

[0829] Step 10:

[0830] The server analyzes the collected emotional data using an emotion analysis algorithm to identify the user's current emotional state, where the input is the collected emotional data and the output is the analyzed emotional state.

[0831] Step 11:

[0832] The server recommends the best songs to the user based on the analyzed emotional state. The input is the analyzed emotional state, and the output is a list of recommended songs.

[0833] Step 12:

[0834] The server stores the user's emotion history in a database and uses it to analyze long-term emotion patterns. The input is the analyzed emotion data, and the output is the emotion history stored in the database.

[0835] (Application example 2)

[0836] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0837] Traditional music streaming services lack personalized song recommendations based on the user's emotional state and new music experiences using AI-generated singers. This makes it difficult to respond to diverse user emotions and preferences, and has made it difficult to improve user satisfaction and provide novel music experiences. Another problem is the lack of transparency regarding revenue distribution to music creators.

[0838] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0839] In this invention, the server includes means for registering copied voice data, means for searching a music database, means for streaming music data, means for recording the number of plays, means for calculating and distributing revenue, means for collecting user emotion data, means for analyzing the emotion data to identify the user's emotional state, means for recommending songs based on the identified emotional state, and means for storing and analyzing the emotion history over time. This enables personalized song recommendations based on the user's emotional state, improving user satisfaction and providing a novel music experience. It also enables transparent and fair revenue distribution to music creators.

[0840] "Duplicated voice data" refers to voice data that reproduces a singer's voice using generative AI technology.

[0841] A "music database" is a system that systematically stores information about multiple songs.

[0842] "Streaming" is a technology that distributes stored music data in real time over the Internet.

[0843] "Play count" is data that records the number of times a particular song has been played by a user.

[0844] "Revenue sharing" is the process of appropriately distributing revenue earned based on the number of times a song is played among the parties involved.

[0845] "Emotional data" is data that indicates the user's current emotional state, and is mainly collected from voice, visual information, physiological data, and the like.

[0846] "Emotion analysis" is the process of identifying a user's emotional state based on collected emotional data.

[0847] "Music recommendation" is the act of suggesting the most suitable music based on the analyzed emotional state of the user.

[0848] "Emotion history" is data that records changes in the user's emotional state over a long period of time.

[0849] An "AI model" is a data model that is optimized for a specific task using machine learning algorithms.

[0850] The present invention relates to a music streaming platform using replicated voice data and an emotion engine, and can be implemented by applying the following elements and means:

[0851] 1. System Configuration

[0852] Hardware

[0853] Camera: Captures the user's facial expressions.

[0854] Microphone: Collects the user's voice and physiological data.

[0855] Smartphone: Used as an application execution environment.

[0856] software

[0857] OpenCV: Used as an image processing library to acquire camera input and perform facial expression recognition.

[0858] Librosa: A library for music analysis. It extracts features such as tempo from music data.

[0859] TensorFlow / Keras: Load and run the emotion engine model.

[0860] Requests: Used to send HTTP requests to the music recommendation service.

[0861] 2. System Operation

[0862] The server implements various functions using the following means:

[0863] Registering duplicated voice data

[0864] Singers upload their own voice data via a terminal. The server receives this data and stores it in a database. The server then integrates this data into an AI model and generates new songs using generative AI technology.

[0865] Music database management

[0866] The server searches the music database based on the user's search query, retrieves the relevant songs, and provides them to the user. It also adds new songs and updates existing songs.

[0867] Streaming music data

[0868] When a user searches for and plays a song, the server quickly streams the corresponding music data to the device, allowing the user to enjoy music in real time.

[0869] Recording views and calculating revenue

[0870] The server records the number of plays each time a song is played, and calculates revenue based on this data, which is then distributed equally among all parties involved (singers, lyricists, composers, etc.).

[0871] Use of emotion engine

[0872] The server collects users' emotional data and analyzes it using an emotion engine. It recommends songs based on the user's emotional state and analyzes their emotional history over time to provide a more personalized music experience.

[0873] 3. Specific Examples

[0874] Emotion data collection and analysis

[0875] While the user is listening to music, the device's camera and microphone work together to capture the user's facial expressions and vocal characteristics, which are then analyzed by the emotion engine to determine the user's emotional state.

[0876] Song recommendations

[0877] The server recommends songs based on the identified emotional state, for example, if the user is in a sad state, the emotion engine selects and recommends comforting songs.

[0878] Example: Using a prompt statement

[0879] The implicit knowledge is that when the user is in a sad state, the prompt sentence is:

[0880] "The user is in a sad state, so please recommend some songs that are slow and comforting."

[0881] Long-term emotion history analysis

[0882] The server stores the user's emotional history in a database and analyzes it over the long term, allowing it to understand the user's emotional patterns and reflect them in future song recommendations.

[0883] By combining these elements and methods, it becomes possible to provide a personalized music streaming service that responds to the user's emotional state, thereby improving user satisfaction and providing a novel music experience.

[0884] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0885] Step 1:

[0886] The user launches a music application.

[0887] Input: User launches application.

[0888] Output: The application home screen is displayed.

[0889] Step 2:

[0890] The user uploads the singer's voice data.

[0891] Input: User selection and upload of voice data.

[0892] Output: Duplicate voice data stored on the server.

[0893] Specific behavior:

[0894] The terminal receives selected voice data from the user.

[0895] The received data is sent to the server and stored in a database.

[0896] The voice data is used as training data to integrate into generative AI models.

[0897] Step 3:

[0898] The server searches the music database.

[0899] Input: A search query by the user.

[0900] Output: A list of songs as search results.

[0901] Specific behavior:

[0902] The user types a query into the search bar and sends it to the server.

[0903] The server searches the music database based on the received query and retrieves the corresponding songs.

[0904] Step 4:

[0905] The user plays a song.

[0906] Input: User clicks the play button.

[0907] Output: The streaming data to be played.

[0908] Specific behavior:

[0909] The server delivers streaming data of the selected music to the terminal.

[0910] The terminal plays the received streaming data, providing the user with a musical experience.

[0911] Step 5:

[0912] The device collects the user's emotional data.

[0913] Input: User's facial and voice data.

[0914] Output: Collected emotion data.

[0915] Specific behavior:

[0916] The device's camera and microphone are activated to capture the user's facial expressions and voice.

[0917] The acquired data is preprocessed and sent to the server.

[0918] Step 6:

[0919] The server analyzes the emotional data to determine the user's emotional state.

[0920] Input: Collected emotion data.

[0921] Output: The identified emotional state of the user.

[0922] Specific behavior:

[0923] The server analyzes the emotion data using machine learning algorithms.

[0924] An emotion engine identifies the user's current emotional state.

[0925] Step 7:

[0926] The server recommends songs based on the identified emotional state.

[0927] Input: The identified emotional state of the user.

[0928] Output: A list of recommended songs.

[0929] Specific behavior:

[0930] The server searches a music database for songs that correspond to the identified emotional state.

[0931] A list of recommended songs is presented to the user.

[0932] Step 8:

[0933] The server records the number of plays and calculates the revenue.

[0934] Input: Song playback event.

[0935] Output: Calculated revenue and distribution results.

[0936] Specific behavior:

[0937] The server records the number of times a song is played in a database each time it is played.

[0938] Revenue is calculated based on the number of plays and distributed to each party.

[0939] Step 9:

[0940] The server stores emotion history and analyzes it over the long term.

[0941] Input: Identified emotional state.

[0942] Output: The stored emotion history and its analysis results.

[0943] Specific behavior:

[0944] The server stores the history of the user's emotional state in a database.

[0945] Conduct long-term analysis to understand user sentiment patterns.

[0946] Through the above steps, a music streaming experience based on the user's emotional state can be provided.

[0947] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0948] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0949] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0950] [Third embodiment]

[0951] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0952] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0953] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0954] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0955] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0956] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0957] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0958] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0959] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0960] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0961] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0962] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0963] The present invention relates to a music streaming platform that uses singer voices replicated using generative AI technology. The system has the following components and functions:

[0964] 1. Registering duplicated voice data

[0965] Singers upload their own voice data via a terminal. The server receives this data and stores it in a database. This data is then integrated into an AI model and used as training data to be reflected in the generated music.

[0966] 2. Music database management

[0967] A music database is a system that stores song information in an organized manner. The server searches the music database based on a user's search query, retrieves the relevant songs, and provides them to the user. It also adds new songs and updates existing songs as needed.

[0968] 3. Streaming music data

[0969] When a user searches for and plays a song, the server quickly streams the corresponding music data to the device, allowing the user to enjoy music in real time. The device then plays the received streaming data, providing the user with a music experience.

[0970] 4. Recording Plays and Calculating Revenue

[0971] The server records the number of plays each time a song is played. Based on this play count data, revenue for each song is calculated and distributed appropriately. Specific rules and algorithms are used to calculate revenue, and the revenue is distributed fairly to each party involved (singer, lyricist, composer, etc.).

[0972] Specific examples

[0973] As an example, consider a scenario in which user A wants to listen to a new song that has been created using the voice of a particular singer B.

[0974] 1. Uploading voice data

[0975] Singer B uploads his / her voice data to the server via his / her device, which receives this data and integrates it into the AI ​​model.

[0976] 2. Music Generation

[0977] Composer C creates a new song and requests to use the voice of singer B. The server uses the AI ​​model to generate the new song in singer B's voice.

[0978] 3. Search and play songs

[0979] User A searches for new songs through a device. The device sends a query to the server, which searches the music database and finds the relevant new songs. The server streams the new songs and delivers them to the device. The device plays the received streaming data, allowing User A to listen to the new songs.

[0980] 4. Playback count and revenue sharing

[0981] The server records the number of plays each time User A plays a new song. Based on the play count data, the server calculates and distributes revenue to Singer B, Composer C, and other parties.

[0982] This system will ensure that singers who use AI-generated voices receive fair revenue, and that all creators involved receive appropriate recognition and compensation. This will improve fairness and transparency throughout the music industry, and promote greater creativity and innovation.

[0983] The processing flow will be explained below.

[0984] 1. Registering duplicated voice data

[0985] Processing Steps

[0986] Step 1:

[0987] The terminal accepts input from the user (singer). The singer selects their own voice data through the terminal and performs the upload operation.

[0988] Step 2:

[0989] The device sends voice data to the server, which receives and temporarily stores the data.

[0990] Step 3:

[0991] The server validates the voice data it receives, checking that the data format and quality meet the standards.

[0992] Step 4:

[0993] The server stores the verified voice data in a voice database, where it associates it with the singer's ID.

[0994] Step 5:

[0995] The server integrates the stored voice data into the AI ​​model, adding new voice data as training data for future music generation.

[0996] 2. Music database management

[0997] Processing Steps

[0998] Step 1:

[0999] The device accepts the user's request to upload music. The user (artist) selects their own music data through the device and performs the upload operation.

[1000] Step 2:

[1001] The device sends the music data to the server, which analyzes the received music data and temporarily stores it.

[1002] Step 3:

[1003] The server validates the song data, checking file integrity, format, and data completeness.

[1004] Step 4:

[1005] The server stores the verified music data in a music database, and associates the music data with artist information and metadata.

[1006] Step 5:

[1007] The server regularly backs up the music database to ensure data safety.

[1008] 3. Streaming music data

[1009] Processing Steps

[1010] Step 1:

[1011] The device accepts the user's song search request. The user enters a search query and clicks the search button.

[1012] Step 2:

[1013] The device sends a search query to the server, which receives it and searches its music database.

[1014] Step 3:

[1015] The server generates the search results, creates a list of matching songs, and returns it to the device.

[1016] Step 4:

[1017] The terminal displays the search results to the user.

[1018] Step 5:

[1019] The user selects the song they want to play and clicks the play button. The device sends a playback request to the server.

[1020] Step 6:

[1021] The server generates streaming data for the specified music piece and transmits it to the terminal.

[1022] Step 7:

[1023] The terminal plays the received streaming data and provides music to the user.

[1024] 4. Recording Plays and Calculating Revenue

[1025] Processing Steps

[1026] Step 1:

[1027] The server detects when a song starts playing. Each time a request to start playing is received, it records the play count.

[1028] Step 2:

[1029] The server collects the play count data and stores it in a database.

[1030] Step 3:

[1031] The server periodically aggregates the play count data and calculates revenue, which is then distributed based on specific rules and algorithms.

[1032] Step 4:

[1033] The server reflects the revenue data in the accounts of each party (singer, lyricist, composer, etc.).

[1034] Step 5:

[1035] Each party checks their own earnings through the terminal, which retrieves the earnings data from the server and displays it to the user.

[1036] This detailed processing flow allows the system to efficiently register the duplicated voice data, stream the song, record the number of plays, and calculate and distribute revenue.

[1037] Example 1

[1038] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1039] In conventional music streaming systems, music data is provided only using existing audio data, making it difficult to generate new audio data. Furthermore, there is no system in place for managing the appropriate use of copied audio data or for revenue distribution. Therefore, there is a need to improve fairness and transparency throughout the music industry and promote new creativity and innovation.

[1040] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1041] In this invention, the server includes a means for registering the replicated audio data, a means for managing a music database, a means for streaming the music data, a means for recording the number of plays, a means for calculating and distributing revenue, and a means for using a generative AI model to generate the replicated audio data. This not only allows users to enjoy high-quality audio data, but also enables fair revenue distribution based on the usage. Furthermore, utilizing the generative AI model enables the generation of new audio data and the provision of songs based on that data, promoting new creativity and innovation in the music industry.

[1042] "Duplicated voice data" is new voice data generated based on the original voice data using a generative AI model.

[1043] A "music database" is a database for systematically storing and managing song information and music files.

[1044] "Streaming" is a technology that distributes music data in real time over the Internet, allowing it to be played instantly on the user's device.

[1045] The "number of plays" is data indicating the number of times a particular song has been played by a user.

[1046] "Revenue" refers to income generated based on the playback of music data, which must be distributed appropriately to the parties involved.

[1047] A "generative AI model" is a model that uses machine learning algorithms to generate new data based on specific input data (in this case, voice data).

[1048] "Training data" refers to the dataset used to train a generative AI model, in this case the original audio data needed to generate the replicated audio data.

[1049] "Interface" refers to the operation screen and input means that a user uses to search for and play music.

[1050] The present invention relates to a music streaming platform that uses replicated audio data using generative AI technology. The system has the following components and functions:

[1051] 1. Registering duplicated audio data

[1052] Artist users upload their audio files using their devices. The devices then send the files to the server using HTTP POST requests. The server validates the received audio data and stores it in a secure database. The server then integrates the audio data into a generative AI model (e.g., WaveNet or Tacotron2) and uses it as training data to influence the generated music.

[1053] 2. Music database management

[1054] The server stores song information and music files in a music database in an organized manner. Song information includes title, artist, album, genre, and release year. The server also searches the database based on user search queries to extract and provide matching songs. New songs are added and existing songs are updated as needed.

[1055] 3. Streaming music data

[1056] When a user searches for and plays a song through a device, the server quickly streams the selected song to the device using HTTP or RTMP protocols. The device buffers the received streaming data and plays it using the built-in music player application, providing the user with a real-time music experience.

[1057] 4. Recording Plays and Calculating Revenue

[1058] The server records the number of plays as a log each time a song is played. This log includes information such as the song ID, playback start time, playback end time, and user ID. The log data is sent to a big data platform (e.g., Hadoop or Spark) for analysis. Based on the analysis, the server uses the play count data to calculate revenue and distributes it fairly to each relevant party (artist, lyricist, composer). Pre-set rules and algorithms are applied to calculate the revenue.

[1059] Specific examples

[1060] A specific example is given below: For example, consider a scenario in which user A listens to a new song created using the voice of a particular artist.

[1061] 1. Uploading voice data

[1062] Artists record their own voices and upload the audio data via their devices to a server, which receives the data and integrates it into a generative AI model.

[1063] 2. Music Generation

[1064] A composer creates a new song and requests that the artist's voice be used, and the server uses a generative AI model to generate the new song in the artist's voice.

[1065] 3. Search and play songs

[1066] User A searches for new songs through a device. The device sends a query to the server, which searches the music database to find the new songs. The server streams the new songs, and the device plays the received streaming data, allowing User A to listen to the new songs.

[1067] 4. Playback count and revenue sharing

[1068] The server records the number of plays each time User A plays a new song. Based on this data, the server calculates revenue and distributes it to the artists, composers, and other parties involved.

[1069] This system will enable efficient management and utilization of duplicated audio data using generative AI models, enabling fair revenue distribution, improving fairness and transparency across the music industry and encouraging greater creativity and innovation.

[1070] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1071] Step 1: Upload your voice data

[1072] The user, an artist, uses a device to record his or her own voice and uploads the audio data file to the server. Specifically, the user uses the device's recording function to create an audio file (e.g., WAV or MP3 format) and sends it to the server via an upload interface. The input is the audio data file, and the output is that file being saved on the server.

[1073] Step 2: Saving and integrating audio data

[1074] The server receives the uploaded audio data and verifies its integrity and quality. After verification is complete, it stores it in a database and performs preprocessing for incorporation into the generative AI model. Specifically, it performs noise removal and sampling rate adjustment. The input is the uploaded audio data file, and the output is the preprocessed audio data stored in the database.

[1075] Step 3: Store song information

[1076] The server adds new and updated song information to the music database. Song information includes title, artist, album, genre, and release year. Specifically, it creates database entries based on information provided by administrators and stakeholders. The input is song information and music files, and the output is that they are accurately stored in the database.

[1077] Step 4: Search for songs

[1078] The server uses the search query received from the user to search the music database and extract the corresponding songs. Specifically, it uses a full-text search engine (e.g., ElasticSearch) to quickly search for songs that match the query. The input is the search query, and the output is a list of songs as search results.

[1079] Step 5: Request a song to play

[1080] When a user uses a terminal to select a particular song, the terminal sends the selection to the server. The input is the user's song selection, and the output is a play request sent to the server.

[1081] Step 6: Stream your music

[1082] The server receives playback requests from users and streams the specified music. Specifically, it sends music files to terminals at a certain bit rate via HTTP or RTMP protocol. The input is the playback request, and the output is the streaming data.

[1083] Step 7: Playing a song

[1084] The device buffers the streaming data received from the server and plays it using the built-in music player application. The input is streaming data and the output is audio playback.

[1085] Step 8: Record play counts

[1086] The server records a playback event as a log each time a song is played. This log includes the song ID, playback start time, playback end time, user ID, etc. The input is the playback event information, and the output is the recorded log.

[1087] Step 9: Calculate and distribute revenue

[1088] The server analyzes the play count log and calculates revenue based on the number of plays of a specific song. This calculation is based on pre-defined rules and algorithms. Based on the analysis results, the revenue is distributed to the relevant parties (artist, lyricist, composer). The input is the play count log, and the output is the calculated revenue and its distribution information.

[1089] This enables the system to efficiently manage and utilize duplicated audio data using generative AI models, enabling fair revenue distribution.

[1090] (Application example 1)

[1091] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1092] On today's music distribution platforms, it is difficult for users to create and enjoy songs using the voice of a specific artist, and there is a demand for proper management of the number of plays and distribution of revenue for the created songs. Therefore, it is necessary to create songs using copied voices and provide them to users while realizing fair and transparent revenue distribution.

[1093] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1094] In this invention, the server includes a means for registering the copied audio data, a means for searching a music database, a means for streaming the music data, a means for users to generate music based on specific requests using a generative AI model, a means for recording the number of plays, and a means for calculating and distributing revenue. This allows users to freely generate new music using artists' voices and enjoy them in real time, while also enabling fair and transparent revenue distribution to artists and related parties.

[1095] "Duplicated audio data" refers to audio data that has been duplicated or generated using artificial intelligence technology from data recorded from the original artist's voice.

[1096] A "music database" is a collection of digital data that organizes and stores information about songs, and can be accessed by users through searches.

[1097] "Music Data" means data, including audio and metadata, of a musical composition stored in digital format.

[1098] "Streaming" is a method of playing music data in real time over the Internet.

[1099] "Plays" is a record of the number of times a particular song has been played by a user.

[1100] "Revenue sharing" is the process of fairly distributing revenues earned as a result of the playback of a generated song among the parties involved.

[1101] A "generative AI model" is an algorithm or system that uses artificial intelligence technology to generate new audio data or music.

[1102] The "user interface" is the part of the software that provides an operation screen for the user to search for and play music.

[1103] This invention provides a music streaming platform that uses duplicated audio using generative AI technology. The system includes functions such as registration of duplicated audio data, management of music data, streaming, recording of play counts, and revenue sharing.

[1104] System Configuration

[1105] 1. Server:

[1106] Audio data registration:

[1107] The server receives the voice data uploaded via the device and stores it in a database. Voice providers provide their own voice, and the server integrates this data into the AI ​​model.

[1108] Music database management:

[1109] The server manages a database that systematically stores song information, receives search queries from users, searches for and provides matching songs, and also adds and updates new songs.

[1110] Streaming:

[1111] The server streams the music selected by the user in real time and delivers it to the device, allowing the user to enjoy the music instantly.

[1112] Playback Tracking and Revenue Sharing:

[1113] The server records the number of times a song is played and calculates and distributes revenue based on that number, using specific rules and algorithms to ensure fair distribution among all parties involved.

[1114] Use of generative AI models:

[1115] It receives prompts based on the user's request and uses a generative AI model to generate new music, which is then instantly added to a database and made available for streaming.

[1116] 2. Device (smartphone):

[1117] Upload audio data:

[1118] The terminal provides an interface for the user to record the voice of the voice provider and upload it to the server.

[1119] Search and play songs:

[1120] The device provides an interface for users to search for and play music. Users can search for music within the app and enjoy music by pressing the play button.

[1121] Music generation using prompts:

[1122] The device provides an interface for users to input specific prompts to generate new music, which are then sent to a server and processed by a generative AI model.

[1123] Hardware and software used

[1124] The hardware includes servers (cloud-based servers or dedicated servers) and user devices (smartphones or tablets). As for software, the server side runs a web application based on Flask and a generative AI model using TensorFlow. The client side runs a smartphone application that provides the user interface.

[1125] Specific examples

[1126] For example, if a user wants to generate a new pop song using the voice of a particular singer, they might enter the prompt text as follows:

[1127] "Generate love songs with pop rhythms using singer's voices"

[1128] By passing this prompt to a generative AI model, a new pop song using the specified voice is generated and instantly added to the database, where users can then search for and stream the new song.

[1129] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1130] Step 1:

[1131] The user uses a device to record the voice of the voice provider and uploads the voice data from the device to the server. The server receives this voice data and stores it in a database as training data for the generative AI model. The input is the voice data, and the output is the voice data stored in the database.

[1132] Step 2:

[1133] The server integrates the received voice data into the generative AI model. The data processing performed here involves converting the voice data into an appropriate format and training the model so that it can generate new voices. The input is the voice data, and the output is updating the generative AI model.

[1134] Step 3:

[1135] A user inputs a search query for a song through a terminal. The terminal sends the query to a server, which then searches a music database to find the relevant song. The input is a search query, and the output is a list of songs as search results.

[1136] Step 4:

[1137] The user selects a specific song from the search results and instructs it to be played. The server streams the corresponding music data and delivers it to the device. The device receives this streaming data and plays it for the user. The input is the song ID, and the output is the song data played in real time.

[1138] Step 5:

[1139] The server records the user's playback actions and updates the database with the number of times the song has been played. This process takes the user ID and song ID as input, increments the number of times the song has been played, and updates the database. The input is the user ID and song ID, and the output is the updated number of times the song has been played.

[1140] Step 6:

[1141] The user uses a terminal to input a specific prompt sentence and request the generation of a new song. The server receives this prompt sentence and passes it to the generative AI model, which then generates a new song. The input is the prompt sentence, and the output is the generated song data. As a concrete example, we use a prompt sentence such as "Generate a love song with a pop rhythm and a singer's voice."

[1142] Step 7:

[1143] The server registers the generated new music in a database, allowing users to search and play it. The input is the generated music data, and the output is the new music registered in the database.

[1144] Step 8:

[1145] The server calculates and distributes revenue. It calculates revenue based on the number of views using specific rules and algorithms and distributes it to the parties involved. The input is the number of views, and the output is the calculated revenue and distribution results.

[1146] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1147] The present invention combines an emotion engine with a music streaming platform that uses singer voices replicated using generative AI technology. The system has the following components and functions:

[1148] 1. Registering duplicated voice data

[1149] Singers upload their own voice data via a terminal. The server receives this data and stores it in a database. This data is then integrated into an AI model and used as training data to be reflected in the generated music.

[1150] 2. Music database management

[1151] A music database is a system that stores song information in an organized manner. The server searches the music database based on a user's search query, retrieves the relevant songs, and provides them to the user. It also adds new songs and updates existing songs as needed.

[1152] 3. Streaming music data

[1153] When a user searches for and plays a song, the server quickly streams the corresponding music data to the device, allowing the user to enjoy music in real time. The device then plays the received streaming data, providing the user with a music experience.

[1154] 4. Recording Plays and Calculating Revenue

[1155] The server records the number of plays each time a song is played. Based on this play count data, revenue for each song is calculated and distributed appropriately. Specific rules and algorithms are used to calculate revenue, and the revenue is distributed fairly to each party involved (singer, lyricist, composer, etc.).

[1156] 5. Use of Emotion Engine

[1157] The system is equipped with an emotion engine for recognizing a user's emotions, which analyzes the user's voice, visual data, text input, and other physiological data to identify the user's current emotional state.

[1158] 1. Collecting Emotional Data

[1159] The device collects emotional data while the user listens to the music being played, including facial expression recognition using a camera, voice analysis using a microphone, and physiological data acquisition using sensors.

[1160] 2. Emotion Analysis

[1161] The server analyzes the collected emotion data to identify the user's emotional state. The emotion engine uses machine learning algorithms to identify the user's emotion based on the collected data.

[1162] 3. Song Recommendations

[1163] The server recommends songs to the user based on the identified emotional state, for example, if the user is in a sad state, the emotion engine selects and suggests comforting songs to the user.

[1164] 4. Emotion history storage and analysis

[1165] The server stores the user's emotional history in a database and performs long-term analysis, which allows it to understand the user's emotional patterns and provide a more personalized music experience.

[1166] Specific examples

[1167] 1. Uploading voice data and integrating AI models

[1168] The terminal accepts input for uploading singer's voice data.

[1169] The server receives the voice data, stores it, and integrates it into an AI model, which then uses it to generate new songs.

[1170] 2. Search and play songs

[1171] The user searches for a new song and clicks the play button.

[1172] The server searches for the music, generates streaming data, and delivers it to the device.

[1173] The terminal receives the streaming data and plays it back to the user.

[1174] 3. Recording of play counts and revenue sharing

[1175] The server records the number of times the song is played and calculates the revenue.

[1176] Profits are distributed among the parties according to certain rules.

[1177] 4. Utilizing the Emotion Engine

[1178] While the user is listening to music, the device collects emotion data.

[1179] The server uses an emotion engine to analyze the user's emotions and recommend appropriate songs.

[1180] It analyzes emotional history over the long term to provide users with the optimal music experience.

[1181] This system ensures that singers who use AI-replicated voices receive fair revenue, and that all creators involved receive appropriate recognition and compensation. Furthermore, the emotional engine can provide users with a more personalized music experience. This will improve fairness and transparency throughout the music industry, promoting greater creativity and innovation.

[1182] The processing flow will be explained below.

[1183] 1. Registering duplicated voice data

[1184] Processing Steps

[1185] Step 1:

[1186] The terminal accepts input from the user (singer). The singer selects his / her own voice data via the terminal and performs the upload operation.

[1187] Step 2:

[1188] The device sends voice data to the server, which receives and temporarily stores the data.

[1189] Step 3:

[1190] The server validates the voice data it receives, checking that the format and quality of the voice data meets standards.

[1191] Step 4:

[1192] The server stores the verified voice data in a voice database, and when it is saved, it associates the voice data with the singer's ID.

[1193] Step 5:

[1194] The server integrates the stored voice data into the AI ​​model, adding new voice data as training data for future music generation.

[1195] 2. Music database management

[1196] Processing Steps

[1197] Step 1:

[1198] The device accepts a song upload request from the user (artist). The artist selects their own song data through the device and performs the upload operation.

[1199] Step 2:

[1200] The device sends the music data to the server, which analyzes the received music data and temporarily stores it.

[1201] Step 3:

[1202] The server validates the song data, checking its integrity, format, and quality.

[1203] Step 4:

[1204] The server stores the verified music data in a music database, and associates the music data with artist information and metadata.

[1205] Step 5:

[1206] The server regularly backs up the music database to ensure data safety and availability.

[1207] 3. Streaming music data

[1208] Processing Steps

[1209] Step 1:

[1210] The device accepts the user's song search request. The user enters a search query and clicks the search button.

[1211] Step 2:

[1212] The device sends a search query to the server, which receives it and searches its music database.

[1213] Step 3:

[1214] The server generates the search results, creating a list of matching songs and sending it back to the device.

[1215] Step 4:

[1216] The terminal displays the search results to the user.

[1217] Step 5:

[1218] The user selects the song they want to play and clicks the play button. The device sends a playback request to the server.

[1219] Step 6:

[1220] The server generates streaming data for the specified music piece and transmits it to the terminal.

[1221] Step 7:

[1222] The terminal plays the received streaming data and provides music to the user.

[1223] 4. Recording Plays and Calculating Revenue

[1224] Processing Steps

[1225] Step 1:

[1226] The server detects when a song starts playing. Each time a request to start playing is received, it records the play count.

[1227] Step 2:

[1228] The server collects play count data and stores it in a database.

[1229] Step 3:

[1230] The server periodically aggregates the play count data and calculates revenue, which is then distributed based on rules and algorithms.

[1231] Step 4:

[1232] The server reflects the revenue data in the accounts of each party (singer, lyricist, composer, etc.).

[1233] Step 5:

[1234] Each party checks their own earnings through the terminal, which retrieves the earnings data from the server and displays it to the user.

[1235] 5. Use of Emotion Engine

[1236] Processing Steps

[1237] Step 1:

[1238] The device collects the user's emotional data. While the user is playing music, it uses a camera to recognize facial expressions, a microphone to analyze voices, and sensors to acquire physiological data.

[1239] Step 2:

[1240] The device sends the collected emotion data to a server, which receives it and temporarily stores it.

[1241] Step 3:

[1242] The server analyzes the emotion data, and the emotion engine uses machine learning algorithms to identify the user's emotion based on the collected data.

[1243] Step 4:

[1244] The server recommends songs based on the analyzed emotional data, suggesting songs that best suit the user's state of mind.

[1245] Step 5:

[1246] The server stores the user's emotional history in a database and performs long-term analysis, analyzing the user's emotional patterns to provide a more personalized music experience.

[1247] This detailed processing flow allows the system to efficiently register the duplicated voice data, stream the songs, record the number of plays, calculate and distribute revenue, and even recommend songs using an emotion engine.

[1248] Example 2

[1249] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1250] While conventional music streaming systems can play songs and distribute revenue, they lack the functionality to recommend songs based on user emotions or replicate artists' voices using AI. This makes it difficult to provide a music experience tailored to each user's emotions, and also poses the problem of insufficient fairness in revenue distribution for songs that use replicated voices.

[1251] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1252] In this invention, the server includes means for registering the duplicated voice data, means for searching a music database, means for streaming the music data, means for recording the number of plays, means for calculating and distributing revenue, means for collecting user emotion data, and means for analyzing the collected emotion data and recommending songs to the user. This makes it possible to provide a personalized music experience according to the user's emotion while maintaining fairness in the distribution of revenue from songs based on the duplicated voice using AI.

[1253] "Replica voice data" refers to audio information that has been recorded from an artist's voice and reproduced using a generative AI model.

[1254] A "music database" is an information system that efficiently manages metadata and audio files related to music, allowing them to be searched and stored.

[1255] "Streaming" is a technology that distributes and plays music data in real time to a user's terminal via the Internet.

[1256] "Plays" is an indicator that records the number of times a particular song has been played by a user.

[1257] "Revenue sharing" is the process of fairly distributing revenue from music content to stakeholders based on data such as the number of plays.

[1258] "Emotion data" is data that indicates the user's current emotional state, collected from the user's facial expressions, voice, physiological data, and the like.

[1259] A "generative AI model" is a collection of algorithms that perform specific tasks or generate results based on large amounts of data, and in this case refers specifically to models used to reproduce an artist's voice.

[1260] "Recommendation" refers to the act of selecting the most suitable music piece based on the analyzed emotional data of the user and suggesting it to the user.

[1261] This invention provides a personalized music experience for users by integrating an emotion engine into a music streaming system that uses replicated voice data. The main hardware used includes a database server, a terminal (including a camera, microphone, and sensors) that collects emotion data, and the user's device. The software used includes a generative AI model, an emotion analysis algorithm, streaming server software, and a user interface.

[1262] Singers record their own voice data on the device and upload the data to the server. The server receives this voice data, stores it in a database, and integrates it into the generative AI model. For example, a singer can record themselves singing specific lyrics on the device and send it to the server via a dedicated app. The server receives the voice data and stores it as replicated voice data. When integrated into the generative AI model, it is used as new voice data added to the existing voice dataset.

[1263] The server then manages a music database, which stores music metadata (such as title, artist name, and album name) and audio files in an organized manner. When a user searches for a song, the server searches the database based on the search query and returns the relevant songs. For example, if a user enters a prompt such as "Add a new song to the server," the server adds the new song's metadata and audio files to the database.

[1264] In music streaming, when a user searches for a specific song and clicks the play button, the server prepares the song data in real time and delivers it to the user's device as streaming data. The device receives the streaming data and immediately starts playing it. For example, the server can generate streaming data and provide it to the user by using a prompt such as "Please search for a specific song and play it."

[1265] Regarding the recording of play counts and revenue calculation, the server records the number of plays in real time each time each song is played. Based on this data, revenue is calculated and fairly distributed to the parties according to specific rules. The revenue calculation takes into account the number of plays, song information, related contract information, etc. For example, upon the prompt "Please calculate and distribute revenue based on the number of song plays," the server will perform the revenue calculation and save the result in the database.

[1266] When using an emotion engine, the user's device uses a camera, microphone, and sensors to collect the user's facial, voice, and physiological data. The collected emotion data is sent to a server and analyzed using an emotion analysis algorithm. Based on the analysis results, the server recommends music that best suits the user's emotional state. For example, a prompt such as "Analyze the user's emotion data and recommend music that matches that emotion" allows the server to analyze the emotion data and recommend appropriate music.

[1267] As described above, this invention provides a music streaming system that uses replicated voice data to recommend songs based on the user's emotions, enabling users to enjoy a personalized music experience. Furthermore, fairness in revenue distribution is maintained, making this a system that benefits all parties involved.

[1268] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1269] Step 1:

[1270] The terminal provides a user interface for singers to record their voice data using an input device. The user presses the record button, and after recording is completed, clicks the upload button. The input is the recorded voice data, and the output is an upload request to the server.

[1271] Step 2:

[1272] The server receives the voice data sent from the terminal and performs pre-processing such as data formatting and format conversion.The server then stores the formatted voice data in a voice database.The input is the uploaded voice data, and the output is the voice data stored in the database.

[1273] Step 3:

[1274] The server adds the stored voice data to a training dataset for integration into a generative AI model. This training dataset is used to generate replicated voices. The input is the voice data stored in the database, and the output is the training data integrated into the generative AI model.

[1275] Step 4:

[1276] A user enters a specific song title into the search bar of their device and presses the search button. A song search request is sent from the device to the server. The input is the search query entered by the user, and the output is the song search request.

[1277] Step 5:

[1278] The server searches the music database based on the user's search query, retrieves the corresponding song data, and returns the retrieved song data to the device. The input is the search query, and the output is the corresponding song data.

[1279] Step 6:

[1280] The device receives the music data sent from the server and launches the streaming player. When the user clicks the play button, the music starts playing. The input is the music data from the server, and the output is the music being played.

[1281] Step 7:

[1282] The server records the number of times a song is played in real time. This play count data is later used to calculate revenue. The input is the song play event, and the output is the recorded play count data.

[1283] Step 8:

[1284] The server uses a revenue calculation algorithm to calculate revenue based on the number of times a song is played and distribute it to the parties involved. The input is the recorded number of times a song is played and the revenue distribution rule, and the output is the revenue data distributed to each party.

[1285] Step 9:

[1286] The device uses sensors, cameras, and microphones to collect emotional data while the user is listening to music. The collected emotional data is sent to a server. The input is the user's physiological response, and the output is the collected emotional data.

[1287] Step 10:

[1288] The server analyzes the collected emotional data using an emotion analysis algorithm to identify the user's current emotional state, where the input is the collected emotional data and the output is the analyzed emotional state.

[1289] Step 11:

[1290] The server recommends the best songs to the user based on the analyzed emotional state. The input is the analyzed emotional state, and the output is a list of recommended songs.

[1291] Step 12:

[1292] The server stores the user's emotion history in a database and uses it to analyze long-term emotion patterns. The input is the analyzed emotion data, and the output is the emotion history stored in the database.

[1293] (Application example 2)

[1294] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1295] Traditional music streaming services lack personalized song recommendations based on the user's emotional state and new music experiences using AI-generated singers. This makes it difficult to respond to diverse user emotions and preferences, and has made it difficult to improve user satisfaction and provide novel music experiences. Another problem is the lack of transparency regarding revenue distribution to music creators.

[1296] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1297] In this invention, the server includes means for registering copied voice data, means for searching a music database, means for streaming music data, means for recording the number of plays, means for calculating and distributing revenue, means for collecting user emotion data, means for analyzing the emotion data to identify the user's emotional state, means for recommending songs based on the identified emotional state, and means for storing and analyzing the emotion history over time. This enables personalized song recommendations based on the user's emotional state, improving user satisfaction and providing a novel music experience. It also enables transparent and fair revenue distribution to music creators.

[1298] "Duplicated voice data" refers to voice data that reproduces a singer's voice using generative AI technology.

[1299] A "music database" is a system that systematically stores information about multiple songs.

[1300] "Streaming" is a technology that distributes stored music data in real time over the Internet.

[1301] "Play count" is data that records the number of times a particular song has been played by a user.

[1302] "Revenue sharing" is the process of appropriately distributing revenue earned based on the number of times a song is played among the parties involved.

[1303] "Emotional data" is data that indicates the user's current emotional state, and is mainly collected from voice, visual information, physiological data, and the like.

[1304] "Emotion analysis" is the process of identifying a user's emotional state based on collected emotional data.

[1305] "Music recommendation" is the act of suggesting the most suitable music based on the analyzed emotional state of the user.

[1306] "Emotion history" is data that records changes in the user's emotional state over a long period of time.

[1307] An "AI model" is a data model that is optimized for a specific task using machine learning algorithms.

[1308] The present invention relates to a music streaming platform using replicated voice data and an emotion engine, and can be implemented by applying the following elements and means:

[1309] 1. System Configuration

[1310] Hardware

[1311] Camera: Captures the user's facial expressions.

[1312] Microphone: Collects the user's voice and physiological data.

[1313] Smartphone: Used as an application execution environment.

[1314] software

[1315] OpenCV: Used as an image processing library to acquire camera input and perform facial expression recognition.

[1316] Librosa: A library for music analysis. It extracts features such as tempo from music data.

[1317] TensorFlow / Keras: Load and run the emotion engine model.

[1318] Requests: Used to send HTTP requests to the music recommendation service.

[1319] 2. System Operation

[1320] The server implements various functions using the following means:

[1321] Registering duplicated voice data

[1322] Singers upload their own voice data via a terminal. The server receives this data and stores it in a database. The server then integrates this data into an AI model and generates new songs using generative AI technology.

[1323] Music database management

[1324] The server searches the music database based on the user's search query, retrieves the relevant songs, and provides them to the user. It also adds new songs and updates existing songs.

[1325] Streaming music data

[1326] When a user searches for and plays a song, the server quickly streams the corresponding music data to the device, allowing the user to enjoy music in real time.

[1327] Recording views and calculating revenue

[1328] The server records the number of plays each time a song is played, and calculates revenue based on this data, which is then distributed equally among all parties involved (singers, lyricists, composers, etc.).

[1329] Use of emotion engine

[1330] The server collects users' emotional data and analyzes it using an emotion engine. It recommends songs based on the user's emotional state and analyzes their emotional history over time to provide a more personalized music experience.

[1331] 3. Specific Examples

[1332] Emotion data collection and analysis

[1333] While the user is listening to music, the device's camera and microphone work together to capture the user's facial expressions and vocal characteristics, which are then analyzed by the emotion engine to determine the user's emotional state.

[1334] Song recommendations

[1335] The server recommends songs based on the identified emotional state, for example, if the user is in a sad state, the emotion engine selects and recommends comforting songs.

[1336] Example: Using a prompt statement

[1337] The implicit knowledge is that when the user is in a sad state, the prompt sentence is:

[1338] "The user is in a sad state, so please recommend some songs that are slow and comforting."

[1339] Long-term emotion history analysis

[1340] The server stores the user's emotional history in a database and analyzes it over the long term, allowing it to understand the user's emotional patterns and reflect them in future song recommendations.

[1341] By combining these elements and methods, it becomes possible to provide a personalized music streaming service that responds to the user's emotional state, thereby improving user satisfaction and providing a novel music experience.

[1342] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1343] Step 1:

[1344] The user launches a music application.

[1345] Input: User launches application.

[1346] Output: The application home screen is displayed.

[1347] Step 2:

[1348] The user uploads the singer's voice data.

[1349] Input: User selection and upload of voice data.

[1350] Output: Duplicate voice data stored on the server.

[1351] Specific behavior:

[1352] The terminal receives selected voice data from the user.

[1353] The received data is sent to the server and stored in a database.

[1354] The voice data is used as training data to integrate into generative AI models.

[1355] Step 3:

[1356] The server searches the music database.

[1357] Input: A search query by the user.

[1358] Output: A list of songs as search results.

[1359] Specific behavior:

[1360] The user types a query into the search bar and sends it to the server.

[1361] The server searches the music database based on the received query and retrieves the corresponding songs.

[1362] Step 4:

[1363] The user plays a song.

[1364] Input: User clicks the play button.

[1365] Output: The streaming data to be played.

[1366] Specific behavior:

[1367] The server delivers streaming data of the selected music to the terminal.

[1368] The terminal plays the received streaming data, providing the user with a musical experience.

[1369] Step 5:

[1370] The device collects the user's emotional data.

[1371] Input: User's facial and voice data.

[1372] Output: Collected emotion data.

[1373] Specific behavior:

[1374] The device's camera and microphone are activated to capture the user's facial expressions and voice.

[1375] The acquired data is preprocessed and sent to the server.

[1376] Step 6:

[1377] The server analyzes the emotional data to determine the user's emotional state.

[1378] Input: Collected emotion data.

[1379] Output: The identified emotional state of the user.

[1380] Specific behavior:

[1381] The server analyzes the emotion data using machine learning algorithms.

[1382] An emotion engine identifies the user's current emotional state.

[1383] Step 7:

[1384] The server recommends songs based on the identified emotional state.

[1385] Input: The identified emotional state of the user.

[1386] Output: A list of recommended songs.

[1387] Specific behavior:

[1388] The server searches a music database for songs that correspond to the identified emotional state.

[1389] A list of recommended songs is presented to the user.

[1390] Step 8:

[1391] The server records the number of plays and calculates the revenue.

[1392] Input: Song playback event.

[1393] Output: Calculated revenue and distribution results.

[1394] Specific behavior:

[1395] The server records the number of times a song is played in a database each time it is played.

[1396] Revenue is calculated based on the number of plays and distributed to each party.

[1397] Step 9:

[1398] The server stores emotion history and analyzes it over the long term.

[1399] Input: Identified emotional state.

[1400] Output: The stored emotion history and its analysis results.

[1401] Specific behavior:

[1402] The server stores the history of the user's emotional state in a database.

[1403] Conduct long-term analysis to understand user sentiment patterns.

[1404] Through the above steps, a music streaming experience based on the user's emotional state can be provided.

[1405] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1406] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1407] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1408] [Fourth embodiment]

[1409] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1410] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1411] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1412] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1413] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1414] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1415] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1416] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1417] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1418] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1419] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1420] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1421] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1422] The present invention relates to a music streaming platform that uses singer voices replicated using generative AI technology. The system has the following components and functions:

[1423] 1. Registering duplicated voice data

[1424] Singers upload their own voice data via a terminal. The server receives this data and stores it in a database. This data is then integrated into an AI model and used as training data to be reflected in the generated music.

[1425] 2. Music database management

[1426] A music database is a system that stores song information in an organized manner. The server searches the music database based on a user's search query, retrieves the relevant songs, and provides them to the user. It also adds new songs and updates existing songs as needed.

[1427] 3. Streaming music data

[1428] When a user searches for and plays a song, the server quickly streams the corresponding music data to the device, allowing the user to enjoy music in real time. The device then plays the received streaming data, providing the user with a music experience.

[1429] 4. Recording Plays and Calculating Revenue

[1430] The server records the number of plays each time a song is played. Based on this play count data, revenue for each song is calculated and distributed appropriately. Specific rules and algorithms are used to calculate revenue, and the revenue is distributed fairly to each party involved (singer, lyricist, composer, etc.).

[1431] Specific examples

[1432] As an example, consider a scenario in which user A wants to listen to a new song that has been created using the voice of a particular singer B.

[1433] 1. Uploading voice data

[1434] Singer B uploads his / her voice data to the server via his / her device, which receives this data and integrates it into the AI ​​model.

[1435] 2. Music Generation

[1436] Composer C creates a new song and requests to use the voice of singer B. The server uses the AI ​​model to generate the new song in singer B's voice.

[1437] 3. Search and play songs

[1438] User A searches for new songs through a device. The device sends a query to the server, which searches the music database and finds the relevant new songs. The server streams the new songs and delivers them to the device. The device plays the received streaming data, allowing User A to listen to the new songs.

[1439] 4. Playback count and revenue sharing

[1440] The server records the number of plays each time User A plays a new song. Based on the play count data, the server calculates and distributes revenue to Singer B, Composer C, and other parties.

[1441] This system will ensure that singers who use AI-generated voices receive fair revenue, and that all creators involved receive appropriate recognition and compensation. This will improve fairness and transparency throughout the music industry, and promote greater creativity and innovation.

[1442] The processing flow will be explained below.

[1443] 1. Registering duplicated voice data

[1444] Processing Steps

[1445] Step 1:

[1446] The terminal accepts input from the user (singer). The singer selects their own voice data through the terminal and performs the upload operation.

[1447] Step 2:

[1448] The device sends voice data to the server, which receives and temporarily stores the data.

[1449] Step 3:

[1450] The server validates the voice data it receives, checking that the data format and quality meet the standards.

[1451] Step 4:

[1452] The server stores the verified voice data in a voice database, where it associates it with the singer's ID.

[1453] Step 5:

[1454] The server integrates the stored voice data into the AI ​​model, adding new voice data as training data for future music generation.

[1455] 2. Music database management

[1456] Processing Steps

[1457] Step 1:

[1458] The device accepts the user's request to upload music. The user (artist) selects their own music data through the device and performs the upload operation.

[1459] Step 2:

[1460] The device sends the music data to the server, which analyzes the received music data and temporarily stores it.

[1461] Step 3:

[1462] The server validates the song data, checking file integrity, format, and data completeness.

[1463] Step 4:

[1464] The server stores the verified music data in a music database, and associates the music data with artist information and metadata.

[1465] Step 5:

[1466] The server regularly backs up the music database to ensure data safety.

[1467] 3. Streaming music data

[1468] Processing Steps

[1469] Step 1:

[1470] The device accepts the user's song search request. The user enters a search query and clicks the search button.

[1471] Step 2:

[1472] The device sends a search query to the server, which receives it and searches its music database.

[1473] Step 3:

[1474] The server generates the search results, creates a list of matching songs, and returns it to the device.

[1475] Step 4:

[1476] The terminal displays the search results to the user.

[1477] Step 5:

[1478] The user selects the song they want to play and clicks the play button. The device sends a playback request to the server.

[1479] Step 6:

[1480] The server generates streaming data for the specified music piece and transmits it to the terminal.

[1481] Step 7:

[1482] The terminal plays the received streaming data and provides music to the user.

[1483] 4. Recording Plays and Calculating Revenue

[1484] Processing Steps

[1485] Step 1:

[1486] The server detects when a song starts playing. Each time a request to start playing is received, it records the play count.

[1487] Step 2:

[1488] The server collects the play count data and stores it in a database.

[1489] Step 3:

[1490] The server periodically aggregates the play count data and calculates revenue, which is then distributed based on specific rules and algorithms.

[1491] Step 4:

[1492] The server reflects the revenue data in the accounts of each party (singer, lyricist, composer, etc.).

[1493] Step 5:

[1494] Each party checks their own earnings through the terminal, which retrieves the earnings data from the server and displays it to the user.

[1495] This detailed processing flow allows the system to efficiently register the duplicated voice data, stream the song, record the number of plays, and calculate and distribute revenue.

[1496] Example 1

[1497] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1498] In conventional music streaming systems, music data is provided only using existing audio data, making it difficult to generate new audio data. Furthermore, there is no system in place for managing the appropriate use of copied audio data or for revenue distribution. Therefore, there is a need to improve fairness and transparency throughout the music industry and promote new creativity and innovation.

[1499] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1500] In this invention, the server includes a means for registering the replicated audio data, a means for managing a music database, a means for streaming the music data, a means for recording the number of plays, a means for calculating and distributing revenue, and a means for using a generative AI model to generate the replicated audio data. This not only allows users to enjoy high-quality audio data, but also enables fair revenue distribution based on the usage. Furthermore, utilizing the generative AI model enables the generation of new audio data and the provision of songs based on that data, promoting new creativity and innovation in the music industry.

[1501] "Duplicated voice data" is new voice data generated based on the original voice data using a generative AI model.

[1502] A "music database" is a database for systematically storing and managing song information and music files.

[1503] "Streaming" is a technology that distributes music data in real time over the Internet, allowing it to be played instantly on the user's device.

[1504] The "number of plays" is data indicating the number of times a particular song has been played by a user.

[1505] "Revenue" refers to income generated based on the playback of music data, which must be distributed appropriately to the parties involved.

[1506] A "generative AI model" is a model that uses machine learning algorithms to generate new data based on specific input data (in this case, voice data).

[1507] "Training data" refers to the dataset used to train a generative AI model, in this case the original audio data needed to generate the replicated audio data.

[1508] "Interface" refers to the operation screen and input means that a user uses to search for and play music.

[1509] The present invention relates to a music streaming platform that uses replicated audio data using generative AI technology. The system has the following components and functions:

[1510] 1. Registering duplicated audio data

[1511] Artist users upload their audio files using their devices. The devices then send the files to the server using HTTP POST requests. The server validates the received audio data and stores it in a secure database. The server then integrates the audio data into a generative AI model (e.g., WaveNet or Tacotron2) and uses it as training data to influence the generated music.

[1512] 2. Music database management

[1513] The server stores song information and music files in a music database in an organized manner. Song information includes title, artist, album, genre, and release year. The server also searches the database based on user search queries to extract and provide matching songs. New songs are added and existing songs are updated as needed.

[1514] 3. Streaming music data

[1515] When a user searches for and plays a song through a device, the server quickly streams the selected song to the device using HTTP or RTMP protocols. The device buffers the received streaming data and plays it using the built-in music player application, providing the user with a real-time music experience.

[1516] 4. Recording Plays and Calculating Revenue

[1517] The server records the number of plays as a log each time a song is played. This log includes information such as the song ID, playback start time, playback end time, and user ID. The log data is sent to a big data platform (e.g., Hadoop or Spark) for analysis. Based on the analysis, the server uses the play count data to calculate revenue and distributes it fairly to each relevant party (artist, lyricist, composer). Pre-set rules and algorithms are applied to calculate the revenue.

[1518] Specific examples

[1519] A specific example is given below: For example, consider a scenario in which user A listens to a new song created using the voice of a particular artist.

[1520] 1. Uploading voice data

[1521] Artists record their own voices and upload the audio data via their devices to a server, which receives the data and integrates it into a generative AI model.

[1522] 2. Music Generation

[1523] A composer creates a new song and requests that the artist's voice be used, and the server uses a generative AI model to generate the new song in the artist's voice.

[1524] 3. Search and play songs

[1525] User A searches for new songs through a device. The device sends a query to the server, which searches the music database to find the new songs. The server streams the new songs, and the device plays the received streaming data, allowing User A to listen to the new songs.

[1526] 4. Playback count and revenue sharing

[1527] The server records the number of plays each time User A plays a new song. Based on this data, the server calculates revenue and distributes it to the artists, composers, and other parties involved.

[1528] This system will enable efficient management and utilization of duplicated audio data using generative AI models, enabling fair revenue distribution, improving fairness and transparency across the music industry and encouraging greater creativity and innovation.

[1529] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1530] Step 1: Upload your voice data

[1531] The user, an artist, uses a device to record his or her own voice and uploads the audio data file to the server. Specifically, the user uses the device's recording function to create an audio file (e.g., WAV or MP3 format) and sends it to the server via an upload interface. The input is the audio data file, and the output is that file being saved on the server.

[1532] Step 2: Saving and integrating audio data

[1533] The server receives the uploaded audio data and verifies its integrity and quality. After verification is complete, it stores it in a database and performs preprocessing for incorporation into the generative AI model. Specifically, it performs noise removal and sampling rate adjustment. The input is the uploaded audio data file, and the output is the preprocessed audio data stored in the database.

[1534] Step 3: Store song information

[1535] The server adds new and updated song information to the music database. Song information includes title, artist, album, genre, and release year. Specifically, it creates database entries based on information provided by administrators and stakeholders. The input is song information and music files, and the output is that they are accurately stored in the database.

[1536] Step 4: Search for songs

[1537] The server uses the search query received from the user to search the music database and extract the corresponding songs. Specifically, it uses a full-text search engine (e.g., ElasticSearch) to quickly search for songs that match the query. The input is the search query, and the output is a list of songs as search results.

[1538] Step 5: Request a song to play

[1539] When a user uses a terminal to select a particular song, the terminal sends the selection to the server. The input is the user's song selection, and the output is a play request sent to the server.

[1540] Step 6: Stream your music

[1541] The server receives playback requests from users and streams the specified music. Specifically, it sends music files to terminals at a certain bit rate via HTTP or RTMP protocol. The input is the playback request, and the output is the streaming data.

[1542] Step 7: Playing a song

[1543] The device buffers the streaming data received from the server and plays it using the built-in music player application. The input is streaming data and the output is audio playback.

[1544] Step 8: Record play counts

[1545] The server records a playback event as a log each time a song is played. This log includes the song ID, playback start time, playback end time, user ID, etc. The input is the playback event information, and the output is the recorded log.

[1546] Step 9: Calculate and distribute revenue

[1547] The server analyzes the play count log and calculates revenue based on the number of plays of a specific song. This calculation is based on pre-defined rules and algorithms. Based on the analysis results, the revenue is distributed to the relevant parties (artist, lyricist, composer). The input is the play count log, and the output is the calculated revenue and its distribution information.

[1548] This enables the system to efficiently manage and utilize duplicated audio data using generative AI models, enabling fair revenue distribution.

[1549] (Application example 1)

[1550] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1551] On today's music distribution platforms, it is difficult for users to create and enjoy songs using the voice of a specific artist, and there is a demand for proper management of the number of plays and distribution of revenue for the created songs. Therefore, it is necessary to create songs using copied voices and provide them to users while realizing fair and transparent revenue distribution.

[1552] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1553] In this invention, the server includes a means for registering the copied audio data, a means for searching a music database, a means for streaming the music data, a means for users to generate music based on specific requests using a generative AI model, a means for recording the number of plays, and a means for calculating and distributing revenue. This allows users to freely generate new music using artists' voices and enjoy them in real time, while also enabling fair and transparent revenue distribution to artists and related parties.

[1554] "Duplicated audio data" refers to audio data that has been duplicated or generated using artificial intelligence technology from data recorded from the original artist's voice.

[1555] A "music database" is a collection of digital data that organizes and stores information about songs, and can be accessed by users through searches.

[1556] "Music Data" means data, including audio and metadata, of a musical composition stored in digital format.

[1557] "Streaming" is a method of playing music data in real time over the Internet.

[1558] "Plays" is a record of the number of times a particular song has been played by a user.

[1559] "Revenue sharing" is the process of fairly distributing revenues earned as a result of the playback of a generated song among the parties involved.

[1560] A "generative AI model" is an algorithm or system that uses artificial intelligence technology to generate new audio data or music.

[1561] The "user interface" is the part of the software that provides an operation screen for the user to search for and play music.

[1562] This invention provides a music streaming platform that uses duplicated audio using generative AI technology. The system includes functions such as registration of duplicated audio data, management of music data, streaming, recording of play counts, and revenue sharing.

[1563] System Configuration

[1564] 1. Server:

[1565] Audio data registration:

[1566] The server receives the voice data uploaded via the device and stores it in a database. Voice providers provide their own voice, and the server integrates this data into the AI ​​model.

[1567] Music database management:

[1568] The server manages a database that systematically stores song information, receives search queries from users, searches for and provides matching songs, and also adds and updates new songs.

[1569] Streaming:

[1570] The server streams the music selected by the user in real time and delivers it to the device, allowing the user to enjoy the music instantly.

[1571] Playback Tracking and Revenue Sharing:

[1572] The server records the number of times a song is played and calculates and distributes revenue based on that number, using specific rules and algorithms to ensure fair distribution among all parties involved.

[1573] Use of generative AI models:

[1574] It receives prompts based on the user's request and uses a generative AI model to generate new music, which is then instantly added to a database and made available for streaming.

[1575] 2. Device (smartphone):

[1576] Upload audio data:

[1577] The terminal provides an interface for the user to record the voice of the voice provider and upload it to the server.

[1578] Search and play songs:

[1579] The device provides an interface for users to search for and play music. Users can search for music within the app and enjoy music by pressing the play button.

[1580] Music generation using prompts:

[1581] The device provides an interface for users to input specific prompts to generate new music, which are then sent to a server and processed by a generative AI model.

[1582] Hardware and software used

[1583] The hardware includes servers (cloud-based servers or dedicated servers) and user devices (smartphones or tablets). As for software, the server side runs a web application based on Flask and a generative AI model using TensorFlow. The client side runs a smartphone application that provides the user interface.

[1584] Specific examples

[1585] For example, if a user wants to generate a new pop song using the voice of a particular singer, they might enter the prompt text as follows:

[1586] "Generate love songs with pop rhythms using singer's voices"

[1587] By passing this prompt to a generative AI model, a new pop song using the specified voice is generated and instantly added to the database, where users can then search for and stream the new song.

[1588] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1589] Step 1:

[1590] The user uses a device to record the voice of the voice provider and uploads the voice data from the device to the server. The server receives this voice data and stores it in a database as training data for the generative AI model. The input is the voice data, and the output is the voice data stored in the database.

[1591] Step 2:

[1592] The server integrates the received voice data into the generative AI model. The data processing performed here involves converting the voice data into an appropriate format and training the model so that it can generate new voices. The input is the voice data, and the output is updating the generative AI model.

[1593] Step 3:

[1594] A user inputs a search query for a song through a terminal. The terminal sends the query to a server, which then searches a music database to find the relevant song. The input is a search query, and the output is a list of songs as search results.

[1595] Step 4:

[1596] The user selects a specific song from the search results and instructs it to be played. The server streams the corresponding music data and delivers it to the device. The device receives this streaming data and plays it for the user. The input is the song ID, and the output is the song data played in real time.

[1597] Step 5:

[1598] The server records the user's playback actions and updates the database with the number of times the song has been played. This process takes the user ID and song ID as input, increments the number of times the song has been played, and updates the database. The input is the user ID and song ID, and the output is the updated number of times the song has been played.

[1599] Step 6:

[1600] The user uses a terminal to input a specific prompt sentence and request the generation of a new song. The server receives this prompt sentence and passes it to the generative AI model, which then generates a new song. The input is the prompt sentence, and the output is the generated song data. As a concrete example, we use a prompt sentence such as "Generate a love song with a pop rhythm and a singer's voice."

[1601] Step 7:

[1602] The server registers the generated new music in a database, allowing users to search and play it. The input is the generated music data, and the output is the new music registered in the database.

[1603] Step 8:

[1604] The server calculates and distributes revenue. It calculates revenue based on the number of views using specific rules and algorithms and distributes it to the parties involved. The input is the number of views, and the output is the calculated revenue and distribution results.

[1605] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1606] The present invention combines an emotion engine with a music streaming platform that uses singer voices replicated using generative AI technology. The system has the following components and functions:

[1607] 1. Registering duplicated voice data

[1608] Singers upload their own voice data via a terminal. The server receives this data and stores it in a database. This data is then integrated into an AI model and used as training data to be reflected in the generated music.

[1609] 2. Music database management

[1610] A music database is a system that stores song information in an organized manner. The server searches the music database based on a user's search query, retrieves the relevant songs, and provides them to the user. It also adds new songs and updates existing songs as needed.

[1611] 3. Streaming music data

[1612] When a user searches for and plays a song, the server quickly streams the corresponding music data to the device, allowing the user to enjoy music in real time. The device then plays the received streaming data, providing the user with a music experience.

[1613] 4. Recording Plays and Calculating Revenue

[1614] The server records the number of plays each time a song is played. Based on this play count data, revenue for each song is calculated and distributed appropriately. Specific rules and algorithms are used to calculate revenue, and the revenue is distributed fairly to each party involved (singer, lyricist, composer, etc.).

[1615] 5. Use of Emotion Engine

[1616] The system is equipped with an emotion engine for recognizing a user's emotions, which analyzes the user's voice, visual data, text input, and other physiological data to identify the user's current emotional state.

[1617] 1. Collecting Emotional Data

[1618] The device collects emotional data while the user listens to the music being played, including facial expression recognition using a camera, voice analysis using a microphone, and physiological data acquisition using sensors.

[1619] 2. Emotion Analysis

[1620] The server analyzes the collected emotion data to identify the user's emotional state. The emotion engine uses machine learning algorithms to identify the user's emotion based on the collected data.

[1621] 3. Song Recommendations

[1622] The server recommends songs to the user based on the identified emotional state, for example, if the user is in a sad state, the emotion engine selects and suggests comforting songs to the user.

[1623] 4. Emotion history storage and analysis

[1624] The server stores the user's emotional history in a database and performs long-term analysis, which allows it to understand the user's emotional patterns and provide a more personalized music experience.

[1625] Specific examples

[1626] 1. Uploading voice data and integrating AI models

[1627] The terminal accepts input for uploading singer's voice data.

[1628] The server receives the voice data, stores it, and integrates it into an AI model, which then uses it to generate new songs.

[1629] 2. Search and play songs

[1630] The user searches for a new song and clicks the play button.

[1631] The server searches for the music, generates streaming data, and delivers it to the device.

[1632] The terminal receives the streaming data and plays it back to the user.

[1633] 3. Recording of play counts and revenue sharing

[1634] The server records the number of times the song is played and calculates the revenue.

[1635] Profits are distributed among the parties according to certain rules.

[1636] 4. Utilizing the Emotion Engine

[1637] While the user is listening to music, the device collects emotion data.

[1638] The server uses an emotion engine to analyze the user's emotions and recommend appropriate songs.

[1639] It analyzes emotional history over the long term to provide users with the optimal music experience.

[1640] This system ensures that singers who use AI-replicated voices receive fair revenue, and that all creators involved receive appropriate recognition and compensation. Furthermore, the emotional engine can provide users with a more personalized music experience. This will improve fairness and transparency throughout the music industry, promoting greater creativity and innovation.

[1641] The processing flow will be explained below.

[1642] 1. Registering duplicated voice data

[1643] Processing Steps

[1644] Step 1:

[1645] The terminal accepts input from the user (singer). The singer selects his / her own voice data via the terminal and performs the upload operation.

[1646] Step 2:

[1647] The device sends voice data to the server, which receives and temporarily stores the data.

[1648] Step 3:

[1649] The server validates the voice data it receives, checking that the format and quality of the voice data meets standards.

[1650] Step 4:

[1651] The server stores the verified voice data in a voice database, and when it is saved, it associates the voice data with the singer's ID.

[1652] Step 5:

[1653] The server integrates the stored voice data into the AI ​​model, adding new voice data as training data for future music generation.

[1654] 2. Music database management

[1655] Processing Steps

[1656] Step 1:

[1657] The device accepts a song upload request from the user (artist). The artist selects their own song data through the device and performs the upload operation.

[1658] Step 2:

[1659] The device sends the music data to the server, which analyzes the received music data and temporarily stores it.

[1660] Step 3:

[1661] The server validates the song data, checking its integrity, format, and quality.

[1662] Step 4:

[1663] The server stores the verified music data in a music database, and associates the music data with artist information and metadata.

[1664] Step 5:

[1665] The server regularly backs up the music database to ensure data safety and availability.

[1666] 3. Streaming music data

[1667] Processing Steps

[1668] Step 1:

[1669] The device accepts the user's song search request. The user enters a search query and clicks the search button.

[1670] Step 2:

[1671] The device sends a search query to the server, which receives it and searches its music database.

[1672] Step 3:

[1673] The server generates the search results, creating a list of matching songs and sending it back to the device.

[1674] Step 4:

[1675] The terminal displays the search results to the user.

[1676] Step 5:

[1677] The user selects the song they want to play and clicks the play button. The device sends a playback request to the server.

[1678] Step 6:

[1679] The server generates streaming data for the specified music piece and transmits it to the terminal.

[1680] Step 7:

[1681] The terminal plays the received streaming data and provides music to the user.

[1682] 4. Recording Plays and Calculating Revenue

[1683] Processing Steps

[1684] Step 1:

[1685] The server detects when a song starts playing. Each time a request to start playing is received, it records the play count.

[1686] Step 2:

[1687] The server collects play count data and stores it in a database.

[1688] Step 3:

[1689] The server periodically aggregates the play count data and calculates revenue, which is then distributed based on rules and algorithms.

[1690] Step 4:

[1691] The server reflects the revenue data in the accounts of each party (singer, lyricist, composer, etc.).

[1692] Step 5:

[1693] Each party checks their own earnings through the terminal, which retrieves the earnings data from the server and displays it to the user.

[1694] 5. Use of Emotion Engine

[1695] Processing Steps

[1696] Step 1:

[1697] The device collects the user's emotional data. While the user is playing music, it uses a camera to recognize facial expressions, a microphone to analyze voices, and sensors to acquire physiological data.

[1698] Step 2:

[1699] The device sends the collected emotion data to a server, which receives it and temporarily stores it.

[1700] Step 3:

[1701] The server analyzes the emotion data, and the emotion engine uses machine learning algorithms to identify the user's emotion based on the collected data.

[1702] Step 4:

[1703] The server recommends songs based on the analyzed emotional data, suggesting songs that best suit the user's state of mind.

[1704] Step 5:

[1705] The server stores the user's emotional history in a database and performs long-term analysis, analyzing the user's emotional patterns to provide a more personalized music experience.

[1706] This detailed processing flow allows the system to efficiently register the duplicated voice data, stream the songs, record the number of plays, calculate and distribute revenue, and even recommend songs using an emotion engine.

[1707] Example 2

[1708] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1709] While conventional music streaming systems can play songs and distribute revenue, they lack the functionality to recommend songs based on user emotions or replicate artists' voices using AI. This makes it difficult to provide a music experience tailored to each user's emotions, and also poses the problem of insufficient fairness in revenue distribution for songs that use replicated voices.

[1710] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1711] In this invention, the server includes means for registering the duplicated voice data, means for searching a music database, means for streaming the music data, means for recording the number of plays, means for calculating and distributing revenue, means for collecting user emotion data, and means for analyzing the collected emotion data and recommending songs to the user. This makes it possible to provide a personalized music experience according to the user's emotion while maintaining fairness in the distribution of revenue from songs based on the duplicated voice using AI.

[1712] "Replica voice data" refers to audio information that has been recorded from an artist's voice and reproduced using a generative AI model.

[1713] A "music database" is an information system that efficiently manages metadata and audio files related to music, allowing them to be searched and stored.

[1714] "Streaming" is a technology that distributes and plays music data in real time to a user's terminal via the Internet.

[1715] "Plays" is an indicator that records the number of times a particular song has been played by a user.

[1716] "Revenue sharing" is the process of fairly distributing revenue from music content to stakeholders based on data such as the number of plays.

[1717] "Emotion data" is data that indicates the user's current emotional state, collected from the user's facial expressions, voice, physiological data, and the like.

[1718] A "generative AI model" is a collection of algorithms that perform specific tasks or generate results based on large amounts of data, and in this case refers specifically to models used to reproduce an artist's voice.

[1719] "Recommendation" refers to the act of selecting the most suitable music piece based on the analyzed emotional data of the user and suggesting it to the user.

[1720] This invention provides a personalized music experience for users by integrating an emotion engine into a music streaming system that uses replicated voice data. The main hardware used includes a database server, a terminal (including a camera, microphone, and sensors) that collects emotion data, and the user's device. The software used includes a generative AI model, an emotion analysis algorithm, streaming server software, and a user interface.

[1721] Singers record their own voice data on the device and upload the data to the server. The server receives this voice data, stores it in a database, and integrates it into the generative AI model. For example, a singer can record themselves singing specific lyrics on the device and send it to the server via a dedicated app. The server receives the voice data and stores it as replicated voice data. When integrated into the generative AI model, it is used as new voice data added to the existing voice dataset.

[1722] The server then manages a music database, which stores music metadata (such as title, artist name, and album name) and audio files in an organized manner. When a user searches for a song, the server searches the database based on the search query and returns the relevant songs. For example, if a user enters a prompt such as "Add a new song to the server," the server adds the new song's metadata and audio files to the database.

[1723] In music streaming, when a user searches for a specific song and clicks the play button, the server prepares the song data in real time and delivers it to the user's device as streaming data. The device receives the streaming data and immediately starts playing it. For example, the server can generate streaming data and provide it to the user by using a prompt such as "Please search for a specific song and play it."

[1724] Regarding the recording of play counts and revenue calculation, the server records the number of plays in real time each time each song is played. Based on this data, revenue is calculated and fairly distributed to the parties according to specific rules. The revenue calculation takes into account the number of plays, song information, related contract information, etc. For example, upon the prompt "Please calculate and distribute revenue based on the number of song plays," the server will perform the revenue calculation and save the result in the database.

[1725] When using an emotion engine, the user's device uses a camera, microphone, and sensors to collect the user's facial, voice, and physiological data. The collected emotion data is sent to a server and analyzed using an emotion analysis algorithm. Based on the analysis results, the server recommends music that best suits the user's emotional state. For example, a prompt such as "Analyze the user's emotion data and recommend music that matches that emotion" allows the server to analyze the emotion data and recommend appropriate music.

[1726] As described above, this invention provides a music streaming system that uses replicated voice data to recommend songs based on the user's emotions, enabling users to enjoy a personalized music experience. Furthermore, fairness in revenue distribution is maintained, making this a system that benefits all parties involved.

[1727] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1728] Step 1:

[1729] The terminal provides a user interface for singers to record their voice data using an input device. The user presses the record button, and after recording is completed, clicks the upload button. The input is the recorded voice data, and the output is an upload request to the server.

[1730] Step 2:

[1731] The server receives the voice data sent from the terminal and performs pre-processing such as data formatting and format conversion.The server then stores the formatted voice data in a voice database.The input is the uploaded voice data, and the output is the voice data stored in the database.

[1732] Step 3:

[1733] The server adds the stored voice data to a training dataset for integration into a generative AI model. This training dataset is used to generate replicated voices. The input is the voice data stored in the database, and the output is the training data integrated into the generative AI model.

[1734] Step 4:

[1735] A user enters a specific song title into the search bar of their device and presses the search button. A song search request is sent from the device to the server. The input is the search query entered by the user, and the output is the song search request.

[1736] Step 5:

[1737] The server searches the music database based on the user's search query, retrieves the corresponding song data, and returns the retrieved song data to the device. The input is the search query, and the output is the corresponding song data.

[1738] Step 6:

[1739] The device receives the music data sent from the server and launches the streaming player. When the user clicks the play button, the music starts playing. The input is the music data from the server, and the output is the music being played.

[1740] Step 7:

[1741] The server records the number of times a song is played in real time. This play count data is later used to calculate revenue. The input is the song play event, and the output is the recorded play count data.

[1742] Step 8:

[1743] The server uses a revenue calculation algorithm to calculate revenue based on the number of times a song is played and distribute it to the parties involved. The input is the recorded number of times a song is played and the revenue distribution rule, and the output is the revenue data distributed to each party.

[1744] Step 9:

[1745] The device uses sensors, cameras, and microphones to collect emotional data while the user is listening to music. The collected emotional data is sent to a server. The input is the user's physiological response, and the output is the collected emotional data.

[1746] Step 10:

[1747] The server analyzes the collected emotional data using an emotion analysis algorithm to identify the user's current emotional state, where the input is the collected emotional data and the output is the analyzed emotional state.

[1748] Step 11:

[1749] The server recommends the best songs to the user based on the analyzed emotional state. The input is the analyzed emotional state, and the output is a list of recommended songs.

[1750] Step 12:

[1751] The server stores the user's emotion history in a database and uses it to analyze long-term emotion patterns. The input is the analyzed emotion data, and the output is the emotion history stored in the database.

[1752] (Application example 2)

[1753] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1754] Traditional music streaming services lack personalized song recommendations based on the user's emotional state and new music experiences using AI-generated singers. This makes it difficult to respond to diverse user emotions and preferences, and has made it difficult to improve user satisfaction and provide novel music experiences. Another problem is the lack of transparency regarding revenue distribution to music creators.

[1755] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1756] In this invention, the server includes means for registering copied voice data, means for searching a music database, means for streaming music data, means for recording the number of plays, means for calculating and distributing revenue, means for collecting user emotion data, means for analyzing the emotion data to identify the user's emotional state, means for recommending songs based on the identified emotional state, and means for storing and analyzing the emotion history over time. This enables personalized song recommendations based on the user's emotional state, improving user satisfaction and providing a novel music experience. It also enables transparent and fair revenue distribution to music creators.

[1757] "Duplicated voice data" refers to voice data that reproduces a singer's voice using generative AI technology.

[1758] A "music database" is a system that systematically stores information about multiple songs.

[1759] "Streaming" is a technology that distributes stored music data in real time over the Internet.

[1760] "Play count" is data that records the number of times a particular song has been played by a user.

[1761] "Revenue sharing" is the process of appropriately distributing revenue earned based on the number of times a song is played among the parties involved.

[1762] "Emotional data" is data that indicates the user's current emotional state, and is mainly collected from voice, visual information, physiological data, and the like.

[1763] "Emotion analysis" is the process of identifying a user's emotional state based on collected emotional data.

[1764] "Music recommendation" is the act of suggesting the most suitable music based on the analyzed emotional state of the user.

[1765] "Emotion history" is data that records changes in the user's emotional state over a long period of time.

[1766] An "AI model" is a data model that is optimized for a specific task using machine learning algorithms.

[1767] The present invention relates to a music streaming platform using replicated voice data and an emotion engine, and can be implemented by applying the following elements and means:

[1768] 1. System Configuration

[1769] Hardware

[1770] Camera: Captures the user's facial expressions.

[1771] Microphone: Collects the user's voice and physiological data.

[1772] Smartphone: Used as an application execution environment.

[1773] software

[1774] OpenCV: Used as an image processing library to acquire camera input and perform facial expression recognition.

[1775] Librosa: A library for music analysis. It extracts features such as tempo from music data.

[1776] TensorFlow / Keras: Load and run the emotion engine model.

[1777] Requests: Used to send HTTP requests to the music recommendation service.

[1778] 2. System Operation

[1779] The server implements various functions using the following means:

[1780] Registering duplicated voice data

[1781] Singers upload their own voice data via a terminal. The server receives this data and stores it in a database. The server then integrates this data into an AI model and generates new songs using generative AI technology.

[1782] Music database management

[1783] The server searches the music database based on the user's search query, retrieves the relevant songs, and provides them to the user. It also adds new songs and updates existing songs.

[1784] Streaming music data

[1785] When a user searches for and plays a song, the server quickly streams the corresponding music data to the device, allowing the user to enjoy music in real time.

[1786] Recording views and calculating revenue

[1787] The server records the number of plays each time a song is played, and calculates revenue based on this data, which is then distributed equally among all parties involved (singers, lyricists, composers, etc.).

[1788] Use of emotion engine

[1789] The server collects users' emotional data and analyzes it using an emotion engine. It recommends songs based on the user's emotional state and analyzes their emotional history over time to provide a more personalized music experience.

[1790] 3. Specific Examples

[1791] Emotion data collection and analysis

[1792] While the user is listening to music, the device's camera and microphone work together to capture the user's facial expressions and vocal characteristics, which are then analyzed by the emotion engine to determine the user's emotional state.

[1793] Song recommendations

[1794] The server recommends songs based on the identified emotional state, for example, if the user is in a sad state, the emotion engine selects and recommends comforting songs.

[1795] Example: Using a prompt statement

[1796] The implicit knowledge is that when the user is in a sad state, the prompt sentence is:

[1797] "The user is in a sad state, so please recommend some songs that are slow and comforting."

[1798] Long-term emotion history analysis

[1799] The server stores the user's emotional history in a database and analyzes it over the long term, allowing it to understand the user's emotional patterns and reflect them in future song recommendations.

[1800] By combining these elements and methods, it becomes possible to provide a personalized music streaming service that responds to the user's emotional state, thereby improving user satisfaction and providing a novel music experience.

[1801] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1802] Step 1:

[1803] The user launches a music application.

[1804] Input: User launches application.

[1805] Output: The application home screen is displayed.

[1806] Step 2:

[1807] The user uploads the singer's voice data.

[1808] Input: User selection and upload of voice data.

[1809] Output: Duplicate voice data stored on the server.

[1810] Specific behavior:

[1811] The terminal receives selected voice data from the user.

[1812] The received data is sent to the server and stored in a database.

[1813] The voice data is used as training data to integrate into generative AI models.

[1814] Step 3:

[1815] The server searches the music database.

[1816] Input: A search query by the user.

[1817] Output: A list of songs as search results.

[1818] Specific behavior:

[1819] The user types a query into the search bar and sends it to the server.

[1820] The server searches the music database based on the received query and retrieves the corresponding songs.

[1821] Step 4:

[1822] The user plays a song.

[1823] Input: User clicks the play button.

[1824] Output: The streaming data to be played.

[1825] Specific behavior:

[1826] The server delivers streaming data of the selected music to the terminal.

[1827] The terminal plays the received streaming data, providing the user with a musical experience.

[1828] Step 5:

[1829] The device collects the user's emotional data.

[1830] Input: User's facial and voice data.

[1831] Output: Collected emotion data.

[1832] Specific behavior:

[1833] The device's camera and microphone are activated to capture the user's facial expressions and voice.

[1834] The acquired data is preprocessed and sent to the server.

[1835] Step 6:

[1836] The server analyzes the emotional data to determine the user's emotional state.

[1837] Input: Collected emotion data.

[1838] Output: The identified emotional state of the user.

[1839] Specific behavior:

[1840] The server analyzes the emotion data using machine learning algorithms.

[1841] An emotion engine identifies the user's current emotional state.

[1842] Step 7:

[1843] The server recommends songs based on the identified emotional state.

[1844] Input: The identified emotional state of the user.

[1845] Output: A list of recommended songs.

[1846] Specific behavior:

[1847] The server searches a music database for songs that correspond to the identified emotional state.

[1848] A list of recommended songs is presented to the user.

[1849] Step 8:

[1850] The server records the number of plays and calculates the revenue.

[1851] Input: Song playback event.

[1852] Output: Calculated revenue and distribution results.

[1853] Specific behavior:

[1854] The server records the number of times a song is played in a database each time it is played.

[1855] Revenue is calculated based on the number of plays and distributed to each party.

[1856] Step 9:

[1857] The server stores emotion history and analyzes it over the long term.

[1858] Input: Identified emotional state.

[1859] Output: The stored emotion history and its analysis results.

[1860] Specific behavior:

[1861] The server stores the history of the user's emotional state in a database.

[1862] Conduct long-term analysis to understand user sentiment patterns.

[1863] Through the above steps, a music streaming experience based on the user's emotional state can be provided.

[1864] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1865] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1866] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1867] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1868] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1869] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1870] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1871] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1872] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1873] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1874] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1875] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1876] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1877] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1878] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1879] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1880] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1881] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1882] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1883] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1884] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1885] The following is further disclosed regarding the above embodiment.

[1886] (Claim 1)

[1887] a means for registering the duplicated voice data;

[1888] a means for searching a music database;

[1889] a means for streaming music data;

[1890] a means for recording the number of plays;

[1891] a means for calculating and allocating revenues;

[1892] A system including:

[1893] (Claim 2)

[1894] Further including the means to integrate artist vocal data into the AI ​​model;

[1895] 10. The system of claim 1.

[1896] (Claim 3)

[1897] further comprising means for providing an interface for a user to search for and play songs;

[1898] 10. The system of claim 1.

[1899] "Example 1"

[1900] (Claim 1)

[1901] means for registering the replicated audio data;

[1902] a means for managing a music database;

[1903] a means for streaming music data;

[1904] a means for recording the number of plays;

[1905] a means for calculating and allocating revenues;

[1906] a means for using a generative AI model to generate replicated audio data;

[1907] A system including:

[1908] (Claim 2)

[1909] further comprising means for providing an interface for a user to search for and play songs;

[1910] 10. The system of claim 1.

[1911] (Claim 3)

[1912] and means for generating the replicated audio data by utilizing the generative AI model to create training data.

[1913] 10. The system of claim 1.

[1914] "Application Example 1"

[1915] (Claim 1)

[1916] means for registering the replicated audio data;

[1917] a means for searching a music database;

[1918] a means for streaming music data;

[1919] a means for recording the number of plays;

[1920] a means for calculating and allocating revenues;

[1921] A means for users to utilize the generative AI model to generate music based on specific requests;

[1922] A system including:

[1923] (Claim 2)

[1924] and means for integrating the speech data of the speech provider into the generative AI model.

[1925] 10. The system of claim 1.

[1926] (Claim 3)

[1927] further comprising means for providing an interface for a user to search for and play songs;

[1928] 10. The system of claim 1.

[1929] "Example 2: Combining Emotion Engines"

[1930] (Claim 1)

[1931] a means for registering the duplicated voice data;

[1932] a means for searching a music database;

[1933] a means for streaming music data;

[1934] a means for recording the number of plays;

[1935] a means for calculating and allocating revenues;

[1936] means for collecting user emotion data;

[1937] A means of analyzing the collected emotional data and recommending songs to users;

[1938] A system including:

[1939] (Claim 2)

[1940] further including means for integrating artist vocal data into the generative AI model;

[1941] 10. The system of claim 1.

[1942] (Claim 3)

[1943] further comprising means for providing an interface for a user to search for and play songs;

[1944] 10. The system of claim 1.

[1945] "Application example 2 when combining emotion engines"

[1946] (Claim 1)

[1947] a means for registering the duplicated voice data;

[1948] a means for searching a music database;

[1949] a means for streaming music data;

[1950] a means for recording the number of plays;

[1951] a means for calculating and allocating revenues;

[1952] means for collecting user emotion data;

[1953] means for analyzing the emotion data to identify the user's emotional state;

[1954] a means for recommending songs based on the identified emotional state;

[1955] A means of storing and analyzing emotion history over time;

[1956] A system including:

[1957] (Claim 2)

[1958] Further including the means to integrate artist vocal data into the AI ​​model;

[1959] 10. The system of claim 1.

[1960] (Claim 3)

[1961] further comprising means for providing an interface for a user to search for and play songs;

[1962] 10. The system of claim 1. [Explanation of symbols]

[1963] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for registering the duplicated voice data; a means for searching a music database; a means for streaming music data; a means for recording the number of plays; a means for calculating and allocating revenues; A system including:

2. Further including the means to integrate artist vocal data into the AI ​​model; The system of claim 1 .

3. further comprising means for providing an interface for a user to search for and play songs; The system of claim 1 .

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A